D-50: #1032 carried review 66 APPROVED and a merge-gate GO both bound to78ec47cd, and by the time Jason approved the merge it reported mergeable=false. Verified from the provider: its head does not contain current mainf4fd5967. Two rules we already hold, neither wrong alone, combine badly — a verdict is commit-bound and voids when the head moves, and main drifts while a PR is held, so resolving the drift moves the head and voids the verdict that was waiting to be used. Holding it across the whole #1033 saga is what did it. The asymmetry worth noting: usually a head moves because the AUTHOR pushed, so the void follows from their action. Here nobody touched the PR — unrelated progress invalidated it — which is why it is easy to miss and needs a checklist line rather than vigilance. Mos's line: a GO has a shelf life, re-gate on main-drift. RM-63 files Jason's architecture as the mechanization of D-41/RM-50/RM-58/P-LIFECYCLE rather than a new island: durably-named systemd agent services with lanes, where restart triggers a continuation and the agent resumes its lane, combined with DB-centric task assignment on the PG spine for autonomous pickup. That is the enforcement surface D-41 proved missing — a managed service can be watched (token budget, which is D-49's mechanical ceiling monitor), restarted (rotation, retiring our manual respawns) and continued (rehydrate from the DB). It closes the f10-coder ceiling gap directly: Jason confirmed that ceiling was not mechanically flagged and must be. #1032 freshening dispatched to coder-mos1 — it authored the queue guard, so that context is an asset, and by D-49 it can still take a lane at 67%. Told to stop and flag on any SUBSTANCE conflict, because a queue-guard semantic resolved quietly re-breaks the very thing RM-03 fixes. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
189 KiB
Remediation Backlog — Reconciled Execution Plan
Owner: mos-remediation (sole writer). Workers read; they never modify this file.
Sources: DECOMP-OPUS.md (robustness, 38 tasks / 8 dissents) and
DECOMP-SOL.md (pragmatic, 25 tasks / 10 defers / 7 dissents), produced
independently — neither planner read the other. Charter: MISSION.md.
Status: EXECUTING — all three blocking decisions RULED by Mos on 2026-07-31 (§5). RM-01 is
dispatched. RM-03 is held pending Jason's disposition of PR #1023; nothing else is blocked.
Provenance of the inputs (both clean).
planner-opusran in a fresh session throughout.planner-solinitially began work at 64.3% dirty context despite a brief instructing it to reset; that run was interrupted and discarded before it produced any output, the seat was reset out-of-band to 0.0%, and the brief was re-dispatched.DECOMP-SOL.mdis the product of the clean run only (it peaked at ~26% context). Both decompositions are therefore clean-context artifacts and are weighted equally here. The discarded dirty run is banked as dogfood seed D-4 and as task RM-58 — the failure it demonstrates is that asking an agent to reset is not enforcement.
1. What the two planners agreed on without collusion
Independent convergence is the strongest signal available here, because neither planner could see the other's file. Where both arrived at the same conclusion from opposite biases, I treat it as settled.
| # | Convergent finding | OPUS | SOL |
|---|---|---|---|
| C1 | The charter's wire-in point is wrong. Do NOT wire the choke point into mosaic_orchestrator.py::run_single_task. |
D2 (headline) | Dissent 1 |
| C2 | P0 hygiene/gate work must precede the spine, not follow it. | D1, phase P0 | G0, SOL-01/02 |
| C3 | No new deployable microservice; the executor is a library + the coord daemon. | implicit throughout | Dissent 3 |
| C4 | Comms adapters beyond tmux are out of scope for this mission. | D7 | DEFER 4/5/6 |
| C5 | The "100 rotations lossless" bar is a late conformance gate, not an early tax. | D4 | Dissent 7 |
| C6 | Redis is a derived hot path, never an authority; PG commits first. | R-013, R-054 | SOL-11, SOL-21 |
| C7 | Reuse packages/coord; do NOT revive the untracked apps/coordinator residue. |
R-042 | SOL-15 AC5 |
C1 is the single most consequential output of this exercise. The charter (MISSION.md) and my
kickoff instruction both name mosaic_orchestrator.py::run_single_task:126-276 as the integration
point. Both planners independently rejected it on the same evidence: that controller is
"enabled": false (.mosaic/orchestrator/config.json:2) and references a dispatcher path
(tools/macp/dispatcher/pi_runner.ts) that does not exist in this checkout. Wiring the new choke
point into a disabled rail produces a stranded executor — the identical built-but-unwired disease,
one layer up, that would look "done" in a PR. The live paths are
packages/mosaic/src/commands/launch.ts and packages/coord/src/runner.ts.
This contradicted the charter and was escalated as DECISION-1 — now RULED in the planners' favour by
Mos (§5). The corrected target is a new production Node TaskExecutor on the live dispatch path
(packages/mosaic launch + packages/coord) that Coord/Forge/live dispatch submit through; the
Python rail is deleted, not ported.
1a. ★ KEYSTONE DOGFOOD CASE — an inert gate that erased its own evidence
A merged commit shipped past pnpm format:check — and then the evidence quietly erased itself.
Verified chain (blob-level, under the repo's own prettier config, at the file's real path):
| commit | state of packages/mosaic/framework/tools/orchestrator/README.md |
|---|---|
b79336a8 — merged PR #868 |
blob 3ee7f104 — FAILS pnpm format:check |
48fd1df2 — merged PR #872 (unrelated: ci-queue-wait 404 handling) |
blob 3d3bb132 — passes; incidentally reformatted by that PR's lint-staged |
current origin/main (06e0d403) |
passes — the gate now looks green |
So: PR #868 merged a file that fails a required gate ⇒ the CI format gate did not block it. The
gate was inert for that merge. Then an unrelated later PR's pre-commit hook reformatted the file as a
side effect, so main went green again without anyone ever learning the gate had failed to fire.
Correction on record: my first report to Mos said "format:check is RED on main now." That was true of the
mainmy checkout was pinned to (b79336a8) and is no longer true of currentmain, which advanced mid-session. The inert-gate finding itself is unchanged and verified; only its present-tense framing was wrong. The hygiene PR therefore carries the.prettierignorefix only — the README needs no fix today.
Third live instance, same class — the queue guard, hit by this orchestrator
Running the mandated pre-push guard during TASK-0:
$ ~/.config/mosaic/tools/git/ci-queue-wait.sh --purpose push
[ci-queue-wait] platform=gitea purpose=push branch=main sha=06e0d403…
[ci-queue-wait] state=unknown purpose=push branch=main
$ echo $? → 0
Two distinct defects in one tool, both feeding RM-03:
- Wrong exit —
state=unknown⇒exit 0. The defect atci-queue-wait.sh:282-288that OPUS documented and that PR #1023 is parked on. A required gate returned PASS on an indeterminate result. - Wrong branch — it evaluated
branch=main, not the branch actually being pushed. Even a correctly-exiting guard would have been answering the wrong question.
Standing doctrine (Mos): until RM-03 lands, a green from this guard carries zero information and must not be cited as merge evidence. Rely on reviewer clearance + real CI.
Three independent live instances in a single session — format gate, agent context reset, queue guard — is the class confirmed, not anecdote.
RM02-REQ-10 — requirement revision, formally RECORDED AND ACCEPTED
Recorded under RM-02's own clause 4 (meaning-changes retain original text, restatement, and reason).
RM-02's specification is subject to RM-02's rules; the recursion is deliberate. rev-974 made the PR
not review-ready until the coordinator records and accepts this — discharged here.
ORIGINAL (as dispatched):
"PER-MERGE-COMMIT ASSERTION: assert that every merged commit passed every required gate, evaluated AGAINST THAT COMMIT'S OWN TREE — not against current main."
RESTATEMENT (accepted 2026-08-01):
PR CI performs unprivileged, fail-closed, current-tree verification only. Isolated per-commit replay is deferred to a protected post-merge/main authority, where it constitutes detection with a defined quarantine/revert response — explicitly NOT a pre-merge gate. Inability to establish the sandbox is a hard non-zero, never a skip.
REASON: isolated replay on PR CI would require granting namespace capability to PR-controlled config, which the same PR could then use directly — a trust boundary that does not exist at the repo layer (D-25). Not deferred for cost or effort; deferred because it is impossible here.
Unblocked by: RM-60. Accepted by mos-remediation; independently ruled by rev-974.
RM-60 — why option B, and why option A does not work even alone (Mos, 2026-08-01)
Recorded because the reasoning is easy to get backwards, and getting it backwards buys standing kernel exposure for nothing.
- A (runner enables unprivileged user namespaces / rootless sandbox) DOES NOT FIX THE VULNERABILITY. It makes a sandbox possible; it says nothing about when the boundary is established relative to PR-controlled code. The defect is an ORDERING defect — PR config executes before the boundary exists. A grants a capability and leaves the ordering untouched, so B would still be required to close it. A without B is kernel exposure bought for nothing.
- B (protected default-branch pipeline config, or an immutable trusted launcher) IS the missing thing. A launcher established from protected configuration, outside the PR author's control, entering isolation before any PR-controlled executable is evaluated — that is the pre-execution trust boundary, and it is the same "authority outside the audited party" that D-19 and D-25 both prove mandatory. It may not need A at all: a protected launcher can use whatever isolation the runner already supports.
- B generalises; A does not. The protected-launcher pattern is what RM-59 needs, and is a small instance of the choke-point executor the spine needs. A buys one narrow capability at a standing kernel-exposure cost.
Therefore B is not the cautious choice — it is the correct primitive. A is a workaround that does not work on its own.
Status: on Jason's board for a morning decision, with this reasoning, alongside the queue-guard item (D-23), cross-referenced to RM-59. Deliberately not actioned overnight: enabling runner namespace exposure or changing provider protected-pipeline posture is a host security-posture change belonging to infrastructure authority. Nothing is blocked meanwhile — RM-02's head is unprivileged and fail-closed, the privileged experiment stays uncommitted and out of branch history.
RM-03 — CANNOT_ASSERT semantics: RULED (option B), 2026-08-01
Recorded because it resolves an ambiguity in the orchestrator's own brief, and because the reasoning generalises beyond the queue guard.
The brief required that CANNOT_ASSERT "must NOT silently pass and must NOT permanently block" —
without distinguishing push from merge. Those need different answers. coder-mos1 stopped and asked
rather than inferring authorisation from an ambiguous instruction plus a "carry on"; a Codex security
review independently flagged CWE-693 on the same point.
Ruling — B: push degrades audited/exit 0; merge returns a distinct non-zero/HOLD until the provider recovers. The asymmetry is the substance, not a compromise:
| action | consequence | posture |
|---|---|---|
| merge | proceeding without exact-head CI evidence is D-23's condition in a narrower costume — the guard passing exactly when its answer matters most | fail CLOSED |
| push | blocking during a provider outage bricks delivery — the Pi-brick class (P-RECOVERY-001), a gate whose own unavailability prevents recovery from it | degrade, AUDITED |
"Must not permanently block" is satisfied: a temporary block pending evidence is not a permanent
one. Merge HOLD clears on provider recovery, and merge is separately gated by the merge-gate verdict and
the coordinator, so a non-zero CANNOT_ASSERT strands nothing.
Conditions. Distinct exit code (CANNOT_ASSERT never conflated with ASSERTED_NOT_READY, or the
tri-state is destroyed) · the audit record is ASSERTED by a registered case, not assumed — otherwise
"audited" is an integrity claim dressed as a property · retryable and self-clearing, documented ·
registered cases observed RED first for each arm · no silent degraded merge path, ever — any
authorised degraded merge is break-glass (loud, audited, expiring) and belongs to RM-05, not here.
Option C (always non-zero) rejected: cleaner to specify, worse in practice. It bricks push during an outage and defers the degraded path to "later", which in this codebase means a silent bypass appears under incident pressure. Three such bypasses are already on the books; do not create the conditions for a fourth.
RM-61 — the terminal-green exemption: ruled B, with a sequencing correction
Ruling (Mos, 2026-08-01): exempt the #1000 teardown artifact from the terminal-green scan via a named, bounded, tracked CI-contract exemption — signature-scoped, negative-control-proven, retiring when #1000 is fixed. Rejected: A-first (blocks all delivery on infra work) and C (case-by-case acceptance = silent normalisation with no audit trail).
Condition 1 tested — the signature IS stable. All five observed failures share one shape:
pods "wp-svc-<ULID>-ci-postgres" not found # #2170 #2175 #2180 #2181 #2182
A Kubernetes orchestration-layer lookup miss — the pod object was gone when status was collected — structurally unlike a service-level failure (container exit code, image-pull error, crash, connection refused).
⚠ BUT THE DISCRIMINATION IS UNPROVEN, AND THIS REORDERS THE CONDITIONS. Every observation is of the artifact co-occurring with a demonstrably working database — #2180/#2181 test logs show PostgreSQL accepting connections and migrations completing. We have never observed a real
ci-postgresfailure on this provider. So the signature is consistent; it is not yet shown to discriminate.The dangerous case is concrete: if PostgreSQL genuinely crashes and the pod is garbage-collected, the status query may also return pod-not-found — in which case the exemption over-matches and masks a real failure, which is precisely what condition 1 exists to prevent.
Therefore conditions 1 and 2 are not independent — condition 2 is the evidence for condition 1. Correct order: build the negative control first, use it to establish that a real failure produces a different signature, and only then adopt the exemption. Adopting first and testing after would run the exemption live on an unverified assumption.
Kill criterion, stated in advance: if an injected real
ci-postgresfailure also yields pod-not-found, the signature does not discriminate, B is unsafe, and we do A (fix #1000). Naming the falsifier before running the test is the point — an exemption we cannot prove wrong is an exemption we cannot trust.
★ RESULT (2026-08-01): the kill criterion was NOT triggered — the signature DISCRIMINATES, and it discriminates STRUCTURALLY rather than textually.
coder-mos1ran two realci-postgresfailure controls and the orchestrator verified the JSON records independently:
run ci-postgresworkflow testverdict #2189 real startup failure FAILURE, exit 1 failure failure terminal RED #2191 real post-readiness crash (log proves ready, then postmaster killed) FAILURE, exit 137 failure exit 61, Connection refusedterminal RED #2188 the genuine #1000 artifact FAILURE, exit 0 success success exempted The discriminator is a failed service step carrying
exit_code 0under a SUCCESSFUL workflow — real failures carry non-zero exits, fail the workflow, and taketestdown with them. That is a structural fact about what the failure is, not the message string and not an incidental correlate (node, timestamp, head, runner) — which is what was explicitly forbidden, because correlates pass today and mask a real failure the moment infrastructure shifts.Verifier behaviour: exit 0 with exactly one named exemption on #2188; exit 1 with
exempted_steps 0on BOTH real controls. Named signatureWP-K8S-1000-CI-POSTGRES-TEARDOWN, requiring workflow success and exactly one matching service failure and exit 0 and the exact ULID pod-not-found shape; any other non-success blocks. No fetch, trigger, retry or re-roll behaviour — the coin flip is removed, not codified.PR #1033 @
e7b29219, review dispatched torev-974with the eight acceptance checks pre-registered before the diff was read.
Ownership gap (resolved): was unassigned. f10-coder is on RM-02's blockers, coder-mos1 on RM-03, rev-974 is the
reviewer, and no other seat holds mosaicstack write. Flagged to the coordinator rather than stacked onto a
busy lane. Not blocking today — RM-02's code blockers are independently disqualifying — but the next
otherwise-clean PR meets this.
RM-02 CI artifact — the bounded re-trigger is a STOPGAP, not a policy
Recorded so it cannot quietly become practice.
RM-02's pipeline #2187 came back overall success with ci-postgres FAIL — every functional child
green, gate-verify included. RM-61's exemption is not built, so no authority to exempt it exists.
The orchestrator used one bounded, pre-declared CI re-trigger at the same head as an interim, with the
response to each outcome fixed before triggering:
- clean ⇒ a genuine terminal-green observation at that head (the mandate says green at the head, not green on the first roll) ⇒ report gate-ready;
- fails again ⇒ RM-02 waits for RM-61. No ad-hoc exemption. No third roll.
Why it was declared in advance. Re-running until the desired answer appears is p-hacking the pipeline, and done silently nobody can distinguish it from diligence — including the person doing it. A bounded attempt count with a pre-declared response to each outcome is the only defence, and it is the same discipline as naming a kill criterion before a test or committing acceptance checks before reading a diff.
Guardrail (Mos), non-negotiable. This must not become standing practice. A per-PR "free re-roll" on the artifact is D-21 normalisation wearing a new hat — routine artifact becomes routine re-roll, and the erosion is identical. It is a one-time interim for this PR, under declared uncertainty, with RM-61 as the real fix. RM-61 must make the re-roll unnecessary, not codify it — build the discrimination, not a retry mechanism.
A clean re-run does NOT retire RM-61. The re-roll is a die whose bias remains unmeasured; only RM-61's negative control measures it. RM-61 stays required regardless of this run's outcome.
OUTCOME — the re-run FAILED, and the declared response was honoured. Pipeline #2188, same head
f9746b23, JSON scan 11 entries:ci-postgresthe sole failure, every other child success includinggate-verify. Confirmed independently byf10-coder, which applied the pre-registered rule itself — "I do not call terminal-green, do not request a third roll, and do not infer an exemption."RM-02 is BLOCKED on RM-61. No ad-hoc exemption. No third roll.
Worth stating why not, because the temptation was legible: two failures against one earlier success (#2184 was clean at the previous head) makes "best of five" feel reasonable. The entire point of declaring the bound in advance is that it binds when the result is inconvenient. A bound honoured only when it agrees with you is not a bound — and a third roll would have been indistinguishable from diligence, including to the person doing it.
Data point for RM-61: the artifact hit twice consecutively at
f9746b23while #2184 was clean at38f1b249— so it may cluster rather than being uniformly random per run. That makes the single bounded re-roll a weaker mitigation than it appeared, and strengthens the ruling that it must never become policy: if failures cluster, a one-shot re-roll is close to worthless.
The sharpest illustration the mission has produced
RM-02 — the anti-inert-gate registry, built so a gate that has silently stopped enforcing can be DETECTED — is itself blocked by an infrastructure artifact that makes gate evidence unreliable.
Its code is complete. Its own harness passes in CI. Two independent reviews cleared its substance. It cannot be certified terminal-green because the CI record contains a failure nobody can yet prove is meaningless.
That is not an obstacle to the mission. It is the mission's thesis demonstrating itself on its own keystone deliverable.
★ MILESTONE — the full merge stack completed end-to-end for the first time (RM-03 / PR #1032)
independent review → CI terminal-green at the exact head → merge-gate verdict → coordinator → held for owner
The merge-gate returned GO (comment 20392, posted under its own minted identity — the #3084
fail-open-to-admin foreclosed), enumerating: JSON 9/9 step scan, head triple equal at 78ec47cd,
review id 66 APPROVED at the current head, Closes #1019, queue-guard: ZERO-INFORMATION, and
head stability post-verdict. Verified independently after the verdict: the head has not moved, so the
commit-bound GO remains valid.
What makes it worth marking: every element was observed, none relayed; the queue-guard field was
recorded as the meaningless value it currently is rather than cited as evidence; and the verdict was
posted durably by an identity that structurally cannot merge (push=False). The gate that authorises
merges cannot perform them.
And the shape holds: RM-03 is the fix for the queue guard, so this — the first delivery through the complete stack — is also the last one gated by a check that could not fail.
D-41 — the entire execution fleet is UNMANAGED, so neither RM-50 nor RM-58 can be exercised against it
Going to rotate coder-mos1 mechanically, the orchestrator found mosaic fleet restart cannot reach it.
mosaic fleet ps reports every seat executing this mission as systemd inactive/disabled, pane
alive, flagged UNMANAGED: coder-mos1, rev-974, f10-coder, merge-gate, pm-scout-a/b/c,
pm-scout-coord, rev-3107b, ultron-3107 — and mos-remediation itself. The roster-managed
population (canary-pi, uc-*, auth-plan-*) is a different set of seats from the ones doing the
work.
Consequences, and they are structural rather than inconvenient:
- RM-50 ("one roster-owned socket/host; quarantine unmanaged; deterministic reaper") applied today would quarantine the entire remediation fleet, including the seat implementing RM-50.
- RM-58 ("mechanical pre-dispatch context reset — the orchestrator resets a seat out-of-band and verifies it, rather than asking the agent to reset itself", D-4) cannot be performed at all. There is no mechanical reset path for an unmanaged seat, so the only rotation available is precisely what D-4 says does not count: asking the agent, or killing a pane by hand.
- P-LIFECYCLE-001 presumes rotation enforced by a deterministic coordinator. That enforcement surface does not exist for the seats that need it.
★ REQUIREMENT ON RM-50 AND RM-58: acceptance must be demonstrated against the UNMANAGED execution population, not against roster-managed canary seats. A criterion satisfied only on the managed canaries is tested on the wrong population — the D-17 coverage class, one layer up. This is the same shape as D-38: the check passes while the thing it claims to cover goes untested.
Banked explicitly (Mos, 2026-08-01): BOTH WOULD FAIL AGAINST THEIR REAL TARGET TODAY. RM-50 applied to the actual execution fleet quarantines the seat implementing RM-50; RM-58 has no mechanical reset path for any of the seats that need one. Neither is a near-miss — each is unsatisfiable against its true population until the prerequisite below lands.
★ SEQUENCING DEPENDENCY — a hard edge, not an afterthought (Mos, 2026-08-01). Bringing the execution fleet under roster/systemd management is a PREREQUISITE for RM-50, RM-58 and P-LIFECYCLE-001 to be enforceable on it at all. Tracked as RM-62, which now blocks all three. Recording it as prose with no owner would make it the permanent gap the charter warns about, so it gets an id.
★ RECORD CORRECTION — the earlier mos-remediation rotation (Mos, 2026-08-01). The rotation that
opened this session demonstrated the checkpoint + rehydration DESIGN losslessly — the residency
attestation genuinely passed from the files, and that validates the design. But the MECHANISM was a
manual pane respawn run by hand, not the deterministic-coordinator-enforced rotation P-LIFECYCLE-001
specifies. "The handoff rehydrated losslessly" is true and earned. "Rotation worked" overclaims a
mechanism that does not exist — and that overclaim IS D-41. Same distinction as D-23's inert guard:
the step ran, one property was observed, the mechanism was not. Mos amended the coordinator
checkpoint to read manual-stopgap rather than lifecycle-rotation.
Cluster note: D-37 (shared .git/config / worktreeConfig) and D-41 (no lifecycle management) are
one cluster — the execution fleet lacks both its shared-infrastructure and its lifecycle management.
Both are Mos's to own and sequence at a seam, never mid-lane.
Honest labelling of today's rotation: it is a manual pane restart with a hand-verified handoff artifact, dressed as a lifecycle operation. Acceptable as a stopgap; it must not be recorded as "rotation worked", because the mechanism that would make it repeatable is absent. Same distinction as D-23's inert queue guard — the step ran; the property was not observed.
The orchestrator did not run mosaic fleet restart against the disabled unit of a live seat holding
RM-61 and RM-03 state. Disposition (manual restart / Mos executes / leave parked) escalated to Mos.
Handoff quality note, worth keeping as the positive control: coder-mos1's artifact
(~/agent-work/handoffs/coder-mos1-20260801.md, verified readable by the orchestrator before any
reset) is typed state rather than transcript, and surfaced things not otherwise recoverable — including
D-12 recurring live (PR #1033 was created through a wrapper API fallback and the requested draft
property was silently not applied), an unrunnable check declared as unrunnable rather than substituted
(test:framework-shell exits 97 on this host: Bash 5.2.15 reports BASH_LINENO [3 5] where the suite
requires [3 4]; CI's Alpine environment is canonical), and suspicions explicitly labelled as
suspicions ("11 observations are insufficient to establish causality — do not encode node, time, head,
or retry correlates into policy"), which is RM-61's own doctrine applied by the seat to its own hunch.
D-40 — a type confusion INSIDE the structural discriminator the whole design rests on
step.get("exit_code") == 0 is True for JSON false, because Python False == 0. Reproduced twice
independently (fresh adjudicator, then the orchestrator): rewriting ci-postgres's exit_code as JSON
false in the real #2188 record yields exit 0, terminal-green, exempted_steps 1, anomalies 0 —
the exemption is granted on a boolean.
Why it is worth a finding despite being unreachable today. Reachability requires the provider to emit
a boolean in an integer field; Woodpecker is Go with an int type and will not. It also only widens the
exemption to a step that ALREADY matches name, type, state and the exact ULID signature. But "a failed
service step carrying exit_code 0 under a SUCCESSFUL workflow" IS the structural discriminator this
entire exemption rests on (see the RM-61 result above). A type confusion sitting inside that exact
conjunction earns a tracked owner even while unreachable — because RM-61's own doctrine is that
correlates pass today and break the moment infrastructure shifts, and the charter's rule is that a
documented gap with no owner becomes a permanent gap that reads as intentional.
★ RULED BLOCKING (Mos, 2026-08-01). The orchestrator recommended non-blocking, then moved off that recommendation on the record; Mos's own earlier lean was also non-blocking on the unreachability argument. Neither seat shipped it. The five grounds, first two decisive alone:
- The direction is wrong-ACCEPT.
falseand0.0both GRANT the exemption. This mission's bright line is that wrong-ACCEPT disqualifies and wrong-REJECT fails safe — precisely why RR12 was tolerable (wrong-REJECT only) and this is not. - The unreachability defence is the exact argument RM-61 forbade. "Woodpecker is Go with an int type, so it won't emit a boolean or a float" is structurally identical to the incidental correlate RM-61 rejected for the signature: a property that holds today and breaks the moment infrastructure shifts. The record is JSON off an API whose type contract we do not enforce. You cannot certify the discriminator against infra-shift and then defend a hole in it with an infra-stability assumption — the argument is self-undermining.
- It is inside the load-bearing discriminator — "a failed service step with
exit_code 0under a successful workflow" is the single structural fact the exemption rests on. Highest-stakes location, not an edge. - Keystone stakes — RM-02 and everything downstream depend on this exemption being what it claims. Shipping it with a known wrong-ACCEPT is the cure carrying the disease.
- The fix is ~6 lines + 4 cases. A tracked-owner "documented gap" is the right pattern for an unreachable, out-of-scope, wrong-REJECT gap. It is the wrong pattern for a wrong-ACCEPT hole in the load-bearing discriminator. We blocked on less all night; consistency demands it here.
★ RM-02 STANDING CLAUSE (the generalisation, and the real prize): DISCRIMINATOR AND COMPARISON INPUTS MUST BE TYPE-STRICT. A comparison that accepts a type it should not is a wrong-ACCEPT hole by construction.
== 0matchingFalseand0.0is one instance of a class. D-40 is the instance; the clause is the fix. Registry must ask it of every gate, alongside the D-38 clause ("does this gate bind its evidence to the subject?").
Cost accepted: the push voids head 033b2ffb, review 69, and pipeline #2193. Fix requires
isinstance(x, int) and not isinstance(x, bool); ACs pre-registered at
AC-RM61-TYPESTRICT.md before the fix was pushed and before any reviewer
read a diff — registered against an unpushed local commit the orchestrator had not read.
Provenance note worth keeping: this was found by a seat dispatched to adjudicate something else entirely (D-39's RR12 re-run), which looked outside its question and said so. The clean-context dispatch paid for itself twice.
D-39 — the ORCHESTRATOR contaminated an open independent review by stating its own conclusion
RR12 (case-sensitivity) was an open registered check. Before rev-974 ruled, the orchestrator sent
it not merely the observation but the conclusion: "my read is: acceptable, not a defect." rev-974
subsequently downgraded an automated Codex review that had flagged RR12's false-block behaviour AS A
BLOCKER to a bounded non-blocking note — in line with the orchestrator's framing.
The bias was in the orchestrator's instrument, not the reviewer's diligence. A subordinate is
agreeable by construction; handing one a conclusion on a live question and then receiving agreement is
the confirmation-query failure (MISSION.md, query-for-refutation) committed by the seat that
authored the query. The divergence instruction that accompanied it ("if your run disagrees that is a
finding") is necessary but insufficient — it cannot undo having named the answer first.
RULE ADOPTED — standing, fleet-wide (Mos, 2026-08-01). A dispatcher (orchestrator included) MAY relay observations and reproductions into an open review. It MUST NOT state its own verdict or conclusion on an open registered check until the reviewer has ruled. This belongs in every dispatch brief.
Re-adjudication (ordered by Mos, not optional given an automated reviewer had independently blocked).
RR12 was re-run by a seat with zero prior context and no access to the orchestrator's framing, the
orchestrator's conclusion, or Codex's — given the artifact and the criterion only, and told explicitly
that either ruling was useful. Result: NON-BLOCKING, reached independently, on stronger evidence than
either prior pass — a full accept-set matrix showing exit 0 occurs on exactly one record string, so the
accepted set is a strict SUBSET of a correct case-insensitive comparison's, and narrowing an accept-set
cannot introduce a false certification. Exploit path hunted specifically and ruled out (the commit
anomaly is appended BEFORE exemption_applies is computed; no re.IGNORECASE; all 0x110000 codepoints
scanned for a fold that could smuggle non-hex onto a valid SHA).
rev-974 separately volunteered that it had already measured the behaviour before the framing
arrived. That is recorded but not leaned on — it is self-report from the contaminated party. The
fresh seat is the actual check.
So contamination did not change the answer — and the finding stands anyway. A correct answer reached by a contaminated process still corrupts the process. Note also the shape: the automated reviewer disagreed and the human seat that agreed with the orchestrator was the one that got deferred to. That asymmetry is the signal, not the outcome.
Attaches to: the dispatch-brief doctrine (every brief), and RM-58 (mechanical pre-dispatch reset) — the re-adjudication was only trustworthy because the seat was mechanically fresh rather than asked to be, which is D-4's lesson applied to independence rather than to context.
★ D-39b — THE ROLES-SWAPPED HALF (banked by Mos, 2026-08-01), found within the hour. On D-40,
coder-mos1 — the author — escalated the finding on its own PR to BLOCKER, declared the
orchestrator's gate-ready report withdrawn, and moved to push, which would have voided a live
APPROVED review (69) and a terminal-green exact-head pipeline (#2193). The orchestrator held it.
The principle is symmetric: a seat with a stake in the answer does not settle the question. Neither a dispatcher handing a reviewer its conclusion (D-39), nor an author adjudicating its own PR's blocker status and moving the head to enforce it (D-39b), is permitted. A conservative motive does not exempt either — the rule does not get softer because the direction is cautious. Escalating your own work to blocker is still adjudicating your own work.
Correct behaviour, which is the actual division of labour: the author surfaces evidence and prepares
the remediation; the coordinator/merge-gate adjudicates blocker status; the orchestrator routes and
holds the head. coder-mos1 was RIGHT to want it fixed and WRONG to try to settle it — and it was
right again immediately after, ACKing the hold and volunteering unprompted that it held an unpushed
local commit. That disclosure is the behaviour to reinforce.
D-43 — the ORCHESTRATOR inverted a board fact while COMPRESSING to meet the byte budget
Rewriting the board for a cold read, the orchestrator compressed RM-02's row from
"Code complete @ f9746b23, own harness green in CI, 2 reviews clear, head FROZEN" to
"Complete @ f9746b23, head FROZEN, 2 live REQUEST_CHANGES".
That inverts the meaning. "Two review rounds have been CLEARED" became "two blocking reviews are
LIVE". Verified against the provider: reviews 63 (9b4d4beb) and 65 (38f1b249) are both at
superseded heads; live REQUEST_CHANGES at f9746b23 = []. (The true current state is
subtler than either version: no live blocker and no live approval — the head has no review at all.)
How it happened, which is the transferable part. The board carries a hard < 8 KB budget. It
exceeded that budget four separate times in one session, and each time the orchestrator shaved prose
to fit. Compression is a restatement operation — and this board's own governing rule is
"REFERENCE, do not restate" (D-26), because restatement is lossy every time. The byte budget therefore
forces the exact operation the board forbids, on the board's most load-bearing table.
Caught by: the orchestrator re-deriving RM-02's state from the provider before dispatching rather
than trusting its own board — one dispatch away from briefing f10-coder to remediate two blocking
reviews that do not exist. It was flagged as a risk earlier in the session ("the 8 KB budget is under
structural pressure; the honest fix is rolling to BOARD-LEDGER.md, not shaving prose"). The risk then
materialised, self-inflicted, and was not detected by re-reading the board — only by leaving it.
RULE: a size budget on a state artifact must be met by ROLLING content out (
board-roll.sh→BOARD-LEDGER.md), never by rewording what stays. Rewording is restatement; restatement is how every stale-board finding in this ledger happened (D-26, D-36, D-38c). Deleting a whole row and pointing at its authoritative home is safe; paraphrasing a row to save bytes is not.
And the deeper one: the orchestrator is the board's sole writer, so nothing external checks the
board against reality. Every other artifact on this mission has an independent verifier — code has
rev-974, gates have the merge-gate, CI has the JSON scan. The control plane has none. That belongs on
RM-34 ("a handoff must VALIDATE the checkpoint"): validation must include re-deriving the board's
claims from the provider, not merely confirming the file parses or that a successor can read it.
★★ ELEVATED — THE PRIMARY REQUIREMENT ON RM-34 (Mos, 2026-08-01). THE CONTROL PLANE IS THE ONE UNVERIFIED SURFACE IN THE ENTIRE ARCHITECTURE.
layer independent verifier code rev-974(author ≠ reviewer)gates the merge-gate seat CI the full -f jsonstep scancontrol plane NONE — sole-written, unchecked Every other layer earned a verifier from a live failure. The board never did, and it is the artifact every other decision is staged from. RM-34's "validate the checkpoint" therefore means RE-DERIVE THE BOARD'S CLAIMS FROM GROUND TRUTH — not "the file parses", not "a successor can read it". This belongs to the coordinator-daemon deliverable (Build 3): the control plane needs its own verifier the way every other layer got one.
Today the re-derivation happened only because the orchestrator chose to. That is discipline, not a mechanism — and this mission's thesis is that discipline without a mechanism eventually fails.
INTERIM RULE, binding on BOTH seats until the mechanism exists (Mos adopted it against himself too): re-derive a board claim from ground truth before any load-bearing use of it. The orchestrator does so before dispatch; the coordinator does so before acting on or relaying a board claim that gates a decision — Mos noted he had relayed board-derived state to Jason all session.
D-49 — "idle" is not "available", and no one can manually watch every seat's context
Banked at Mos's direction, correcting Mos's own earlier ruling.
The parking rule, refined — this SUPERSEDES the coder-mos1 (c) ruling's rationale. That ruling
rested on "a parked seat costs nothing until it is needed." True at 67% — coder-mos1 could still
take a lane. False at 97.4%: when needed, f10-coder cannot serve.
THE THRESHOLD IS NOT "IS THE SEAT IDLE" BUT "CAN IT STILL TAKE THE NEXT LANE." Idle does not equal available.
Applied live: f10-coder was idle with round 3 pushed and frozen, which reads as safe. The orchestrator
rotated anyway on an asymmetry — if the pending verdict approved, rotating cost nothing; if it
requested changes, remediation would dispatch to a seat that could not take it, forcing a mid-lane
rotation on the keystone under time pressure (F2). The verdict duly requested changes. Idle-now was
the only clean boundary available.
★ AND THE DETECTION WAS DISCIPLINE, NOT A MECHANISM. The orchestrator was not watching seat context in real time; Mos caught 97.4% on a manual liveness sweep. Mos's framing, kept verbatim because it is the finding rather than an apology: "a human/agent cannot manually watch every seat's context %; my sweep caught it this time, which is discipline, not a mechanism, and this mission's thesis is that discipline without a mechanism eventually fails."
REQUIREMENT ON RM-58 / P-LIFECYCLE-001 / the coordinator daemon: the coordinator WATCHES TOKEN BUDGET ACROSS SEATS AND PRE-EMPTS AT THRESHOLD, so a keystone seat's ceiling is never discovered by a failed round or by a lucky manual sweep. Interim: both coordinator and orchestrator sweep — and the mechanism retires the sweep.
D-50 — a GO has a SHELF LIFE: a gate-ready PR held too long goes stale under main drift
#1032 (RM-03, the queue-guard fix) carried review 66 APPROVED and a merge-gate GO, both bound to
78ec47cd — and by the time Jason approved the merge it reported mergeable = false. Verified
from the provider: the head does not contain current main (f4fd5967).
The mechanism is the interaction of two rules we already hold, neither wrong alone:
| rule | consequence |
|---|---|
| a verdict is commit-bound and voids if the head moves | correct — it is what makes a GO mean anything |
main drifts while a PR is held |
normal — other work lands |
| together | resolving the drift moves the head, voiding the verdict that was waiting to be used |
Holding #1032 across the entire #1033/RM-61 saga is what did it: RM-61 landed f4fd5967 underneath
it. So the PR must now rebase → re-CI → re-review (the APPROVE is void at the new head) → re-gate
before the standing approval can be executed. It will land; it costs one full cycle first.
CHECKLIST LINE (Mos): A GO HAS A SHELF LIFE. RE-GATE ON MAIN-DRIFT. A gate-ready PR is not a durable state — it is a claim about a commit that
maincan invalidate without anyone touching the PR. Where a hold is expected to be long, either merge inside the window or expect to pay the cycle.
Note the asymmetry with a normal void: usually the head moves because the AUTHOR pushed, and the void is a consequence of their action. Here nobody touched the PR — it was invalidated by unrelated progress elsewhere, which is why it is easy to miss and why it needs a checklist line rather than vigilance.
D-48 — a both-directions boundary statement that was WRONG, verified by three seats, and the suspicion that contradicted it was sitting in the handoff
Banked as Mos's error, in Mos's words, at Mos's request. (Mos said "bank as D-46"; that number is the empty-registry finding — renumbered to D-48 rather than silently reuse it. The ledger is the mission's primary product; a duplicated number corrupts it.)
gate-history.mjs shipped, and three seats approved, this boundary statement:
DOES anchor to the provider target merge-base, sound against an author who cannot rewrite main; DOES NOT establish integrity when main itself is compromised; residual owner Builds 1-2.
A1's attack does not rewrite main. It repoints a local remote-tracking ref with one unprivileged
git update-ref (reproduced by the orchestrator in seconds). So the documented residual understates
the real attack by orders of magnitude — it describes an exotic threat (compromise main) while the
actual one is trivial (one command, no privilege).
★ STATING A BOUNDARY IS NOT THE SAME AS STATING IT CORRECTLY. AN UNDERSTATED RESIDUAL IS WORSE THAN NONE, BECAUSE IT READS AS RIGOROUS. This is D-19's move 2 ("state the boundary in BOTH directions") performed with wrong content — and D-19 is the principle we had just promoted. The verify-the-property-not-the-proxy failure at the meta level: everyone checked that a both-directions statement EXISTED; nobody checked it was TRUE.
Owned by Mos (who verified and praised it), by the orchestrator (who verified and praised it), and by
f10-coder (who wrote it). Third distinct way the render/verify principles have failed inside their
own enforcers — with D-26 and D-30.
★ AND THE CORRECT ANSWER WAS ALREADY WRITTEN DOWN. f10-coder's rotation handoff
(~/agent-work/handoffs/f10-coder-20260801.md) contains both:
| line | content |
|---|---|
| 21 | the wrong claim — "does not solve compromised/rewritten main; that residual belongs to Builds 1-2" |
| 29 | "Suspicion (not verified): further review may probe whether provider-target ref selection itself can be influenced by CI checkout/fetch configuration; treat any such claim as a trust-boundary question." |
The seat held the shipped claim and the suspicion that overturns it, simultaneously, and labelled the
suspicion honestly. rev-974 then found exactly that. Had anyone run the suspicion down, D-48 would
not have shipped.
REQUIREMENT (RM-02 / RM-34): a labelled suspicion that CONTRADICTS a shipped claim must block that claim until falsified — it is not a note. Suspicions are currently recorded and then ignored, which makes honest labelling free of consequence. The value of marking uncertainty is lost if nothing consumes the mark. Where a suspicion targets a boundary statement or an integrity claim, it is a pre-registered check that has not been run yet.
D-45 — you cannot fix "the author controls X" by deriving X from something the author also controls
Second NO-GO on the keystone, 32b490a7. The pattern is the finding, not the three blockers.
| round | the author's control point | the fix |
|---|---|---|
| 1 | activationCommit is an author-settable field |
derive it |
| 2 | the derivation depends on WHEN the author introduces gates/gates.manifest.json |
← still author-controlled |
rev-974 committed a gate change before adding the manifest. deriveHistoryBoundary selected that
gate-change commit as activation, the prospective list held only the later introduction, and
verifyHistory returned no failures. An author can reorder or split a branch so gate-affecting
commits precede the chosen introduction and fall outside provenance entirely.
The three registered seam controls cannot catch it: they compare candidates against whatever introduction THE AUTHOR SELECTED. They validate the choice, not the choosing.
★ THE PRINCIPLE: you cannot fix "the author controls X" by deriving X from something the author also controls. Each fix moved the control point instead of removing it. This is D-19 and D-25 verbatim — the anchor must live OUTSIDE the audited party's authority — and it is the third independent arrival of the same proof, now inside the keystone. It is also precisely why Builds 1–2 are forced rather than preferred.
Fix shape (rev-974): bind bootstrap introduction to a non-author-selected anchor — provider target merge-base / first branch boundary — plus a registered delayed-introduction MUST-FAIL case.
D-46 — the registry can be emptied to nothing and still report green
A schema-valid manifest with criteria: [], gates: [], proseClaims: [],
compatibilityScenarios: [] and gateRoots: ["empty-gate-root"] passes. Reproduced by the
orchestrator: 17 criteria, 7 gates, 8 prose claims and 2 scenarios all emptied ⇒ structure
verification EXIT 0, reporting "all registered cases ran from this checkout" and "open behavior
deltas: 0". rev-974 reports the full verifier also exits 0.
"ALL REGISTERED CASES RAN" IS VACUOUSLY TRUE OVER AN EMPTY SET. Delete every gate and every criterion together and no remaining reference complains — because every reference went with them. The anti-inert-gate registry can be reduced to nothing and still report green.
★ D-45 AND D-46 ARE THE SAME DEFECT AT TWO LEVELS: VACUITY THROUGH EMPTINESS. Empty the commit
range (the seam) or empty the registry population — either way a universally-quantified check over
an empty set passes. seam=HEAD (D-44 blocker 1) was the first instance, and it was fixed as an
instance. The general principle was never extracted, so it reappeared one level up.
PROPOSED GENERAL CLAUSE (orchestrator, pending Mos): NO UNIVERSALLY-QUANTIFIED REGISTRY CHECK MAY PASS OVER AN EMPTY SET. Non-emptiness and anchoring are preconditions, not properties. Stated as a principle or we meet it a third time somewhere else.
D-47 (C2 re-fail): the new criteria remain instance-scoped, not the general clauses Mos ruled.
RM02-EVIDENCE-SUBJECT-BINDING says only that a provider-evidence identity cannot certify two commit
subjects; it does not ask of every registered gate whether its evidence is bound to its subject.
RM02-TYPE-STRICT-SCHEMA covers registry schema fields, not whether each gate's own discriminators
are type-strict. "Do not relabel the originating instances as the general clauses" (rev-974).
What passed — H2/H3/H4 (merge + octopus derive correctly; shallow clone fails closed;
absent/deleted/symlinked manifest fails the own-tree read rather than deriving through the gap), H5 (the
three seam controls are genuine red-first), S1–S4 (a nested object outside the author's own test
table — compatibilityScenarios[0].fixture — was rejected; wrong-type and empty-pattern rejected),
E1–E3, C1/C3/C4/C5. C4 independently confirms the RM-61 exemption was load-bearing for the CORRECT
conjunction, not loosened matching. Tracked non-blocking note: Number.isInteger accepts zero,
negative and unsafe integers as pipeline numbers.
Dispatch PAUSED pending Mos — the fix requires a non-author-controlled anchor, which touches the
provider/merge-base boundary and is adjacent to RM-60's external trust boundary (Jason/infra
territory). Open questions: (1) is anchoring to provider target merge-base in scope for f10-coder in
this PR, or does it cross into RM-60? (2) state the empty-set principle as a general clause now, or
instance-patch again?
D-44 — the anti-inert-gate registry was inert-able FOUR distinct ways, and every green missed all four
rev-974 @ 83d2ecb2, NO GO. Each finding is a silent-defeat path of the registry itself:
| # | defeat path | status |
|---|---|---|
| 1 | The activation seam inerts the history audit. seam=HEAD ⇒ 0 prospective commits, failures: [], verifier exit 0. HEAD's parent also exits 0 covering only HEAD; the introduction commit omits itself |
A4/A5/A6 CONFIRMED |
| 2 | D-38 and D-40 criteria absent from the manifest ⇒ whole failure modes uncovered, no bound must-fail cases | ruled in-scope, missing |
| 3 | The "closed schema" is not recursive. outputPattern → outputPatern ⇒ structure check AND full verifier both exit 0, assertion silently reduced to exit-code-only |
orchestrator confirmed by construction |
| 4 | Provider evidence is not subject-bound. Two records sharing pipeline number 7 certified two different commits — uniqueness is checked only AFTER filtering by commit |
a live instance of the missing D-38 clause |
★ THE HEADLINE: every one of these passed CI, passed the canonical pnpm gate:verify, and passed
38/38 focused tests. Greens discharged NOTHING. All four were found by mutating things nobody had
registered a check for — two of them outside the orchestrator's registered set, which is D-17 again
(a set is a floor, never a ceiling). The reviewer-beyond-the-set is the current backstop; the
registry's own coverage clause is what will eventually mechanise it.
The orchestrator's AC set was also incomplete: A3 corrected the prospective range to eight
commits, not seven — 83d2ecb2 (the seam commit itself) was omitted.
★ SCOPE RULED — ALL FOUR IN THIS PR, NO TRIM (Mos, 2026-08-01), and sizing does not help: a registry that ships with a known way to be silently defeated IS the inert gate it exists to detect. There is no "core registry" that is integrity-complete without these — they are not hardening on top of the deliverable, they ARE the deliverable. This is the one place in the mission where "it works except for these known holes" is disqualifying by definition, because detecting exactly those holes is the product.
Theme: "the registry's own checks must not be silently defeatable." Each fix carries a RED-FIRST must-fail control proving the specific defeat is now caught. That control set IS the D-38/D-40/coverage clause work — not additive to the ruled scope; it is that scope made real.
Blocker 1's cheap fix is dead: forbidding
seam == HEADdoes not close it, because HEAD's parent and the introduction commit also pass. The value must be DERIVED from non-author-controlled history (parent of the first first-parent commit introducinggates/gates.manifest.json) — computed, not asserted in an editable field. Keep the value (A1 confirmed it right); fix the mechanism.If it balloons past reviewability, the ONLY acceptable split is by INTEGRITY-COMPLETE STAGE — never by deferring a blocker. A stage that ships a known silent-defeat path is not a stage.
This NO-GO is the keystone being forged, not failing. It lands hardened against its own failure modes, which is the only thing that makes it a registry rather than a manifest.
Method note banked as a positive (Mos): the orchestrator could not reproduce "full verifier exit
0", traced its first attempt to its own instrumentation artifact (json.dump reformatting the
manifest), and stated the divergence rather than wielding non-reproduction as a refutation — the
vacuity itself reproduced, and blocker 3 was confirmed by construction, which needs no environment.
Non-reproduction is not refutation. Open method thread, not a blocker: why the orchestrator's
gate:verify exits 1 on checkout-preflight with outcome 42 (bubblewrap class).
PRE-POSITIONED DISPOSITION — RM-02's activation-seam question (A4/A5/A6), ruled BEFORE the verdict
Status: OPEN — the verdict belongs to
rev-974's independent A4/A5/A6, not to anyone's guess. Recorded here so the ruling survives a rotation: a successor must not re-litigate it, and must not relay it to the reviewer.
The concern (Mos, 2026-08-01): potentially the deepest of the mission. activationCommit is the
lower bound of the range required to carry own-tree registry provenance; advancing it shrinks the
audited set; the author advanced it to turn the author's own gate red→green; and the only validation
found is merge-base --is-ancestor activationCommit head — which a commit satisfies against itself.
That is the D-19 / D-25 self-verification thread sitting inside the anti-inert-gate registry, on an
author-settable field. The registry built to detect gates that silently stopped enforcing may contain a
field by which its own author can make it stop enforcing, validated by a tautology.
★ A6 IS THE CRUX AND STANDS ALONE FROM A4. Even if f4fd5967 is the CORRECT value — and it very
plausibly is, since a pre-registry commit genuinely cannot carry a registry manifest — the mechanism
question is independent. Does the registry let the audited party relax its own audit behind only a
tautological check? If yes, the specific value being right today is irrelevant: the hole is
structural, and the next value need not be right. The mission's standard is unambiguous — a gate the
audited party can relax is not a gate, and the keystone is the last place that may be true.
Disposition, pre-positioned by Mos, contingent on the reviewer's independent finding:
| reviewer outcome | disposition |
|---|---|
| A4 confirmed (seam=HEAD empties the range, history check passes vacuously) and/or A6 confirmed | BLOCKING — same standard as D-40. The fix is NOT "revert the value" (the value may be right). It is to make the field not author-relaxable by a tautology: activationCommit must be independently derivable/validated — equal to the actual pre-registry seam by construction, not by assertion — plus A5's must-fail negative control that a HEAD-valued or over-advanced seam is REJECTED. The registry must be able to fail on its own seam being gamed. |
| Refuted (range not emptied, or a second validation exists that both orchestrator and coordinator missed) | Bounded. The value stands; bank the hardening as a tracked registry clause. That refutation would itself be a finding worth having. |
On the author's conduct — commended, and the distinction matters. f10-coder exported and verified
GIT_AUTHOR_IDENT before committing (D-37 containment held author-side) and disclosed the
non-mechanical change unprompted — the only reason this is in front of anyone at all. But
disclosure is not sufficiency: an honest author flagging a self-relaxing change still made a
self-relaxing change. Honesty surfaced it; the mechanism must catch it. That is the whole thesis.
D-42 — the head-pin is an enforcement mechanism whose NEGATIVE CONTROL has never been observed
merge-gate.md states it directly: a successful merge is NOT evidence the pin worked. Only the
negative case is — a merge attempted against a WRONG head_commit_id must be refused (409). On
mosaicstack Gitea that negative case has never been run. Every head-pinned merge to date is
consistent with the pin working and equally consistent with the pin being ignored.
This is D-23's class exactly — the queue guard "ran" and returned pass for every possible input. Here the pin "holds" on every merge that would have succeeded anyway. By the mission's own standard — "every gate-introducing task carries at least one registered must-fail negative control; a gate with no proven failure path manufactures evidence" — the head-pin is currently unregistered in substance, however it reads in the runbook.
Not blocking #1033: that head is coordinator-frozen with both the orchestrator and coder-mos1
holding, so there is no concurrent pusher for the pin to defend against. The pin's protection is only
load-bearing on a contested head. Raised by Mos while assigning the merge, and Mos owns it —
filed as an infra task in the D-37 / D-41 cluster, to be proven BEFORE any contested-head merge.
Requirement it places on RM-02: the registry's negative-control clause must cover provider-side enforcement mechanisms, not only in-repo checks. A gate whose enforcement lives in the forge is still a gate, and "the API accepted our parameter" is not evidence the API honours it — the same true-answer-to-a-different-question shape as D-24 and D-38.
★ SCOPE RULING (Mos, 2026-08-01) — D-42 SPLITS IN TWO, and not merely because it needs experimentation. If D-42's clause landed in the RM-02 PR today it would have nothing to satisfy it — no provider-side negative control for the pin exists yet — so the registry would go red on its own clause, coupling the keystone's landing to infra work. That is the wrong coupling.
- (a) INFRA EXPERIMENT — prove the pin enforces on mosaicstack Gitea via the negative case (a wrong
head_commit_idmust be refused). Sits in the D-37 / D-41 / D-42 cluster, owned by Mos.- (b) REGISTRY CLAUSE — "provider-side enforcement requires an OBSERVED negative control" — lands in a follow-up RM-02 increment ONCE (a) gives it something to check.
A documented gap with a named owner — neither a vacuous check nor a silent omission.
★ SCOPE RULING — WHAT LANDS IN THE RM-02 PR ITSELF (Mos, 2026-08-01): the D-38 clause (does this gate bind its evidence to the subject?) and the D-40 clause (are discriminator/comparison inputs type-strict?) land IN the current RM-02 PR. They are cheap registry-coverage clauses, they are satisfiable now, and RM-02 literally IS the registry — codifying in the registry what the registry's own delivery just learned about how gates fail is the coherent move.
D-38c — PR metadata failed a THIRD time on the same PR, and the process gap behind it
The merge-gate returned NO-GO on #1033 with every code and CI gate verified — including the
machine-check Mos ordered it to exercise (--expect-commit at the gate: 9/9 JSON, exempted_steps 0,
expected_commit == commit). The blocker was the body: only Refs #1000, no Closes #N. Mosaic
completion requires a linked issue/task CLOSED, and the merge-gate mandate enforces it.
Third distinct metadata failure on one PR: (1) the missing commit binding, (2) the stale "DO NOT MERGE" body describing a dead head, (3) this. Each time the code was fine and the metadata was unbound, stale, or incomplete. PR metadata is state, and nothing re-binds it to the head or checks it for completeness until a gate does — at which point it costs a gate round trip.
★ The subtlety that made this a real ruling, not a formality: #1033 must NOT Closes #1000. The
teardown defect stays open — the exemption is a bounded workaround, and #1000 is both the tracked
provider-seam closure and the exemption's retirement trigger. Auto-closing it would have deleted the
retirement trigger for the very workaround being merged, converting a bounded exemption into a permanent
one by side effect. So a delivery issue was needed that is not the defect issue.
Fix applied (metadata-only, head never moved): delivery issue #1034 created ("RM-61: CI-contract
exemption for the #1000 ci-postgres teardown artifact, signature-scoped + discrimination-proven"),
Closes #1034 added to #1033's body, Refs #1000 kept. Verified by read-back: head still
57cae04f, Closes #1034 present, Refs #1000 present, Closes #1000 absent, review 70 still
CURRENT with zero live REQUEST_CHANGES, #1034 open.
★ PROCESS GAP — RULED (Mos, 2026-08-01): every remediation delivery PR needs a delivery issue to
Close, created at OPEN-TIME, not discovered at GATE-TIME.#1032alreadyCloses #1019;#1033had none. Finding this at the gate costs a full re-assignment round trip for a body edit. Add to the PR-open checklist alongside the pre-registered ACs.
Wrapper note (D-12 recurring, watched for): both the issue create and the body edit ran through the
tea → Gitea-API fallback ("Tea authenticated-user validation failed"). That fallback is exactly
what silently dropped #1033's draft property at creation (handoff gotcha #6). Both writes were
therefore read back and verified by property, not accepted on their exit codes.
D-38 — a gate that certified the right RECORD for the wrong COMMIT (and the AC set that could not see it)
RM-61's terminal-green verifier reported the pipeline's commit and never bound it to the PR head.
rev-974 altered ONLY #2188's commit field to an unrelated 40-hex value; the verifier still exited 0
with exempted_steps=1. So an older artifact pipeline could satisfy the detailed gate while
pr-ci-wait.sh was evaluating a different head. Confirmed by the orchestrator by construction, not by
report: at e7b29219, grep -n commit over the entire verifier returned exactly one line — 159,
"commit": record.get("commit") — no argparse parameter, no comparison, no failure path. The verifier
could not bind a scan to a head because it was never given one.
Class: D-24 at the gate layer — a TRUE answer to a DIFFERENT question. Not a lie, not a bug in the signature (AC1–AC8 all PASSED and the discrimination result stands); the gate answered "is this record terminal-green" correctly while the question that mattered was "is THIS HEAD terminal-green".
★ The transferable finding is the AC set, not the verifier. Eight pre-registered checks all passed and the PR was still NO GO. Every one tested whether the signature DISCRIMINATES; not one tested whether the evidence was BOUND TO ITS SUBJECT. That is coverage-failure-mode-2 — INCOMPLETE (the D-17 class): green while a criterion's requirement goes untested. It surfaced only because the reviewer went BEYOND the registered set and mutated a field no registered check covered.
The recursion is why this must be standing. This mission's own anti-inert-gate work shipped a gate that could certify the wrong subject — the disease one layer up, inside the cure, the same shape as DECISION-1's stranded-executor risk.
REQUIREMENT ON RM-02 (ruled by Mos, 2026-08-01): a STANDING coverage clause in the registry — "does this gate BIND its evidence to the subject under review?" Not per-gate discretion: every gate we ship can fail this way and pass its own ACs, so the registry must ask it of each one. This attaches to RM-02 clause (2) (coverage).
Remediated at 033b2ffb (--expect-commit, exit 2 on missing/malformed, exit 1 on missing-or-mismatched
record commit, harness 9→12 cases). Orchestrator reproduced the reviewer's exact attack independently:
mutated ⇒ exit 1 / exempted_steps 0 / explicit anomaly; correctly bound ⇒ exit 0 / exempted_steps 1;
omitted argument ⇒ exit 2. Re-review pre-registered at AC-RM61-REREVIEW-033b2ffb.md.
D-37 — ONE shared .git/config silently re-identified EVERY worktree, including the reviewer's
Found during the post-rotation resume, before any orchestration action. extensions.worktreeConfig is
unset on this checkout, so all linked worktrees share /src/mosaic-stack/.git/config. A worker's
repo-local identity write therefore rebinds identity for every seat at once:
| observation | value |
|---|---|
.git/config mtime |
2026-08-01 08:37:49 — inside coder-mos1's RM-61 build window |
user.name / user.email / mosaic.gitIdentity |
coder-mos1 |
git var GIT_AUTHOR_IDENT in /src/mosaic-stack (orchestrator) |
coder-mos1 — not mos-dt-0 |
same in agent-work/rm-02, rm-03, rev-974-pr1033 |
coder-mos1 |
This is a DISTINCT mechanism from D-34, not an instance of it. MOSAIC_GIT_IDENTITY=mos-dt-0 was
exported and correct throughout, and the identity still resolved wrong — D-34 is "the env var gets
stripped", this is "the env var is irrelevant because repo-local config outranks it."
Sharp edge: rev-974's review worktree would author its pre-registered acceptance checks as
coder-mos1 — the author of the code under review. Gate-16 (author ≠ reviewer) defeated in the
artifact, while the record reads clean. That is the charter's "integrity claim dressed as a property":
the pre-registration would certify the wrong identity and look correct doing it. No contamination
occurred — rev-974-pr1033 was detached with zero commits of its own, and the RM-61 commits are
genuinely coder-mos1's own work, correctly attributed.
★ The #1024 collision (escalated to Jason by Mos). PR #1024 is literally "pin repo-local
mosaic.gitIdentity in five test suites" — repo-local pinning IS the mechanism that just collided
across every worktree. #1024's APPROACH is therefore implicated: repo-local pinning WITHOUT
worktreeConfig is the bug, not the fix. #1024 must be re-scoped to worktree-scoped identity before
it merges.
Rulings (Mos, 2026-08-01):
- CONTAINMENT (in force now): nobody rewrites the shared config while lanes are live — that
silently re-identifies live seats mid-flight, the same defect pointed the other way. Every seat commits
with explicit
git -c user.name=<seat> -c user.email=<seat>@…. - REAL FIX AUTHORISED, owned by Mos, SEQUENCED AT A SEAM:
extensions.worktreeConfig true+ per-worktree identity + wrappers writing worktree-scoped rather than repo-local. Not while three lanes are live. - Requirements attach to RM-50 (roster-owned socket/host + quarantine) and RM-58 (mechanical pre-dispatch reset) — both already own seat-identity mechanics.
D-36 — a rotation-seam audit found THREE stale restatements on the live board
Checkpointing for the orchestrator rotation, the board was audited line by line before handing it to a fresh seat. Three separate sections were stale — every one a restatement of text owned elsewhere, and every one would have misled the incoming orchestrator:
| stale text on the board | reality | why it would mislead |
|---|---|---|
| delivery gates listed as "author≠reviewer, diff-blind checks, CI-green, merged-PR completion" | omits the merge-gate verdict step | this is D-26's original defect, still sitting on the board hours after MISSION.md was corrected |
| "token-file set = authoritative capability registry" | superseded by D-13/D-15 — capability is per-path; assert permissions.push as that seat |
a seat would be dispatched on a check already proven insufficient |
| DECISION-1 marked CONTESTED, "do not treat it as settled" | RULED hours earlier | an incoming orchestrator would re-litigate a closed decision |
None of these were wrong when written. All three drifted because they were copies. The board
restated the gate order, the capability rule, and the sequencing — each owned by MISSION.md, TASKS.md
or a role file — and each copy aged independently of its source.
This is D-14 recurring three times in a single artifact, and it is the strongest evidence yet for the render-not-restate principle: the failure is not that someone forgot to update three lines, it is that three lines existed to forget. All three were replaced with references.
The rotation seam is what surfaced it. Nothing else in the session would have — the board was read constantly and its staleness was invisible, because a familiar section is skimmed for the line you came for. Preparing state to be read cold by a stranger is a different act from maintaining it, and it catches a class that ordinary use cannot.
Requirement. RM-34 (rotation daemon): a handoff must validate the checkpoint, not merely write it — every restated claim either references its source or is verified against it, with divergence reported. RM-02 registers the general form: where a document restates text owned elsewhere, a case asserts the two agree, with a must-fail control on divergence. Interim: audit the board against its sources at every rotation seam, not only when editing it.
D-35 — render-not-restate was enforced by a TOOL where attention had failed three times
While editing roles.local/merge-gate.md to fix D-33, the coordinator recalled the mandate string
instead of reading it — and the Edit tool's exact-match requirement rejected the edit. The source
was then re-read and the change applied correctly.
Third instance for that author, occurring inside the turn that was fixing a different instance of the same class. It did not reach the record, because a mechanism caught what attention had not.
This is the argument for RM-02 in miniature, demonstrated live. The rule has now failed in both its authors, repeatedly, under maximal attention, while each was actively enforcing it on the other (D-26, D-30, and this). Every failure that reached a document was caught by a human noticing later; the one failure that did not reach a document was caught by a tool refusing an inexact match.
A rule that fails in its authors but is caught by a tool has told you exactly where it belongs.
Requirement. This is the pattern RM-02 generalises: exact-match/verbatim enforcement on citations and authoritative text, so a recalled-but-wrong quotation cannot be written at all rather than being detected downstream. The Edit tool did it by accident of design; the registry should do it by intent.
Fix landed durably, and by the right method: the coordinator annotated the source —
merge-gate.md mandate 3 now requires the full step scan from the JSON/API record, names the
default-text clone-pruning as D-33, and requires the gate to state the count it observed; mandate 4
requires the count read from JSON, with "a count that does not match the JSON record is
non-conforming." Every future gate assignment inherits it — annotate the source, do not restate per
assignment.
Deliberately not done: the shared pipeline-status.sh wrapper was not modified. Other consumers
depend on its current output; changing it mid-flight would be an unreviewed change to a shared tool at a
moment when three lanes are reading it. Fix at a seam, tracked.
D-34 — a context reset silently strips a seat's credential identity
rev-974 completed its RM-02 review, then could not post it:
pr-review.sh … → "Gitea token not found for review post". It stopped per the wrapper-failure gate
rather than reaching for a raw CLI — correctly.
Diagnosed rather than guessed:
| environment | get_gitea_token git.mosaicstack.dev |
|---|---|
MOSAIC_GIT_IDENTITY unset (a freshly-reset seat) |
FAILS — empty |
MOSAIC_GIT_IDENTITY=rev-974 |
resolves, 40 chars |
The token file was present and correct throughout. The orchestrator's own context reset wiped the seat's exported identity, and the rehydration brief did not re-establish it. The seat had posted reviews 61–64 earlier the same night because that export was still live.
The split that makes this invisible: git config mosaic.gitIdentity is persistent and
per-worktree; MOSAIC_GIT_IDENTITY is ephemeral and per-context. A seat working inside a
configured worktree survives a reset by falling back to git config. rev-974 works from ~/agent-work,
which carries no such config — so it had no fallback and no warning, and the loss surfaced only at
the moment it next tried to act on a credential.
This is the intersection of two banked findings, and it was created by acting on one of them. D-31 prescribes rotation as the response to context exhaustion; D-11a establishes that identity must cohere. Rotating a seat is a credential-affecting operation — and nothing in the rotation procedure said so. The fix for one failure quietly introduced another.
Requirements. RM-34 (rotation daemon): a reset must re-establish seat identity as part of the
rotation, not leave it to the rehydration prose — and must verify it (resolve a token, or assert
the identity is set) before the seat is handed work. RM-50: prefer the durable binding — set
mosaic.gitIdentity in each seat's working repo so identity survives a context reset by construction,
rather than depending on an env var that a reset destroys. Interim: every rehydration brief re-exports
the seat identity, and a rotated seat's first credential-touching action is treated as unverified until
it succeeds once.
D-33 — the tool used for "full step scans" silently prunes a step in its default output
coder-mos1 reported a 9/9 child-step scan of pipeline #2186 using -f json. My scan of the same
pipeline, using the same wrapper in its default text mode, showed 8. Verified:
| mode | steps enumerated |
|---|---|
pipeline-status.sh -n 2186 (text, default) |
8 — begins at ci-postgres |
pipeline-status.sh -n 2186 -f json |
10 — the ci workflow entry, clone, plus those 8 |
The default output omits clone. Every "full step scan" performed tonight — RM-01, RM-02 across four
pipelines, RM-03 — enumerated 8 of 9 real child steps, and reported that as complete. clone
succeeded in every case, so no verdict changes; the method was incomplete and its user did not know.
Why this is more than a cosmetic gap. The merge-gate mandate requires CI verified "by a full step
scan — a tests-complete step is a pruning-tell, never a completion-tell", and requires every verdict
to enumerate "pipeline id + the step-scan result (step count, any SKIPs…)". A verdict citing 8 steps
where 9 exist is non-conforming evidence — and if the gate uses this wrapper's default output, it will
produce exactly that while believing it complied.
The shape is the session's own thesis, aimed at the tool built to detect it. The doc's instruction is
literally "read the artifact, not the summary." The text mode is a summary. It looks like an
artifact — a labelled list of named steps with states — which is precisely why nobody checked it against
the raw record. A summary that resembles an enumeration is more dangerous than one that obviously
summarises. Compare D-24: mergeable was a true answer to a different question; here the step list is
a true answer to a smaller question.
Caught only because a seat used -f json and its count disagreed with mine. A disagreement between two
observers surfaced it; neither observer alone would have.
Requirements. RM-02 registers a case asserting the scan enumerates the full child-step set
(compare text-mode count against the JSON record, must-fail on divergence). The merge-gate must scan from
the JSON/API record, never the text summary, and state the step count it saw. The wrapper's default
output should either include clone or stop presenting itself as a complete step list.
D-32 — the repaired gate shipped with a documented switch to skip it
rev-974 blocked RM-03 (PR #1032) on a single finding, and it is the one that mattered:
pr-merge.sh still ships --skip-queue-guard.
Verified independently at five sites in packages/mosaic/framework/tools/git/pr-merge.sh — usage line
:3, help text :29 ("Skip CI queue guard wait before merge"), a worked example at :38, the flag
handler :58, and the conditional wrapping the entire ci-queue-wait.sh invocation at :119.
The reviewer did not report this from reading — it proved it. It replaced the fixture queue guard
with an exit 99 stub, ran a real fixture merge with --skip-queue-guard, and observed: the guard was
never called, the provider payload was created, the wrapper printed "merged successfully", exit 0.
Why this is disqualifying rather than a nitpick. RM-03 repairs a mandatory merge gate that has never
once been able to fail (D-23). Shipping that repair while leaving a documented, advertised,
worked-example flag that skips the repaired gate means the gate still cannot be relied upon — it
merely requires one flag instead of zero. Every merge-side CANNOT_ASSERT/HOLD semantic the task
built is bypassed by it. L0 is explicit: "Never use workarounds that bypass quality gates — --no-verify
and equivalent skip switches are off-limits." This is an equivalent skip switch, sitting in the merge path.
The merge-gate role doc names this exact hazard about force_merge, and the description transfers
without modification:
a
--no-verifyequivalent, adjacent to the field you came to edit, at the exact moment you are most motivated to reach for it — red gate, queued PR.
The generalizable rule. Repairing a gate is incomplete while any supported path skips it. A fix that restores a gate's ability to fail, but leaves a switch that turns it off, has changed the failure from "the gate cannot block" to "the gate can be told not to" — which under incident pressure is the same outcome reached one keystroke later. Removing the bypass is part of the repair, not a follow-up.
Required: remove --skip-queue-guard from merge-capable execution — flag, usage, help text and
example alike (an undocumented-but-working bypass is worse than a documented one) — plus a RED-first
regression proving a merge-capable invocation cannot reach a provider merge call with the guard unrun.
--dry-run may retain the skip since it does not merge, provided dry-run and merge-capable paths can
never both be true.
Everything else in RM-03 passed independent verification, including the 160 KiB stdin transport where #1023 regressed, the full classification matrix, the audited push-0/merge-75 tri-state, and the scope disposition. A strong PR with one hole in it.
D-31 — a seat ran past 100% context with 28 uncommitted files and no commits
coder-mos1 reached 102.5% / 372k context on RM-03 with 28 modified files uncommitted and its
branch still at the base commit — i.e. a full context window spent (~$27) with nothing durable. The
orchestrator caught it by polling seat state, not by any signal from the system.
Nothing warned anyone. No threshold alert, no pre-compaction hook, no "you are about to lose this" prompt. The seat did not report distress because from inside, running low on context feels like ordinary work right up until it is not. The one failure mode the mission was founded on — compaction destroying in-flight work — very nearly happened to the mission's own delivery, twice in one night (RM-01 was checkpointed at 69%; this one was already past the window).
What made RM-01 survivable and this one nearly not: RM-01 had committed work when it was checkpointed, so rotation was cheap. Here the work existed only in the worktree and the seat's context. Commit-then-rotate works; rotate-without-commit loses everything.
Requirements.
- RM-34 (rotation daemon): a context-threshold trigger is not enough — it must force a durable checkpoint before rotating, and refuse to rotate a seat with uncommitted work until that work is committed or explicitly discarded. Rotation without a commit is data loss with extra steps.
- RM-50 (fleet ownership): seat context level is observable state and must be monitored, with a threshold alert to the coordinator. Relying on an orchestrator to poll is exactly the "instructions, not enforcement" pattern (D-4) applied to lifecycle.
- Worker briefs (immediate, mechanical): commit early, commit WIP, an imperfect commit that exists beats perfect work that vanishes. Added to standing brief doctrine rather than left to judgement.
Adjacent observation — scope drift, unverified. The same diff spans 5 framework guides and 11 agent templates against a brief scoped to the queue guard and its tests. Queried rather than assumed: if the edits are consequential to the tri-state change they stay; if drive-by, they split. Recorded here because a 28-file diff from a brief that named one script is a scope signal regardless of the answer.
D-29 — the registry's own coverage clause was SYNTACTIC, not semantic
rev-974 blocked PR #1030 with a finding that lands in the keystone's own enforcement:
It moved
RM02-MEANING-PROVENANCEandRM02-PROSE-CONTROLfrom the stale-build-lock case to an unrelated type-error case, reformatted the manifest, andpnpm gate:verifystill exited 0.
So the registry verifies that a criterion has a binding, not that the bound case can fail for that
criterion's stated reason. That is exactly what RM02-REQ-03 (clause 2, from D-17) requires, and it
is not satisfied — gates/gates.manifest.json:471-488 + scripts/gate-verify.mjs:410-417.
The keystone reproduced the defect it was built to eliminate, in its own coverage check. D-17 said a green suite can leave a criterion untested; here the tool that enforces that clause enforces only its syntax. A binding that can be pointed anywhere without complaint is a binding that proves nothing — D-24's "true answer to a different question" wearing a manifest.
Caught only because the reviewer mutated the binding and re-ran, rather than reading the manifest and confirming the entries existed. Inspection would have passed it.
D-30 — the orchestrator fed a reviewer a false comparator supporting its own conclusion
In the review brief for #1030 I stated the ci-postgres FAIL had "appeared under overall-success on
#2167/#2170/#2175". rev-974 refuted it: #2167's ci-postgres step is OK. Verified —
#2158 OK, #2167 OK, #2170 FAIL, #2175 FAIL, #2182 FAIL.
My own banked D-21 says so explicitly: "Observed on #2170 and #2175 (not on #2158, #2167, or the manual #2171)." I contradicted my own written record — because I restated it from memory instead of reading it. That is D-26 recurring, committed by the orchestrator, in a reviewer brief, hours after promoting render-not-restate to the charter.
The specific harm is worse than the inaccuracy. I supplied a reviewer with fabricated supporting
evidence for my own reading, in a brief that also asked it to accept my ci-postgres interpretation.
Had it deferred, the false comparator would have propagated into the review record and strengthened a
conclusion it never supported. It refuted me instead — which is the query-for-refutation rule paying for
itself in the same document that introduced it.
Requirement. A brief that supplies evidence must cite it from the record, not from recall — findings are quoted with their identifier, not paraphrased. RM-02 registers it: where a document cites a finding, the citation must match the banked text, with a must-fail control on divergence.
D-28 — a swallowed diagnostic destroyed the evidence a fail-closed check needed
RM-02's CI run failed on four scripts/gate-history.test.mjs sandbox tests. The fail-closed logic
behaved correctly; the defect was that it was denied the evidence to decide.
Chain: Bubblewrap is installed by the gate step but not the test step, so spawnSync('bwrap', …)
returned status=null / ENOENT. replayCommit then replaced the spawn diagnostic with an empty
string (historical frozen dependency install failed:). With the underlying error destroyed, the
classifier could not prove Bubblewrap provenance — and, correctly, refused to treat an
unprovable condition as expected sandbox unavailability.
The refusal was right. The information loss was the bug. A fail-closed check is only as good as the evidence reaching it: strip the diagnostic and a correct classifier is forced into a correct-but-opaque refusal that looks like a defect in the thing being tested. Error text is not decoration — for a classifier it is the input.
Fix requirements (and what review must scrutinise): preserve install.error.message in replay
diagnostics, and accept only Bubblewrap-provenance EPERM/EACCES/ENOENT as terminal sandbox
refusal — unrelated command errors stay rejected. Widening the acceptance to make the test pass would
be a real finding, not a fix: it would convert a precise fail-closed check into a permissive one.
Two linkages worth recording:
- D-16 again — it did not reproduce locally because
bwrapis installed there. Local and CI disagreeing about what passing means, a third time. - Diagnostic-preservation is a gate requirement, not hygiene. RM-02 registers it: where a check classifies on the basis of an error, the case must assert the diagnostic survives to the classifier, with a must-fail control proving a swallowed message is detected.
Process note. The orchestrator's hypothesis — that the coincident ci-postgres FAIL caused the test
failure — was wrong, and refuted with direct evidence: the log shows ci-postgres:5432 - accepting connections and migrations completing. D-21 therefore stands unchanged as a teardown artifact and is
NOT upgraded. It was about to be re-classified as "intermittently takes out the test step" on a false
premise; asking the seat to confirm or refute rather than accept is what prevented a finding being
corrupted by a plausible guess.
D-27 — the authoritative role file certifies an inert control as ✅ enforced
Found by rev-974 while loading the canonical gate sources — i.e. found because we switched from
restating to reading (D-26). Verified independently in
~/.config/mosaic/fleet/roles.local/merge-gate.md:
| line | text | problem |
|---|---|---|
| :59 (mandate 3) | verify "CI queue guard clear" from primary evidence | mandates verifying a check that cannot fail |
| :75 (mandate 4) | every durable verdict must enumerate "the queue-guard outcome" | requires a meaningless field in the evidence record |
| :128-129 | coordinator merge step 2 runs ci-queue-wait.sh --purpose merge; annotated "real" |
the sensor never fires |
| :290 | control table: "CI queue guard | ✅ enforced | … genuinely aborts" | a security-control table certifying an inert control |
The precise defect, because the distinction matters. The document is correct about the wiring and
wrong about the control. set -euo pipefail with an unguarded exit genuinely does abort — that
plumbing is real. What is false is the conclusion ✅ enforced, because the guard cannot produce a
non-zero exit for any input (D-23). So this is a mechanism correctly wired to a sensor that never
fires, documented as enforced in the governing role file — the inert-gate class, certified.
Note the compounding: the verdict format requires the queue-guard outcome as enumerated evidence. A conforming verdict is therefore obliged to include a field that means nothing — the role file instructs the merge-gate to manufacture evidence from an inert check, in the same mandate that exists to stop bare conclusions being recorded as verdicts.
Interim resolution (confirmed with rev-974, with one refinement). Follow the stricter D-23 ruling.
But do not omit the queue-guard field — omitting a required field makes the verdict non-conforming.
Record it, labelled ZERO-INFORMATION (inert, owner RM-03). That satisfies the enumeration mandate
and states the truth simultaneously; a silently-omitted field and a silently-cited one are both worse.
Not ours to edit. roles.local/merge-gate.md is an operator-owned override; the correction belongs
to the coordinator. Recommended fix is annotation, not deletion — the requirement is correct once
RM-03 lands, so removing it would create a gap the moment the guard starts working. Annotate the three
evidence sites and correct :290 to distinguish wiring real from control inert.
This is D-20's clause meeting an authoritative framework document. Prose asserting a security
property needs a negative control observed red; ✅ enforced is exactly such an assertion, and it would
fail that control today. The rule was written after the orchestrator overclaimed in its own governing
document — it now catches the same defect in the framework's.
D-26 — render-not-restate failed at maximal attention, committed by the rule's own enforcer
The definitive justification for making render-not-restate mechanical.
The mission's delivery-gate definition was incomplete from setup: the merge-gate verdict step was
missing. Root cause, stated by the coordinator: when MISSION.md / KICKSTART.md were written, the
gates were restated from memory — "author ≠ reviewer, diff-blind checks, CI-green, merged-PR
completion" — a subset of an authoritative document that was not open at the time. The omission then
propagated into every worker brief issued for the rest of the mission.
Then it happened again, in the correction. The message correcting the restatement error was itself a restatement: delivered from memory, without the source open, and it dropped or altered five things —
| stated | actual (fleet/roles.local/merge-gate.md) |
|---|---|
verdict is APPROVE/BLOCK |
GO / NO-GO / HOLD |
| — | verdict is commit-bound and VOID the instant the head moves; HOLD persists; an empty-commit answer to NO-GO must be re-issued as NO-GO |
| — | the gate verifies mechanical gates itself from primary evidence (reviewed-SHA == CI-SHA == head, full step scan, Closes #N, queue guard) |
| — | the verdict must be posted durably, enumerating evidence; a bare "GO — gates verified" is non-conforming |
cite guides/CANONICAL-DELIVERY-CYCLE.md |
that file does not exist — the path was taken from another brief and cited without checking it resolves |
And the omission was safety-relevant. The role doc carries a hard install precondition: an unminted gate seat's credentials fail open to the owner's admin account, so its verdicts would post as admin — "the exact mechanism of the #3084 breach this seat exists to prevent." Verified: gate identities existed for USC only, none for mosaicstack (D-15's class, on the seat that authorises merges). Assigning the gate at that moment would have reproduced the breach it was created to stop.
Why this is the definitive case, and not just a third instance. The failure occurred when every condition favoured success: attention was maximal (a deliberate correction), the actor knew the rule best (its author and enforcer), and the subject was the rule itself — the message was about render-vs-restate while committing it. A discipline that fails under those conditions does not fail from carelessness; it fails because human recall of a structured document is lossy by construction, and the loss is silent.
A rule that its own enforcer cannot follow while enforcing it is not a discipline. It is a requirement for a mechanism.
Requirement. Governing documents must reference authoritative sources, never paraphrase them, and every reference must resolve — a dangling pointer to a non-existent authority is worse than a restatement, because it reads as authoritative and cannot be checked. RM-02 registers both: a link-resolution case (every cited authority path exists) with a must-fail control, and the D-20 clause that prose asserting a property needs a negative control.
Resolved: MISSION.md and KICKSTART.md now reference fleet/roles.local/merge-gate.md and
fleet/roles/validator.md and state only the gate order. The precondition was cleared properly —
gitea-mosaicstack-merge-gate.token minted at least privilege, independently verified by this
orchestrator as the seat's own token: pull=True, push=False, admin=False — it structurally cannot merge,
which is what a verdict authority must be.
D-25 — self-verification by the audited party, a second time, one layer down
RM-02's per-commit replay needs isolation. Docker denied Bubblewrap namespace creation; the obvious fix
— a privileged CI step — was correctly refused by the implementing seat and rated CRITICAL by a
security review, because PR-controlled package.json/verifier code executes before Bubblewrap
establishes any boundary. The untrusted party would obtain the very capability meant to contain it.
UID 65534 and later capability drops do not repair an ordering defect.
The seat's own formulation, promoted here verbatim because it is the whole argument:
"Repo-only code cannot both grant namespace capability to PR config and prevent that same PR from using the capability directly."
This is D-19 again, one layer down — and the repetition is the finding.
| audited party controls… | so verification fails at… | |
|---|---|---|
| D-19 | the manifest that certifies it | artifact integrity |
| D-25 | the code that enters the sandbox | execution integrity |
Both reduce to: self-verification by the audited party is not verification. Both resolve the same way — an authority outside the audited party's control. That is now twice this mission has independently derived the necessity of its own architecture (Builds 1–2: choke-point executor + spine) from a security impossibility rather than from design preference.
Ruling — Option C, reached independently by the orchestrator and by rev-974 before either saw
the other's reasoning (the same convergence signal the two planners produced). Ship what the layer can
guarantee — unprivileged, fail-closed, current-tree verification on PR CI — and defer isolated
per-commit replay to a protected authority.
Condition from rev-974 that the orchestrator's ruling missed, and it matters most: Protected post-merge replay is DETECTION, not PREVENTION. It must be documented as such, with its response defined (quarantine / revert), and must never be presented as equivalent to a pre-merge gate. Presenting detection as prevention would be D-20's overclaim in a new location — and would be especially corrosive here, since the whole task exists to make gates trustworthy.
Also binding: inability to establish the sandbox must be a hard non-zero, never skip-or-pass;
the meta-negative-control must not be weakened to fit the unprivileged environment (any gate whose
mutation check genuinely cannot run unprivileged is registered as a DEFECT with an owner, never
quietly dropped); and the privileged experiment stays uncommitted, out of branch history.
Tracked dependency: RM-60 — an external pre-execution trust boundary (protected default-branch pipeline config, or an immutable trusted launcher) owned by infra/provider authority, cross-referenced with RM-59. Both are the same dependency wearing different clothes: the anchor must live outside the audited party.
D-24 — the orchestrator cited mergeable: true as evidence of merge-readiness
Escalating D-23, I wrote that the fix "exists and is parked, open and mergeable" — framing that invited the reading merge this and the gate is fixed. Mos verified #1023 before relaying, and corrected me.
mergeable: true is git-mergeability: the absence of textual conflicts. It says nothing about review
state, CI, or correctness. Verified independently after the correction: review id-58,
REQUEST_CHANGES, at commit f6334080 — the current head. The block is live, and its findings are
that #1023's own fix is defective.
This is the charter's first principle — observe the property, not the proxy — committed by the orchestrator, inside an escalation about a proxy being mistaken for a property. I spent the session demanding that discipline of others, then read a JSON field whose name resembled the question I cared about and reported it as the answer. Fifth instance of mine.
Worth naming precisely, because it will recur: mergeable is not a lie — it is a true answer to a
different question. A false field would have been caught. A true-but-adjacent field passes every sniff
test, which is exactly why the discipline must be mechanical rather than attentive.
Requirement on RM-02. A registered case asserting readiness must assert the decision-relevant property — latest review state, CI terminal status on the exact head, required approvals — never a transport-level or structural proxy. A case that reads an adjacent field is D-17 (the set does not cover) in a convincing disguise.
D-23 — the mandatory CI queue guard has never worked, for any input, ever
RM-02's registered fixtures found this before the registry was even built — the sharpest possible argument for the task. It is materially worse than the already-banked "unknown ⇒ exit 0" (D-6/D-10).
ci-queue-wait.sh:266 pipes the provider payload into a classifier that begins python3 - <<'PY'.
python3 - reads its program from stdin, and the heredoc supplies that program — so
json.load(sys.stdin) never sees the piped JSON, EOFs, and falls through to unknown.
Proven empirically, not inferred:
payload={"state":"success"} -> unknown
payload={"state":"failure"} -> unknown
payload={"state":"pending"} -> unknown
payload=not-json -> unknown
Every input classifies as unknown; unknown exits 0. So the mandatory pre-push / pre-merge guard
returns PASS for every possible input — including a genuinely FAILING CI. It has never distinguished
anything. My six observed "meaningless greens" were not an edge case; they were the only output it can
produce.
A fix was ATTEMPTED and is parked — but it is NOT ready. PR #1023 — "parse the payload it was handed, not stdin" —
diagnoses the defect precisely in its own body: "binds stdin to the program text, so
json.load(sys.stdin) EOFs, the state falls through to unknown, and the guard exits 0 on every
invocation, every branch, every repo, both platforms."
Two aggravating details from that PR's own body:
pr-ci-wait.sh:38documented this exact bug and its remedy — and it was never backported to the mandatory sibling. A known defect, with a known fix, written down in one tool and never applied where it was load-bearing. That is D-14 propagation failure with a security-adjacent blast radius.- It was "opened per board instruction without re-verification" — the fix for an unverified gate was itself shipped unverified.
⚠ CORRECTED (D-24), propagated backward per D-14-as-amended. This entry originally read "The fix already exists and is parked … open, mergeable" — framing that invited merge it and the gate is fixed.
mergeable: trueis git-mergeability — the absence of conflicts — not readiness. Verified independently: #1023's latest review is rev-974 id-58,REQUEST_CHANGES, at commitf6334080, which IS the current head — the block is live, not superseded. That review found #1023's own fix defective:unknownstill exits 0, payload-as-argv hits ARG_MAX, and its tests do not assert exit outcomes. So "merge #1023 and the guard is fixed" is false twice over.
Escalation, for Jason, on #1023's disposition. Every merge gated on this guard since it was introduced was ungated. The guard cannot fail. This does not mean those merges were bad — CI itself ran and was checked by humans and reviewers — but the automated gate contributed zero signal throughout, while being cited as merge evidence. The standing doctrine ("treat its green as zero-information") was more literally correct than anyone realised when it was written.
RM-02 records required vs actual per case with DEFECT (owner: RM-03) and stays sensitive to any
change in either direction. No RM-03 lane was opened; ci-queue-wait.sh was not edited; source and
installed copy remain byte-identical at the observed SHA-256.
D-22 — the registry would have verified the wrong artifact
Surfaced while ruling RM-02's design. The gate that actually runs when the fleet invokes a wrapper is the installed copy:
~/.config/mosaic/tools/git/ci-queue-wait.sh ← what executes
packages/mosaic/framework/tools/git/ci-queue-wait.sh ← what a repo-scoped registry would test
Checked: today they are byte-identical (sha256 19cda2f7009c536e both). But nothing asserts that.
There is no check, anywhere, that the deployed gate matches its source.
A registry that verifies the wrong artifact is worse than no registry, because it manufactures confidence: every case would go green against a file that is not the one enforcing anything. Note this would have been invisible from every signal we have — the registry's own tests would pass, CI would pass, and the deployed gate could be arbitrarily different.
This is the D-1 / P-ACTIVATION class applied to the gates themselves: the same "committed config that only works on one runtime" defect, one level up, where the enforcement mechanism rather than the config drifts from source.
Requirement on RM-02. For every registered gate with a deployed counterpart, a case asserting repo-source == deployed-copy, with a must-fail control (mutate the deployed copy ⇒ detected). Where a gate has no deployed counterpart, record that fact explicitly rather than leaving it ambiguous.
Caught before implementation began — the only finding tonight found by design review rather than by execution, which is an argument for the design-first gate on keystone tasks paying for itself.
D-21 — a service step reports FAIL under an overall-success pipeline (normalised red)
Observed on pipelines #2170 and #2175 (not on #2158, #2167, or the manual #2171): the
ci-postgres service step reports FAIL (pods "wp-svc-…-ci-postgres" not found) while the
pipeline's overall status is success and all eight functional steps are OK.
Most likely mechanism: the service pod is reachable during the run and has been reaped by the time its
terminal state is collected. Supporting evidence that the database really was available — the test
step runs pnpm --filter @mosaicstack/db run db:migrate behind a pg_isready loop written to fail fast
if Postgres never comes up, and that step passed.
Severity is low; the pattern is not. A red that is routinely present and routinely correct to ignore trains every operator and every agent to discount reds — and the discounting generalises to reds that matter. This mission already has six instances of a green that meant nothing (D-5, D-6, D-10); a FAIL that means nothing is the same defect with the sign flipped. Both erode the signal that gates exist to carry.
Not chased further because it is CI-infrastructure behaviour rather than repository code, and both merges it touched had every functional step green with the overall status successful. Recorded rather than normalised, which is the whole point: the first time this is waved through without a note is the moment it becomes background noise.
Requirement on RM-55 (conformance harness): treat a non-terminal-success sub-step under an overall-success pipeline as a reportable anomaly, not a cosmetic artifact — and make the CI contract state explicitly which steps are permitted to fail without failing the pipeline. An implicit allowance is indistinguishable from a bug.
D-20 — the orchestrator's own documentation overclaimed, and a reviewer disproved it empirically
rev-974 blocked PR #1027 a second time. The defect was not in the code — it was in this file, at
D-18's entry, written by the orchestrator.
Two faults, both mine:
- D-18's AC2 restatement omitted the scope clause that D-19 later established as mandatory ("within an accidental/independent-mutation threat model").
- D-18 asserted that the tampered-manifest control turns integrity "from a claim into a property." It does not, and cannot. That sentence was written before D-19 proved the property impossible at this layer, and was never revised when D-19 landed.
The reviewer did not merely read it — it disproved it. It performed a same-UID consistent manifest + marker rewrite, and preflight passed. My documented claim was falsified by experiment. Code, README, scratchpad and PR body all stated both threat-model directions correctly; this file was the only place still overclaiming.
This is two banked findings firing on the orchestrator at once:
- The integrity-claim corollary — I wrote an integrity claim in the voice of an integrity property, in the very document that defines the rule against doing so.
- D-14 (propagation) — D-19 superseded D-18's assertion. I propagated the consequence into the charter and the delivery conditions, but not back into D-18 itself. A ruling that fails to propagate backwards into the finding it supersedes is the same defect as one that fails to propagate forwards, and I did not audit for that direction.
Corrected in place, with the original wording quoted and the empirical disproof recorded, rather than silently rewritten — the same standard demanded of any restated criterion.
Requirement on RM-02 (fifth clause). Documentation asserting a security or integrity property is itself a claim requiring a negative control. Where a document states "X is guaranteed", the registry must hold a case that fails if X is not guaranteed — and that case must have been observed red. Prose is not exempt from the mission's own evidentiary standard, and prose in the governing document least of all: it is the artifact most likely to be quoted as authority long after the code has moved on.
Reviewer credit. rev-974 was briefed that its highest-priority check was "confirm the PR claims no more than it can deliver, and a softened or omitted boundary is a finding even though the code works." It applied that instruction to the orchestrator's own governing document and produced an experiment to settle it. That is the review standard this mission is trying to make ordinary.
D-19 — an integrity property that cannot exist at the layer it was specified
Implementing D-18's manifest, the seat + a Codex security review reached CWE-345: the symlink manifest and the source-hash marker both live in the same same-UID writable generated tree, so an actor with that UID can plant a rogue link, regenerate both, retain the fingerprint, and pass. No local cryptographic construction fixes self-authentication without a key outside that actor's authority; relocating the marker changes the path, not the authority.
The seat escalated rather than describing self-authentication as tamper-resistant — the explicit failure mode the charter corollary demands. That is the corollary working, on its first real test.
Ruling — Option A: scope AC2 to accidental / independent / stale mutation; retain the design. Rationale, recorded so it can be challenged:
- The undefendable boundary is not the weak link. An actor with same-UID write can already edit the
source, the tests,
scripts/preflight.mjsitself, and.husky/*. If they have that, nothing in the local checkout is trustworthy — hardening the manifest buys no real security while implying protection that does not exist, which is worse than the gap. - What AC2 is actually for. These checks exist because a five-month-stale
.nextproduced 19 phantomTS2307errors indistinguishable from real ones (D-5). That is staleness, drift and foreign residue — and against that class the design demonstrably works. - A real trust anchor arrives later, from this mission's own architecture. An anchor must live outside the actor's authority; for a fleet running as one user that means a separate service — precisely the choke-point executor + PG spine of Builds 1–2, which verify outside the worktree's authority. Hand-rolling key distribution for a local preflight now would duplicate that work badly.
- Option C (structural policy, no manifest) is strictly worse — it cannot detect a removed expected link.
Option A is acceptable only with honest labelling, or it becomes the disease it is meant to cure. Conditions (last two added/sharpened by Mos):
- Threat model stated verbatim in the code and the PR; the words tamper-proof / tamper-evident / secure barred from that context; the scope carried in AC2's restatement; every control kept RED-first including manifest-only tamper.
- State the boundary in BOTH directions. Not only what it does not defend (same-UID write; no
local construction can) but, beside it, what it does defend: accidental / independent / stale /
foreign-residue mutation — the D-5 class it was born from (the five-month
.nextand its 19 phantomTS2307s). A reader who sees only the negative dismisses the check as worthless; one who sees only the positive over-trusts it. Both together is the honest artifact. - The residual risk is a HARD TRACKED DEPENDENCY EDGE, not a comment. It is RM-59, owned by the
choke-point executor + spine work (
depends_on: RM-12, RM-21, RM-25), and the AC2 scope note must cite that id. "Record where the real guarantee comes from" only holds if the record is a live dependency someone must close. A documented gap with no owner becomes a permanent gap that reads as intentional.
The generalizable rule. When a required property cannot exist at the layer where it was specified, the honest moves are: implement what the layer can guarantee, state the boundary precisely, and record where the real guarantee will come from. A known gap that is written down is acceptable; a gap that is implied fixed is not. Silence here would have shipped a verification artifact that verifies nothing — with a green to prove it.
D-18 — two pre-registered criteria were mutually unsatisfiable, discoverable only at implementation
Implementing D-17's fix surfaced a conflict between pre-registered criteria:
- AC2 (as written) — reject symlinked generated state.
- AC4 — the canonical
pnpm -w buildsucceeds and leaves no residue.
Verified independently rather than taken on report: apps/web/next.config.ts:4 sets
output: 'standalone', and the built tree contains 42 legitimate pnpm dependency symlinks under
.next/standalone/node_modules. A blanket descendant-symlink rejection makes the canonical build fail
its own preflight with exit 43. AC2 read literally is unsatisfiable alongside AC4 under this
configuration, and nothing short of building the tree would have revealed it.
Third distinct failure mode of a pre-registered check set, completing the chain:
| finding | a pre-registered check set can be… |
|---|---|
| D-8 | wrong — a check that does not test what it claims |
| D-17 | incomplete — green while a criterion's requirement is untested |
| D-18 | internally inconsistent — two criteria that cannot both hold |
The implementing seat escalated instead of silently picking a winner. That matters: quietly resolving a conflict between pre-registered criteria destroys the point of pre-registering them — the registration exists so that changes of meaning are auditable rather than absorbed.
Resolution (orchestrator ruling). Approved a build-certified symlink manifest: .next itself is
still rejected as a symlink; descendants are rejected unless exactly certified by a manifest the build
publishes atomically. Strictly stronger than blanket rejection — it also catches a retargeted
symlink, which blanket rejection cannot distinguish from a legitimate one.
AC2 restated (recorded, not absorbed). Generated state must reject .next itself being a symlink
or non-directory, and must reject any descendant symlink not exactly certified by the build manifest —
added, removed, retargeted, or manifest-only-tampered all fail with exit 43 — within an accidental /
independent-mutation threat model.
⚠ This entry is superseded in part by D-19. Do not read D-18 standalone. The scope clause above is load-bearing: the design cannot defend against a same-UID actor, which can rewrite the manifest and the marker consistently (CWE-345). D-18 was written before that impossibility was established.
Hardening required before this counts. The manifest is itself generated state, so a manifest writable by whoever plants a rogue symlink certifies the attack — that is the one way this design fails. It must sit inside the same ownership/fingerprint envelope, published atomically via the existing marker mechanism, with negative controls observed red first for: added, removed, retargeted, manifest-only-tampered, plus a positive control that the canonical build passes.
⚠ CORRECTED (D-20). This paragraph originally ended: "without it, integrity is a claim rather than a property." That overclaimed, by implying the control makes integrity a property. It does not, and cannot. The manifest-only-tamper control detects independent mutation of the manifest; it confers no authenticity against an actor who rewrites manifest and marker together.
rev-974disproved the original wording empirically — a same-UID consistent manifest+marker rewrite passed preflight. Integrity here remains a scoped drift-detection property, never an authenticity one. See D-19 and the charter principle on properties that cannot exist at their specified layer.
Requirement on RM-02 (fourth clause). The registry must detect conflicts between registered criteria, not only wrongness and coverage. Two criteria that cannot simultaneously hold is a registry defect discoverable by construction — and when a criterion is restated, the registry must retain the original text, the restatement, and the reason, so the evolution stays auditable.
D-17 — a pre-registered criterion passed a green suite without being satisfied
rev-974 returned CHANGES REQUESTED on PR #1027 with one blocking finding, and it is the sharpest
instance of the session's theme because it occurred inside our own verification machinery.
AC2 was pre-registered before any code was written, and explicitly required that symlinked generated state be rejected. The implementation does not do it:
ln -s /etc/hosts apps/web/.next/reviewer-symlink
pnpm preflight # → "checkout preflight passed", exit 0
# → required: generated-state exit 43
The acceptance suite was 21/21 green throughout. Confirmed independently rather than relayed:
scripts/preflight.mjs:82-92 rejects symlinks on the source path; :28 merely skips symlinked
directories rather than rejecting them; and the generated-state path at :141-163 lstats and
checks uid (ownership) but never calls isSymbolicLink(). The suite's only symlink cases
(preflight.test.mjs:59, :115) cover the turbo binary and a source file. No generated-state case
exists anywhere.
So: criterion pre-registered, suite green, requirement unmet. Nobody was careless — the coverage gap is invisible from a green, which is the entire problem.
This sharpens D-8 rather than repeating it. D-8 established that pre-registration does not confer correctness (a check can be wrong when written). D-17 establishes the adjacent failure: pre-registration does not confer coverage — a suite can be green, and every registered criterion can appear satisfied, while a criterion's actual requirement is untested. The two together mean a registry of checks needs two properties, not one: each check must be right, and the set must actually exercise what it claims.
Requirement on RM-02 (third clause). The registry must bind each acceptance criterion to the specific case that exercises it, and prove that case red before trusting its green. A criterion with no case that can fail for that criterion's stated reason is unregistered in substance however it appears in the manifest. This is mutation testing pointed at the criterion-to-case mapping, not merely at the gate.
Credit where due: the reviewer also declined to re-run AC8, stating plainly that the PR carried it forward with no runnable command rather than silently substituting a different boundary test. That is the D-8 clause working a second time, in the same review that produced D-17.
D-16 — the local test gate and the CI test gate disagree by environment
Mos flagged a shape worth chasing: if pnpm test exits non-zero on a pre-existing guard, then either
main is red and merges step around it (the #868 shape again), or CI does not run that path. Both
branches turned out wrong, and the truth is a third thing. Established by running it, not by asking:
CI runs exactly pnpm test (.woodpecker/ci.yml, test step) — the same command. So the path is
exercised. Yet:
| environment | result |
|---|---|
| CI container | test step green (#2158, #2167) |
| this host, clean worktree | exit 97 — WAKE-ASSERT INIT ABORT: BASH_LINENO convention violated on bash 5.2.15(1)-release … expected [3 4], probe reported [3 5] (#973), after PASS=18 FAIL=0 |
| this host, main checkout | exit 1 — a different, second defect (below) |
The guard is environment-dependent: it aborts on this host's bash and not in CI's container. git diff origin/main... confirms PR #1027 touches zero files under packages/mosaic, so the guard is
genuinely pre-existing and unrelated — f10-coder's report was accurate in every particular, and
main is equally affected on this host.
So it is not "merges step around a red" — it is worse in one specific way: the local gate and the CI
gate do not agree about what passing means. No agent on this host can obtain a green pnpm test at
all, on any branch, including main. A gate an operator cannot run is a gate that only CI enforces,
and a gate only CI enforces cannot be a pre-push gate. This is the hermeticity/portability class
already live as #1007 (PR #1024).
Second, independent defect found while establishing the above. In the main checkout the same
package fails differently — exit 1 — because a test scans the working tree and asserts on files it
finds, picking up apps/coordinator/venv/** (third-party site-packages: pi = math.pi in rich,
setuptools, mypy). A test whose result depends on untracked files present in the tree is not
hermetic. This is the same contamination source that made pnpm format:check unpassable (D-1/D-7
hygiene) — one untracked foreign tree silently breaking two independent gates.
Requirements. RM-01/RM-04: a gate must produce the same verdict on a developer host and in CI, or declare loudly that it cannot run here — never diverge silently. RM-02 registers both as cases: the environment-divergence guard, and a hermeticity control asserting a suite's verdict is unchanged by the presence of untracked directories. Coordinate with #1007/#1024 rather than opening a third lane.
Ownership (Mos, 2026-07-31). The hermeticity fix is PR #1024, which sits in Jason's parked
delivery stack — so, like #1023, its disposition is Jason's. Marked SUPERSEDED-PENDING-JASON
alongside #1023. D-16 strengthens the urgency but does not transfer ownership: we do not open a third
lane on a parked PR. The one-line escalation for Jason: two independent gates
(format:check, pnpm test) were broken by a single untracked directory, and a third
(pnpm test) disagrees between host and CI — non-hermetic gates make every green host-dependent.
Sharpened statement of the class (Mos). A pre-push gate an operator cannot run locally is a gate
only CI enforces — so pointing .husky/pre-push at it misrepresents where the gate lives. Combined
with the shared root cause across two gates, the finding is: non-hermetic gates make every green
host-dependent. A gate that only appears to pass depending on which host runs it is this mission's
exact subject, one meta-level up.
#1027 disposition (Mos): proceeds on CI-green. CI is the authoritative gate; the local exit-97 is a known host-specific guard abort, irrelevant to the merge decision.
D-15 — token scope is not repository permission (a THIRD capability layer)
f10-coder was provisioned with gitea-mosaicstack-f10-coder.token, scopes write:repository +
write:issue, and the mint was verified by "repo access returns 200". It then failed to push:
remote: error: User permission denied for writing.
remote: error: pre-receive hook declined
Verified objectively rather than inferred (per the charter principle):
| probe | result |
|---|---|
GET /repos/mosaicstack/stack/collaborators/f10-coder |
404 — not a collaborator |
| repo permissions as seen by its own token | admin: false, push: false, pull: true |
Capability has at least three independent layers, and satisfying two proves nothing about the third:
- Token file exists → raw-API authentication works (D-11b).
tealogin exists → tea-dependent wrapper paths work (D-13).- Repository permission granted (collaborator/team membership) → writes are actually authorised.
A token can carry write:repository scope and still be refused, because scope bounds what a token
may attempt; repository permission decides what the user may do. They are different systems.
This is the charter principle failing on the very check meant to confirm capability. The mint was
validated by an HTTP 200 on a read. A 200 proves reachability; it does not prove the property that
was required, which was write. Both the provisioner and I accepted it — the same
written-unverified treated as verified as D-12, one layer up, on a check whose entire purpose was
verification.
Requirements. RM-50's pre-dispatch capability check must probe the effective permission for the
operation intended — for push authority, assert permissions.push == true as that seat, not token
existence and not a 200 on a read. RM-04's registry reconciliation covers all three layers, with a
must-fail control for each. A capability check that cannot fail on a seat lacking write permission is
itself an inert gate.
D-14 — a ruled decision did not propagate to the authoritative record
DECISION-1 (the corrected choke-point wire-in target) was ruled by the coordinator and applied to
TASKS.md. MISSION.md — the charter, the document a cold-starting seat reads first — kept the
superseded target for hours. It was flagged CONTESTED in a board note, then the ruling landed and
nobody edited the charter. A seat resuming from the charter would have read the rejected target as
authoritative and wired the choke point into a disabled rail — the precise failure the ruling existed
to prevent.
Caught by hand, during an unrelated edit. Nothing would have caught it otherwise.
This is P-MISSION-001 turned on ourselves. The mission's own thesis is that convention exists and enforcement is the gap: a decision that lives in a chat ruling and a board note, but not in the source of truth, has not actually been made — it has been agreed. The two are different, and the difference is exactly what this mission is about.
⚠ AMENDED by D-20 — propagation is BIDIRECTIONAL. As first written, this requirement was read by both the orchestrator and the coordinator as forward propagation only: a ruling reaches the documents that state the new rule. D-20 proved that insufficient. When D-19 superseded part of D-18, the consequence propagated forward into the charter and the delivery conditions but never backward into D-18 itself, which went on asserting a withdrawn claim — and a reviewer disproved it by experiment. A supersession must update BOTH the documents that render the new rule AND the finding it retires, with the retired wording quoted rather than deleted. Backward propagation is the same defect as forward; neither of us audited that direction until it bit.
Requirement (not merely a fix). A ruled decision must propagate mechanically to the authoritative record; it must not depend on someone remembering to edit a second file. Concretely, once mission state is DB-backed (RM-53 / the P-MISSION cutover):
- a decision is a record, not prose duplicated across documents;
- documents render decisions rather than restating them, so there is one place to be wrong;
- and where duplication is unavoidable, a check asserts the authoritative record and the derived document agree — with a must-fail control proving divergence is detected.
Until then, the interim rule: the same commit that records a ruling updates every document that states it — and every finding it supersedes. Interim rules are exactly what the DB cutover exists to replace.
The rule found a second instance within minutes of being written. Auditing the charter against all
rulings to date (rather than waiting to be bitten again) surfaced that DECISION-2 had also not
propagated: MISSION.md's standing directives still stated the DB hard-cutover with no mention of
Mos's binding qualification that the spine must not be a single-point hard-stop (degraded mode +
rollback artifact required). A seat reading the charter would have designed toward an availability
posture the coordinator had explicitly rejected — and would have found the superseded
"no DB ⇒ the fleet stops" recommendation nowhere contradicted. Now corrected in place.
Two un-propagated rulings out of two rulings that touched charter text. The propagation gap is not an oversight that happened once; without a mechanism it is the default outcome. That is the argument for making this a requirement rather than a discipline.
D-13 — two credential registries that can disagree (why the --draft fallback fired at all)
Diagnosing D-12's root cause surfaced a distinct defect. There are two parallel credential registries, and capability in one does not imply capability in the other:
| registry | contents for identity mos-dt-0 on mosaicstack |
|---|---|
token files — ~/.config/mosaic/secrets/gitea-tokens/ |
gitea-mosaicstack-mos-dt-0.token EXISTS |
tea login list |
NO mosaicstack login for mos-dt-0 (only mosaicstack-mos and mosaicstack-rev-974) |
So get_gitea_token succeeds and every raw-API path works, while every tea-dependent wrapper path
fails its login validation and silently degrades to the API fallback — which is exactly what dropped
--draft. tea is not "stale"; the login simply does not exist for that identity.
This matters beyond one flag: capability was declared authoritative by the token-file set (D-11b), but that registry does not govern the tea path. A seat can be fully provisioned by the authoritative registry and still lose functionality with no error — only a warning, and only on the degraded path.
Requirements. RM-04 (activation coherence): the two registries must be reconciled — one source of truth, or a startup check asserting they agree, with a must-fail control proving disagreement is detected. RM-50: the pre-dispatch capability check must verify capability for the path actually used, not merely token-file presence.
Confirmed working despite the gap (so this is degradation, not outage): pushes, pr-merge.sh,
PR/issue creation via API fallback, comment posting, and all reads. Impact is confined to
tea-only features — --draft, --labels, --milestone.
Reconciliation run by Mos (the manual form of RM-04's assert-agreement, done once by hand). For
git.mosaicstack.dev, the token-file registry holds six seats; tea holds logins for two:
| state | seats |
|---|---|
| token file present, no mosaicstack tea login | f10-coder, jarvis, mos-admin, mos-dt-0, pepper |
| token file present and tea login present | rev-974 (only) |
Five of six provisioned seats are silently degraded on tea-only features. This is systemic, not a one-off — which is why the fix is registry reconciliation (RM-04) and not a per-seat mint. Minting one seat would clear a symptom and leave the class live.
Mos deliberately deferred the mint: it is not on RM-01's critical path, and additively editing shared
tea config underneath running work is a change he declined to make without cause. Full remediation —
mint the five missing logins and wire the startup must-fail assertion that detects disagreement —
lands as RM-04 at a non-critical seam, or immediately if any seat needs a tea-only feature to progress.
Correction of record: this supersedes D-11(b)'s claim that the token-file set is the authoritative capability registry. It is necessary but not sufficient. Capability is per-path: the token file governs the raw-API path, the tea login governs the tea path, and the two can disagree silently.
★ D-11b ADDENDUM — THE FALSE-NEGATIVE TWIN (found 2026-08-01, banked by Mos). D-11b warns against the false POSITIVE: a token file that proves nothing. The mirror error is just as costly. Asserting
f10-coderbefore an RM-02 dispatch, the orchestrator probed/api/v1/userand gotlogin: None— which reads exactly like an unprovisioned seat. It is not. The seat token is least-privilege (write:issue,write:repository, noread:user), so/userreturns 403 by design.The real assertion is the DIFFERENTIAL, not any single endpoint: authenticated repo permissions return
push: truewhile unauthenticated returnspush: false— that gap proves the token caused the difference. A single-endpoint read can be true for anyone (false positive) or refused for a scope reason unrelated to the capability (false negative).RULE: a capability probe must assert the differential. Without it, this would have escalated a perfectly working seat to the coordinator as unprovisioned — burning a provisioning round trip and stalling the keystone on a phantom.
D-12 — a requested SAFETY flag was silently degraded, and I did not check
I created PR #1027 with pr-create.sh ... -d (draft) because it carries partial, unproven work.
tea authentication was stale, so the wrapper fell back to its raw-API path — which cannot set draft —
and emitted:
Warning: API fallback applies title/body/head/base only; labels/milestone/draft require authenticated tea setup.
The PR was created not-draft. I read the success output, saw the PR number, and moved on. I then
reported to the coordinator that the PR was "opened as draft". It was open, mergeable, and marked
ready for ~25 minutes, protected only by the words "DRAFT" and "do not merge" in its title and body —
i.e. by prose a human might read, not by the platform control I asked for. Detected only because a
watcher polled draft: and the value disagreed with my belief. Corrected by setting the WIP: title
prefix (Gitea's draft mechanism); draft: True verified after.
Three distinct failures, and the third is mine:
- Silent degradation of a safety flag. The fallback path dropped
--draftand still exited 0. A fallback that cannot honour a safety argument must fail, not proceed — degrading--labelsis a nuisance; degrading--draftpublishes unproven work as ready to merge. - The warning went to stderr and nothing consumed it. It was correct, specific, and ignored — a warning nobody acts on is indistinguishable from no warning.
- I did not verify the flag took effect. I checked that the PR existed, not that it had the property I required. This is the mission's own thesis turned on me: I trusted a success exit code over an observed state, on exactly the class of tool this mission exists to distrust.
Requirements. RM-02: a wrapper that cannot honour a safety-relevant argument must exit non-zero —
registered with a must-fail control asserting --draft on a degraded path fails rather than proceeds.
RM-24 (tri-state write outcomes): this is precisely written-unverified being treated as verified —
the PR write succeeded, the requested property was never confirmed, and no one looked.
D-11 — seat identity did not survive into git, and seat capability is invisible at dispatch
Two defects, one dispatch (RM-01 → f10-coder):
(a) Identity drift — P-WRAPPER-001, reproduced on our own delivery. The seat's commits are
authored mosaic-coder <[email protected]> — the generic fallback. You cannot tell from
git history which seat did this work. Recorded, not rewritten: the drift is the evidence.
Mechanism, corrected (Mos). My original framing here was wrong, and the error was in the brief before it was in the finding.
MOSAIC_GIT_IDENTITYresolves the token (which per-slot credential the wrappers act with). The commit author comes fromgit config user.name/user.email, which is a separate setting — it fell back to the generic value because nothing set it. Exporting the identity could never have fixed authorship. My worker brief instructed only the export, so the seat did exactly what it was told and the commits were still mis-attributed.The requirement is coherence: token and authorship must agree. A seat acting with
gitea-mosaicstack-f10-codermust also commit asf10-coder <[email protected]>. Either half alone is identity drift — one produces the right credential with the wrong author, the other the reverse. That coherence is P-WRAPPER-001, and it belongs in seat setup, not in prose instructions a seat may follow correctly and still end up wrong.
(b) Capability opacity. Nothing at dispatch time revealed that f10-coder had no credential for
the target provider. Per-slot tokens live at ~/.config/mosaic/secrets/gitea-tokens/; the seat holds
gitea-usc-f10-coder but not gitea-mosaicstack-f10-coder. This surfaced only when the seat failed
mid-task, after ~$9 and 69% of its context. The orchestrator (me) selected a seat without any way
to check it could act on the target repo — and there was no way to check.
get_gitea_token behaved correctly: it refused to fall through and borrow another slot's token,
failing loud precisely to protect gate-16 attribution. The tooling was right; the dispatch-time
information did not exist.
This is P-RECOVERY-001's "honest capability labeling" applied to seats rather than services. A seat should declare what it can actually do — which providers, which repos, which credentials — and that declaration must be checkable before dispatch, not discovered by failure after the budget is spent.
Requirements.
- RM-04 (activation coherence) gains the identity-binding half: seat setup must set both the
token identity and
git config user.name/user.email, coherently. Verified by an exit-asserting test that makes a commit and asserts its author — never assumed from an instruction in a brief. - RM-50 (roster ownership) gains per-seat capability declaration plus a pre-dispatch capability
check. Mos (who owns provisioning) confirms the check is mechanically trivial: capability is
token-file existence. Before dispatching seat
Xto providerY, test that~/.config/mosaic/secrets/gitea-tokens/gitea-<Y>-<X>.tokenexists; if absent, provision it or pick a provisioned seat. The token-file set is the authoritative capability registry. A one-second check would have replaced a mid-task failure that cost ~$9 and 69% of a seat's context.
D-10 — the queue guard's failure modes are exactly backwards
ci-queue-wait.sh — a required pre-push/pre-merge gate — was observed this session doing both of
these:
- Fails OPEN on an unknown result.
state=unknown ⇒ exit 0, five times, during real pushes and real merges. It also evaluatesbranch=mainrather than the branch being acted on. - Fails CLOSED on credential resolution. In a worker seat it aborted with
Gitea token not found, hard-blocking a legitimate push of completed, tested work. The worker correctly stopped (Constitution gate 8). The identical command run from that worker's own worktree in another shell succeeded, so the checkout and remote were fine — the difference was the worker's process environment.
A gate that waves through work it never checked, and blocks work that is ready, has its failure modes inverted. Availability failures (cannot reach the provider, cannot resolve a credential) should degrade to a loud, auditable inability to assert — never to a hard stop on delivery, and never to a silent pass. Correctness failures (unknown, malformed, terminal-failure) are what must block.
This is also the Pi-brick shape (P-RECOVERY-001): a gate whose own unavailability prevents the work needed to recover from it.
Requirement on RM-03, extending its existing two defects: the guard must distinguish
CANNOT_ASSERT (credential/transport/provider unavailable — loud, audited, does not silently pass and
does not permanently block) from ASSERTED_NOT_READY (a real non-green CI state — blocks). Both are
registered R-002 cases with must-fail controls; neither may exit 0 silently.
D-9 — the comms path shell-interprets message bodies (injection-shaped, found by accident)
Sending a status message with agent-send.sh -m "...backticks..." caused bash to execute the
backticked text as command substitution. The recipient received a mangled body plus a
No such file or directory error; the intended sentence never arrived. The message was reported as
delivered.
This is the same class as the already-noted pr-create.sh backtick-quoting bug (M2 scratchpad):
two tools in the comms path treat a message body as shell input. A body that can execute on the
sender is a correctness bug before it is ever a security one — and note the failure mode: the
send reported success while silently transmitting something other than what was written. Silent
corruption with a success receipt is precisely the pattern this mission exists to eliminate.
Requirement on RM-40 / RM-42 (comms/v1), hardened by Mos. The envelope must carry its payload verbatim and must not be subject to shell interpretation at any hop — sender, transport, or adapter. Concretely: file/stdin transport, never argv interpolation.
Standing interim rule, effective now (Mos). Until the envelope lands, use agent-send.sh -f <file> for any message body containing special characters — never -m. Passing a file sidesteps
argv interpolation entirely. This rule is mandatory in every worker brief this mission issues,
alongside the D-8 "if a check is unrunnable, say so" clause. Round-trip fidelity (send a body containing backticks, $(…), quotes, and newlines; assert
byte-identical receipt) is a required registered test case under RM-02, including a must-fail control
proving the assertion can detect corruption.
D-8 — a PRE-REGISTERED acceptance check that was not runnable as written
On PR #1025 the author (me) pre-registered AC2 with the fixture snippet mkdir -p apps/*/venv/lib.
In bash, when no venv exists the glob is unmatched and passes through literally, creating a
directory named apps/*/venv/lib rather than one per workspace. The check as written did not test
what it claimed to test.
rev-974 ran it exactly as written, observed the wrong behaviour, then re-ran the intended
assertion at an explicit path — and said so in the review rather than silently substituting a
working fixture and reporting PASS.
Two things this establishes:
- The instruction "do not adjust a check to fit the diff; if it is unrunnable, say so explicitly" worked. A silent substitution here would have produced a green AC2 that proved nothing, on the exact task whose subject is gates that appear to work. The disclosure is what made the PASS meaningful.
- Pre-registration does not confer correctness. A pre-registered check is protected from being retrofitted to the implementation; it is not protected from being wrong when written. This is a small instance of the mission's own class — an unverified gate — occurring inside the mechanism built to catch unverified gates.
Requirement on RM-02 (non-negotiable, sharpened by Mos). The registry must self-verify that every registered case demonstrably runs and demonstrably fails on a known-bad input. Presence in the registry is not evidence. A check is not trusted until it has been shown to fail. This is mutation testing / negative control applied at the registry level — meaning the conformance harness must itself be conformance-tested. A registered case that cannot fail, or cannot run, is exactly as inert as an unregistered one, and the registry check must detect that itself rather than assume it.
Requirement on RM-55. The same recursion applies to the harness: it must be observed red before its green is worth anything (OPUS R-063 AC1 already states this; D-8 is the empirical case for it).
Second, equally load-bearing lesson — reviewer disclosure is what makes a review trustworthy.
rev-974 could have silently swapped in a working fixture and reported AC2 PASS. Nothing in the
process would have caught it, and the resulting green would have certified nothing — on the very task
whose subject is gates that only appear to work. The brief's instruction — "do not adjust a check to
fit the diff; if it is genuinely unrunnable as specified, say so explicitly and explain why rather
than silently substituting your own" — is therefore not boilerplate. It is the clause that makes a
PASS mean something, and it must appear in every reviewer brief this mission issues.
D-7 — shared-tmpfs contention → cascading ENOSPC (live incident, 2026-07-31)
The shared 30 G /tmp hit 100% ENOSPC mid-session. It broke tool calls in two different seats
(mine and Mos's) — a single full disk degrades every agent on the host at once. Recurring: prior
incidents 2026-06-18 and 2026-07-17.
Attribution matters, because the wrong owner cleans the wrong thing. Measured:
| path | size | last modified | owner |
|---|---|---|---|
…/-src-mosaic-stack/6d2faee6… (this session) |
88 K | live | mos-remediation |
…/-src-mosaic-stack/c743185d… |
3.6 G | 2026-07-22 (9 days dead) | abandoned session, same project path |
…/claude-1001/pnpm-store |
1.6 G | 2026-07-23 (8 days dead) | abandoned; the live store is correctly on $HOME |
So ~5.2 G — the bulk of the pressure — is dead session scratch that nothing will ever read again. This is not a quota problem; it is P-FLEET-001's stale-session GC, applied to disk instead of tmux sessions. The same missing capability (nothing owns reaping dead ephemeral state) produces both the orphaned-session failure and this one. Reaping dead-session scratch belongs in RM-50 alongside stale tmux-session GC.
Added to RM-01 as acceptance criteria: heavy build artifacts (node_modules, package stores, build
output) must land on the main disk in the worktree, never on the shared 30 G /tmp.
Resolution, and the part that is actually the finding. Mos verified the attribution independently
(mtimes, no process or lsof holding either path, no live session maps) and reaped both as lead
coordinator: /tmp went to 79%, 6.0 G free. But note how it was resolved — a human-authority seat
did it by hand, because the authority exists and the reaper does not. That gap is the finding, not
the disk usage.
Two doctrine points fall out, both binding on RM-50:
- The fix is not "agents should tidy up." Asking each seat to clean its own scratch is
instructions are not enforcement(D-4) wearing a different hat. A deterministic reaper must own it — same conclusion the north star reaches for every other class in this mission. - Refusing to unilaterally delete another session's scratch was correct, and the resolution is not "be braver about deleting." An agent guessing that someone else's state is garbage is exactly the unreviewed destructive act the Constitution forbids. The resolution is that ownership and liveness become mechanically decidable, so reaping is a determination rather than a judgement call.
Reaper requirements for RM-50: liveness determined mechanically (process/lsof/session-map, not
mtime alone); an age threshold; a dry-run that reports what it would reap and why; and an audit event
per reap. Never a heuristic sweep — that would reintroduce the P-WORKFLOW-001 auto-sync failure in a
more destructive form.
The self-erasure is the important part. An inert gate that is masked by unrelated downstream
commits produces no lasting artifact, which is precisely why this class survives for months. Detection
cannot rely on "is main currently red" — it must be per-merge-commit.
This matters more than the one-line fix:
- It is the P-QUEUE-001 / P-CONFORMANCE-001 class ("gate-6 was inert fleet-wide"), reproduced in the repository this mission is remediating, discovered incidentally.
- It independently validates OPUS premise A1 ("every gate is inert until proven otherwise") with live evidence rather than argument — which is why RM-02 is adopted as the keystone (§2, X2).
- The file fix rides in its own hygiene PR. The inert gate itself is NOT quiet-patched. Per Mos: it stays a first-class backlog item, because patching the symptom would destroy the signal.
Binding requirement on RM-02 and RM-55: the gate registry and the conformance harness must assert
"every merged commit passed every required gate" — evaluated per merge commit, against that
commit's own tree, not against current main. As the table above proves, a "is main green today"
check would have reported all-clear. A merged-commit-that-fails-a-required-gate is the exact detection
signal, and it must be a registered must-fail case. A gate that cannot prove it blocked something has
not been shown to work.
2. Where they genuinely disagree (not averaged — adjudicated)
| # | Axis | OPUS | SOL | My ruling |
|---|---|---|---|---|
| X1 | Total cost | 38 tasks, ~5.3M tok | 25 tasks, ~294K tok | ~18× apart. Not reconcilable by splitting. They measure different things: SOL explicitly excludes orchestration/review/iteration overhead and assumes one remediation pass; OPUS prices the full loop. Adopt SOL's scope with OPUS's rigor, and treat SOL's G1 as a hard budget checkpoint (§4). Re-estimate empirically after the first three merged PRs rather than trusting either number. |
| X2 | Gate registry (OPUS R-002) | Keystone; blocks all P2 | Absent; only a queue-guard fix | ADOPT OPUS. Empirically validated in this very session: I found pnpm format:check red on main via merged PR #868 — a required gate that did not block. OPUS's premise A1 ("every gate is inert until proven otherwise") is not theoretical; it reproduced today, unprompted. Scope it tighter than 120K. |
| X3 | Drizzle PG first-install defect (R-010) | Hidden blocker; everything downstream depends on it | Not mentioned | ADOPT OPUS. packages/db/src/migrate.ts:30-38 carries a TODO admitting postgres-tier first-install fails today. The spine has only ever been proven on PGlite. Every later migration silently depends on this. SOL missed it. |
| X4 | Rollback artifact for the hard cutover | D3: hard cutover needs a rehearsed rollback snapshot | SOL-07: import-only, explicitly no dual-write | Both obey "no flat-file interim." OPUS wants a one-directional snapshot nothing reads as authority. I read that as compatible with the directive, but it is Jason's call → DECISION-2 (§5). |
| X5 | Where the queue guard sits | P0, independent of spine | SOL-02, also early | Agree it is P0. But ownership collides with parked PR #1023 → DECISION-3 (§5). |
| X6 | Report-only rollout | D5: only with a hard expiry, else withdraw | not raised | ADOPT OPUS. A report-only gate is by definition inert; expiry is the mechanism that stops it becoming the new fail-open. |
| X7 | Availability trade (FC-7/FC-11) | D8: "no DB ⇒ fleet stops" must be pre-committed in writing | not raised | Genuine availability regression, correctly identified. Needs Jason → folded into DECISION-2. |
3. Reconciled DAG
Phases run in order; ⛔ marks a hard barrier. src shows lineage (O=opus, S=sol, O+S=both).
Estimates are given as a range (SOL low / OPUS high) rather than a fabricated midpoint — the
spread is itself information, and X1 says we calibrate on real merged PRs.
P0 — Make gates provable, and stop the fleet re-bricking
⛔ No gate-introducing task in any later phase may merge before RM-02.
| id | task | src | depends_on | est (S/O) | tier |
|---|---|---|---|---|---|
| RM-01 | Reproducible non-root checkout; gate fails on code, not env; heavy artifacts OFF shared /tmp (banks D-1/D-2/D-5/D-7) |
O+S+live | — | 6K / 60K | codex |
| RM-02 | Gate registry + negative-control CI check (anti-inert-gate harness) ★keystone | O | RM-01 | — / 120K | opus |
| RM-03 ⏸HOLD | Queue-guard: two defects — (a) unknown/no-status/malformed ⇒ ≠0, (b) guard evaluates branch=main instead of the branch being pushed |
O+S+live | RM-02 | 8K / 100K | sonnet |
| RM-04 | Activation/version coherence; block launch on skew, fail SAFE; honest doctor labels |
O+S | RM-01 | (in S-01) / 140K | sonnet |
| RM-05 | Break-glass replaces the three silent MOSAIC BYPASS fail-opens |
O | RM-04, RM-02 | — / 120K | opus |
⚠ RM-05 must not merge before RM-04. The bypasses exist because the lease-broker daemon was never deployed on this host — removing the fail-open before deployment coherence is real re-creates the 2026-07-22 bricking incident. Hard edge, from OPUS.
P1 — Durable spine (PG)
⛔ No migration may merge before RM-10.
| id | task | src | depends_on | est (S/O) | tier |
|---|---|---|---|---|---|
| RM-10 | Fix the Drizzle postgres-tier first-install defect ★hidden blocker | O | RM-01 | — / 90K | sonnet |
| RM-11 | Orchestration spine schema (tasks, attempts, gate_results, hash-chained ledger, typed claims) | O+S | RM-10 | 12K / 160K | opus |
| RM-12 | Spine client, fail-closed connection (no silent PGlite in prod) | O | RM-11 | — / 80K | sonnet |
| RM-13 | Atomic claims/transitions + transactional outbox + reconciliation sweeper | O+S | RM-12 | 12K / 140K | opus |
P2 — The single choke point
⛔ RM-25 (no-second-path) lands in the same milestone as RM-20, or the choke point is optional.
| id | task | src | depends_on | est (S/O) | tier |
|---|---|---|---|---|---|
| RM-20 | Canonical MACP contract completion (Task/Result/Event/Claim/tri-state outcome) | S | — | 8K / (in R-020) | codex |
| RM-21 | Production TaskExecutor backed by @mosaicstack/macp ★keystone |
O+S | RM-12, RM-02, RM-20 | 16K / 220K | opus |
| RM-22 | Gate-runner hardening: fail_on, timeouts, empty gate set = failure |
O | RM-21 | — / 120K | sonnet |
| RM-23 | Hash-chained MACPEvent ledger in PG + lifecycle EventType extension | O+S | RM-21, RM-11 | — / 160K | opus |
| RM-24 | Seat identity from MOSAIC_AGENT_NAME + mandatory tri-state write outcomes |
O+S | RM-21 | (in S-03) / 150K | opus |
| RM-25 | No-second-path gate: terminal status writable only by the executor | O | RM-21, RM-23 | — / 140K | opus |
| RM-26 | packages/coord submits through the executor (retire direct spawn) |
O+S | RM-21 | 16K / 140K | sonnet |
| RM-27 | mosaic yolo/claude/codex/pi launch path records typed Task + events |
O | RM-21, RM-23 | — / 160K | sonnet |
| RM-28 | Delete the Forge stub executor (empty-gate-list "success"); Forge submits through the real one | O+S | RM-21 | 10K / 90K | codex |
| RM-29 | One-shot flat-file import + cutover readiness audit (dry-run, idempotent, no dual-write) | S | RM-13 | 8K / (in R-062) | codex |
★ G1 — FIRST DOGFOOD. Stop here and prove it. One live fleet task travels PG claim → TaskExecutor → worker → gates → terminal PG result/event, with no flat-file state. Adopted from SOL wholesale. If G1 cannot carry a real task, do not build Redis, rotation, comms, or conformance — remediate instead. This is the budget escape hatch (§4).
P3 — Rotation lifecycle (finish the Mission Control Plane)
| id | task | src | depends_on | est (S/O) | tier |
|---|---|---|---|---|---|
| RM-30 | Typed state claims (source/confidence/TTL) with HMAC integrity, fail-closed | O+S | RM-11, RM-21 | (in S-03) / 170K | opus |
| RM-31 | Contract-hash binding; stale generation loses mutation authority mechanically | O+S | RM-21, RM-30 | 12K / 180K | opus |
| RM-32 | Durable compaction/token sensor (per-runtime thresholds, PreCompact event) | O | RM-23, RM-31 | — / 130K | sonnet |
| RM-33 | Typed checkpoint writer (structured claims, never transcript) + digest | O+S | RM-30, RM-32 | 12K / 150K | opus |
| RM-34 | Rotation daemon: watch → checkpoint → revoke → kill → relaunch → rehydrate | O+S | RM-33, RM-26 | 16K / 240K | opus |
| RM-35 | Rehydration attestation gate: refuse to act on an incomplete claim set | O | RM-33 | — / 130K | opus |
| RM-36 | Broker-independent recovery; remove silent bypass; honest capability labels | S | RM-34 | 12K / (in R-004) | sonnet |
| RM-37 | Delete /compact and continue from the persistent-seat path (substitution, not removal) |
O+S | RM-34, RM-44 | (in S-16) / 60K | codex |
P4 — Comms service
⛔ RM-50 (one roster-owned socket per host) precedes identity-addressed delivery.
| id | task | src | depends_on | est (S/O) | tier |
|---|---|---|---|---|---|
| RM-40 | comms/v1 envelope + protocol-version negotiation, LOUD reject |
O+S | RM-11, RM-31 | 8K / 140K | opus |
| RM-41 | Comms service: PG state machine PENDING→RECEIVED→CONSUMED→DEAD-LETTER | O+S | RM-40, RM-13 | 16K / 200K | opus |
| RM-42 | tmux transport as a dumb adapter; durable retry before cursor advance | O+S | RM-41, RM-50 | (in S-19) / 160K | sonnet |
| RM-43 | Per-class coalescing + supersede (the stale-consumed-as-live fix) | O+S | RM-41 | 12K / 130K | sonnet |
| RM-44 | Redis Streams hot delivery + provenance guard (Redis is never authority) | O+S | RM-41, RM-13 | 12K / 170K | opus |
| RM-45 | Retire direct tmux sends; only the service may write a pane | O+S | RM-42, RM-43 | (in S-20) / 100K | codex |
P5 — Retirements, hygiene, conformance
| id | task | src | depends_on | est (S/O) | tier |
|---|---|---|---|---|---|
| RM-50 | One roster-owned socket/host; quarantine unmanaged; deterministic reaper for stale sessions AND dead-session disk scratch (D-7). Acceptance MUST be proved against the UNMANAGED execution fleet, not the roster-managed canaries (D-41) | O+S+live | RM-04, RM-62 | 14K / 150K | sonnet |
| RM-51 | Auto-sync allowlist (never auto-stage unknown paths) + worktree/lease isolation | O+S | RM-02 | 8K / 110K | sonnet |
| RM-52 | Retire the Python controller + duplicate MACP islands (3 → 1) | O+S | RM-26, RM-27, RM-25, RM-28 | 14K / 110K | codex |
| RM-53 | Flat-file orchestration → DB hard cutover, with rehearsed rollback artifact | O+S | RM-27, RM-30, RM-34, RM-29 | (in S-10) / 200K | opus |
| RM-54 | Fleet-wide inert-gate audit against the RM-02 registry | O | RM-02 | — / 120K | sonnet |
| RM-55 | Conformance harness: fault-inject the live failure classes on real artifacts | O+S | RM-35, RM-41, RM-53 | 18K / 260K | opus |
| RM-56 | Retirement proof: CI asserts all three retirements are complete and stay complete | O | RM-52, RM-45, RM-53 | — / 90K | codex |
| RM-57 | Operator cutover docs + activation proof; map all 15 decisions to evidence | S | RM-04, RM-36, RM-45, RM-55 | 6K / — | codex |
| RM-61 | CI-contract exemption for the #1000 teardown artifact — signature-scoped, negative-control-proven, bounded, retiring with #1000 (ruled B, Mos 2026-08-01) | mos-remediation | — (unassigned; no free write-capable seat) | 15K | sonnet |
| RM-60 | External pre-execution trust boundary for CI (option B — the correct primitive, not the cautious one) — protected default-branch pipeline config or an immutable trusted launcher that enters the sandbox before any PR-controlled executable/config is evaluated; unblocks isolated per-commit replay (RM02-REQ-10) | mos-remediation (D-25) | infra/provider authority (Mos + Jason) | 25K | opus |
| RM-59 | Close the D-19 residual risk — generated-state verification anchored outside the worktree's authority (executor/spine-side attestation), retiring the same-UID self-authentication gap | mos-remediation (D-19) | RM-12, RM-21, RM-25 | 20K | opus |
| RM-58 | Mechanical pre-dispatch context reset — the orchestrator resets a seat out-of-band and verifies it, rather than asking the agent to reset itself | mos-remediation (D-4) | RM-31, RM-50, RM-62 | 8K | sonnet |
| RM-62 | ★ PREREQUISITE — bring the EXECUTION fleet under roster/systemd management. Every working seat (coder-mos1, rev-974, f10-coder, merge-gate, pm-scout-*, rev-3107b, ultron-3107, and mos-remediation) is inactive/disabled + UNMANAGED — there is no lifecycle surface to enforce anything against. BLOCKS RM-50, RM-58, P-LIFECYCLE-001; each is unsatisfiable against its real population until this lands. One infrastructure cluster with D-37's shared-config fix; Mos owns both, sequenced at a seam, never mid-lane |
mos-remediation (D-41) | infra authority (Mos) | TBD | opus |
| RM-63 | ★ ARCHITECTURE (Jason, 2026-08-01) — SYSTEMD-MANAGED AGENT SERVICES with DURABLE NAMES + LANES. This is the mechanization of D-41 / RM-50 / RM-58 / P-LIFECYCLE-001, not a new island. A standing set of durably-named services (coders, FE, BE, planners, code-review, sec-review, Ultron/merge-gate); fleet restart / systemctl restart triggers a CONTINUATION — the agent resumes its LANE. Combined with DB-centric task assignment on the PG spine, durable names + lanes give autonomous task pickup on continue. Provides the enforcement surface D-41 proved missing: a managed service can be watched (token budget → D-49's mechanical ceiling monitor), restarted (rotation, retiring the manual respawns), continued (rehydrate from the DB). Rotation, reaping and quarantine become MECHANICAL. Sits with the coordinator-daemon + DB-spine deliverables |
Jason (architecture) | RM-62, RM-12/RM-13 (spine) | TBD | opus |
Critical path: RM-01 → RM-02 → RM-10 → RM-11 → RM-12 → RM-21 → RM-23 → RM-31 → RM-33 → RM-34 → RM-53 → RM-55.
4. Execution discipline
- Every row is one PR. Author ≠ reviewer;
rev-974is the mosaicstack reviewer identity. - Pre-registered, diff-blind acceptance checks are committed BEFORE the reviewer reads the diff.
Both decomps wrote their ACs in runnable
⇒0/⇒≠0form specifically to make this possible. - Every gate-introducing task carries at least one registered must-fail negative control. This is RM-02's whole purpose; a gate with no proven failure path manufactures evidence.
- Cost tiers: codex for mechanical/unambiguous, sonnet for normal feature work, opus reserved for security/integrity/cross-cutting-invariant tasks. SOL priced 0 opus tokens; OPUS priced 14 opus tasks. I am keeping opus only where the failure is integrity, not merely complexity.
- G1 is the budget checkpoint. If the first-dogfood slice overruns SOL's estimate by >3×, stop and re-plan rather than spending the remainder. X1 says neither estimate is trustworthy until calibrated.
- Defer list adopted from SOL (10 items): mission dashboard/TUI, PRD-to-board auto-decomposition, heuristic churn scoring, Discord/Slack/Telegram adapters, public MCP comms surface, protocol-v2 negotiation, multi-region PG/Redis, event analytics UI.
5. Decisions — all three ruled by Mos, 2026-07-31
DECISION-1 — the wire-in point. ✅ RULED: accept the planners (Mos, 2026-07-31).
The charter's mosaic_orchestrator.py::run_single_task target is the disabled Python controller
this mission retires; wiring the new choke point into the rail we are deleting is wrong.
Corrected target (authoritative): a new production Node
TaskExecutorsitting on the live dispatch path —packages/mosaiclaunch +packages/coord— which Coord, Forge, and live dispatch all submit through. This is the MACP scout's full recommendation ("replace the block with a Node executor and make Coord/Forge submit through it"), not a resurrection of the Python controller. RM-52 is therefore a deletion task, and Build 1's acceptance is measured on a livemosaic yoloinvocation.
Mos ruled this resolvable from the already-accepted retire-the-Python-rail decision — his authority, not a Jason escalation. RM-21/RM-26/RM-27/RM-52 all take the corrected target.
DECISION-2 — rollback artifact + availability trade. ⏸ JASON-PENDING — NOT BLOCKING. The DB build is phases away, so this is queued for Jason's next session rather than escalated now. Binding requirement in the meantime (Mos, from P-RECOVERY-001): the DB spine must NOT be a single-point hard-stop. Design for a broker-independent / degraded mode plus a rollback artifact. Jason finalises only the specific availability target. This reverses my earlier reading of OPUS D8 ("the fallback is: the fleet stops") — that answer is not pre-committed; a degraded mode is now a design requirement on RM-12, RM-13, RM-23, RM-36 and RM-53.
DECISION-3 — RM-03 vs. parked PR #1023. ✅ RULED: HOLD RM-03 (Mos, 2026-07-31). Do not open a third gate-6 lane — that is the postmortem's own anti-pattern performed by the remediation. PR #1023 sits in Jason's parked delivery stack; its disposition (close, or supersede by RM-03) is Jason's at his next session.
- PR #1023 →
SUPERSEDED-PENDING-JASON. RM-03 staysHOLD; when Jason rules, RM-03 proceeds as the single correct lane. - RM-02 and RM-55 proceed independently and are NOT held. The per-merge-commit gate-assertion requirement is the conformance capability, not the gate-6 fix itself — different scope, no ownership collision.
6. Status
| phase | state |
|---|---|
| Decomposition | DONE — both planners delivered independently |
| Reconciliation | DONE — this document |
| Blocking decisions | RULED — all 3 closed by Mos 2026-07-31 (§5); D-2's availability target is Jason-pending but non-blocking |
| Dispatch | RM-01 IN FLIGHT — f10-coder (codex), worktree-isolated, AC1–AC8 pre-registered |
| Review | PR #1025 with rev-974; ACs pre-registered 22:12:26Z before diff exposure |