Files
stack/docs/remediation/TASKS.md
T
mos-dt-0andClaude Opus 5 75dfe2fa75 docs(remediation): bank D-11 — identity drift + seat capability opacity at dispatch
RM-01's seat was briefed to export MOSAIC_GIT_IDENTITY; its commits are authored as
the generic mosaic-coder fallback, so git history cannot say which seat did the work.
P-WRAPPER-001 reproduced on our own delivery.

Separately, nothing at dispatch time revealed the seat lacked a credential for the
target provider — discovered only when it failed mid-task after ~$9 and 69% context.
get_gitea_token behaved correctly by refusing to borrow another slot's token; the
dispatch-time information simply did not exist.

RM-50 gains per-seat capability declaration + pre-dispatch check; RM-04 gains
identity-binding verified by an exit-asserting test rather than assumed from an
export in a brief.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:56:34 -05:00

40 KiB
Raw Blame History

Remediation Backlog — Reconciled Execution Plan

Owner: mos-remediation (sole writer). Workers read; they never modify this file. Sources: DECOMP-OPUS.md (robustness, 38 tasks / 8 dissents) and DECOMP-SOL.md (pragmatic, 25 tasks / 10 defers / 7 dissents), produced independently — neither planner read the other. Charter: MISSION.md. Status: EXECUTING — all three blocking decisions RULED by Mos on 2026-07-31 (§5). RM-01 is dispatched. RM-03 is held pending Jason's disposition of PR #1023; nothing else is blocked.

Provenance of the inputs (both clean). planner-opus ran in a fresh session throughout. planner-sol initially began work at 64.3% dirty context despite a brief instructing it to reset; that run was interrupted and discarded before it produced any output, the seat was reset out-of-band to 0.0%, and the brief was re-dispatched. DECOMP-SOL.md is the product of the clean run only (it peaked at ~26% context). Both decompositions are therefore clean-context artifacts and are weighted equally here. The discarded dirty run is banked as dogfood seed D-4 and as task RM-58 — the failure it demonstrates is that asking an agent to reset is not enforcement.


1. What the two planners agreed on without collusion

Independent convergence is the strongest signal available here, because neither planner could see the other's file. Where both arrived at the same conclusion from opposite biases, I treat it as settled.

# Convergent finding OPUS SOL
C1 The charter's wire-in point is wrong. Do NOT wire the choke point into mosaic_orchestrator.py::run_single_task. D2 (headline) Dissent 1
C2 P0 hygiene/gate work must precede the spine, not follow it. D1, phase P0 G0, SOL-01/02
C3 No new deployable microservice; the executor is a library + the coord daemon. implicit throughout Dissent 3
C4 Comms adapters beyond tmux are out of scope for this mission. D7 DEFER 4/5/6
C5 The "100 rotations lossless" bar is a late conformance gate, not an early tax. D4 Dissent 7
C6 Redis is a derived hot path, never an authority; PG commits first. R-013, R-054 SOL-11, SOL-21
C7 Reuse packages/coord; do NOT revive the untracked apps/coordinator residue. R-042 SOL-15 AC5

C1 is the single most consequential output of this exercise. The charter (MISSION.md) and my kickoff instruction both name mosaic_orchestrator.py::run_single_task:126-276 as the integration point. Both planners independently rejected it on the same evidence: that controller is "enabled": false (.mosaic/orchestrator/config.json:2) and references a dispatcher path (tools/macp/dispatcher/pi_runner.ts) that does not exist in this checkout. Wiring the new choke point into a disabled rail produces a stranded executor — the identical built-but-unwired disease, one layer up, that would look "done" in a PR. The live paths are packages/mosaic/src/commands/launch.ts and packages/coord/src/runner.ts. This contradicted the charter and was escalated as DECISION-1 — now RULED in the planners' favour by Mos (§5). The corrected target is a new production Node TaskExecutor on the live dispatch path (packages/mosaic launch + packages/coord) that Coord/Forge/live dispatch submit through; the Python rail is deleted, not ported.


1a. ★ KEYSTONE DOGFOOD CASE — an inert gate that erased its own evidence

A merged commit shipped past pnpm format:check — and then the evidence quietly erased itself.

Verified chain (blob-level, under the repo's own prettier config, at the file's real path):

commit state of packages/mosaic/framework/tools/orchestrator/README.md
b79336a8merged PR #868 blob 3ee7f104FAILS pnpm format:check
48fd1df2 — merged PR #872 (unrelated: ci-queue-wait 404 handling) blob 3d3bb132 — passes; incidentally reformatted by that PR's lint-staged
current origin/main (06e0d403) passes — the gate now looks green

So: PR #868 merged a file that fails a required gate ⇒ the CI format gate did not block it. The gate was inert for that merge. Then an unrelated later PR's pre-commit hook reformatted the file as a side effect, so main went green again without anyone ever learning the gate had failed to fire.

Correction on record: my first report to Mos said "format:check is RED on main now." That was true of the main my checkout was pinned to (b79336a8) and is no longer true of current main, which advanced mid-session. The inert-gate finding itself is unchanged and verified; only its present-tense framing was wrong. The hygiene PR therefore carries the .prettierignore fix only — the README needs no fix today.

Third live instance, same class — the queue guard, hit by this orchestrator

Running the mandated pre-push guard during TASK-0:

$ ~/.config/mosaic/tools/git/ci-queue-wait.sh --purpose push
[ci-queue-wait] platform=gitea purpose=push branch=main sha=06e0d403…
[ci-queue-wait] state=unknown purpose=push branch=main
$ echo $?   →  0

Two distinct defects in one tool, both feeding RM-03:

  1. Wrong exitstate=unknownexit 0. The defect at ci-queue-wait.sh:282-288 that OPUS documented and that PR #1023 is parked on. A required gate returned PASS on an indeterminate result.
  2. Wrong branch — it evaluated branch=main, not the branch actually being pushed. Even a correctly-exiting guard would have been answering the wrong question.

Standing doctrine (Mos): until RM-03 lands, a green from this guard carries zero information and must not be cited as merge evidence. Rely on reviewer clearance + real CI.

Three independent live instances in a single session — format gate, agent context reset, queue guard — is the class confirmed, not anecdote.

D-11 — seat identity did not survive into git, and seat capability is invisible at dispatch

Two defects, one dispatch (RM-01 → f10-coder):

(a) Identity drift — P-WRAPPER-001, reproduced on our own delivery. The brief instructed the seat to export MOSAIC_GIT_IDENTITY=f10-coder. Its commits are authored mosaic-coder <[email protected]> — the generic fallback. You cannot tell from git history which seat did this work. Recorded, not rewritten: the drift is the evidence.

(b) Capability opacity. Nothing at dispatch time revealed that f10-coder had no credential for the target provider. Per-slot tokens live at ~/.config/mosaic/secrets/gitea-tokens/; the seat holds gitea-usc-f10-coder but not gitea-mosaicstack-f10-coder. This surfaced only when the seat failed mid-task, after ~$9 and 69% of its context. The orchestrator (me) selected a seat without any way to check it could act on the target repo — and there was no way to check.

get_gitea_token behaved correctly: it refused to fall through and borrow another slot's token, failing loud precisely to protect gate-16 attribution. The tooling was right; the dispatch-time information did not exist.

This is P-RECOVERY-001's "honest capability labeling" applied to seats rather than services. A seat should declare what it can actually do — which providers, which repos, which credentials — and that declaration must be checkable before dispatch, not discovered by failure after the budget is spent.

Requirements: RM-50 (roster ownership) gains per-seat capability declaration and a pre-dispatch capability check; RM-04 (activation coherence) gains the identity-binding half — a seat's declared identity must provably reach its commits, verified by an exit-asserting test, not assumed from an export in a brief.

D-10 — the queue guard's failure modes are exactly backwards

ci-queue-wait.sh — a required pre-push/pre-merge gate — was observed this session doing both of these:

  • Fails OPEN on an unknown result. state=unknown ⇒ exit 0, five times, during real pushes and real merges. It also evaluates branch=main rather than the branch being acted on.
  • Fails CLOSED on credential resolution. In a worker seat it aborted with Gitea token not found, hard-blocking a legitimate push of completed, tested work. The worker correctly stopped (Constitution gate 8). The identical command run from that worker's own worktree in another shell succeeded, so the checkout and remote were fine — the difference was the worker's process environment.

A gate that waves through work it never checked, and blocks work that is ready, has its failure modes inverted. Availability failures (cannot reach the provider, cannot resolve a credential) should degrade to a loud, auditable inability to assert — never to a hard stop on delivery, and never to a silent pass. Correctness failures (unknown, malformed, terminal-failure) are what must block.

This is also the Pi-brick shape (P-RECOVERY-001): a gate whose own unavailability prevents the work needed to recover from it.

Requirement on RM-03, extending its existing two defects: the guard must distinguish CANNOT_ASSERT (credential/transport/provider unavailable — loud, audited, does not silently pass and does not permanently block) from ASSERTED_NOT_READY (a real non-green CI state — blocks). Both are registered R-002 cases with must-fail controls; neither may exit 0 silently.

D-9 — the comms path shell-interprets message bodies (injection-shaped, found by accident)

Sending a status message with agent-send.sh -m "...backticks..." caused bash to execute the backticked text as command substitution. The recipient received a mangled body plus a No such file or directory error; the intended sentence never arrived. The message was reported as delivered.

This is the same class as the already-noted pr-create.sh backtick-quoting bug (M2 scratchpad): two tools in the comms path treat a message body as shell input. A body that can execute on the sender is a correctness bug before it is ever a security one — and note the failure mode: the send reported success while silently transmitting something other than what was written. Silent corruption with a success receipt is precisely the pattern this mission exists to eliminate.

Requirement on RM-40 / RM-42 (comms/v1), hardened by Mos. The envelope must carry its payload verbatim and must not be subject to shell interpretation at any hop — sender, transport, or adapter. Concretely: file/stdin transport, never argv interpolation.

Standing interim rule, effective now (Mos). Until the envelope lands, use agent-send.sh -f <file> for any message body containing special characters — never -m. Passing a file sidesteps argv interpolation entirely. This rule is mandatory in every worker brief this mission issues, alongside the D-8 "if a check is unrunnable, say so" clause. Round-trip fidelity (send a body containing backticks, $(…), quotes, and newlines; assert byte-identical receipt) is a required registered test case under RM-02, including a must-fail control proving the assertion can detect corruption.

D-8 — a PRE-REGISTERED acceptance check that was not runnable as written

On PR #1025 the author (me) pre-registered AC2 with the fixture snippet mkdir -p apps/*/venv/lib. In bash, when no venv exists the glob is unmatched and passes through literally, creating a directory named apps/*/venv/lib rather than one per workspace. The check as written did not test what it claimed to test.

rev-974 ran it exactly as written, observed the wrong behaviour, then re-ran the intended assertion at an explicit path — and said so in the review rather than silently substituting a working fixture and reporting PASS.

Two things this establishes:

  1. The instruction "do not adjust a check to fit the diff; if it is unrunnable, say so explicitly" worked. A silent substitution here would have produced a green AC2 that proved nothing, on the exact task whose subject is gates that appear to work. The disclosure is what made the PASS meaningful.
  2. Pre-registration does not confer correctness. A pre-registered check is protected from being retrofitted to the implementation; it is not protected from being wrong when written. This is a small instance of the mission's own class — an unverified gate — occurring inside the mechanism built to catch unverified gates.

Requirement on RM-02 (non-negotiable, sharpened by Mos). The registry must self-verify that every registered case demonstrably runs and demonstrably fails on a known-bad input. Presence in the registry is not evidence. A check is not trusted until it has been shown to fail. This is mutation testing / negative control applied at the registry level — meaning the conformance harness must itself be conformance-tested. A registered case that cannot fail, or cannot run, is exactly as inert as an unregistered one, and the registry check must detect that itself rather than assume it.

Requirement on RM-55. The same recursion applies to the harness: it must be observed red before its green is worth anything (OPUS R-063 AC1 already states this; D-8 is the empirical case for it).

Second, equally load-bearing lesson — reviewer disclosure is what makes a review trustworthy. rev-974 could have silently swapped in a working fixture and reported AC2 PASS. Nothing in the process would have caught it, and the resulting green would have certified nothing — on the very task whose subject is gates that only appear to work. The brief's instruction — "do not adjust a check to fit the diff; if it is genuinely unrunnable as specified, say so explicitly and explain why rather than silently substituting your own" — is therefore not boilerplate. It is the clause that makes a PASS mean something, and it must appear in every reviewer brief this mission issues.

D-7 — shared-tmpfs contention → cascading ENOSPC (live incident, 2026-07-31)

The shared 30 G /tmp hit 100% ENOSPC mid-session. It broke tool calls in two different seats (mine and Mos's) — a single full disk degrades every agent on the host at once. Recurring: prior incidents 2026-06-18 and 2026-07-17.

Attribution matters, because the wrong owner cleans the wrong thing. Measured:

path size last modified owner
…/-src-mosaic-stack/6d2faee6… (this session) 88 K live mos-remediation
…/-src-mosaic-stack/c743185d… 3.6 G 2026-07-22 (9 days dead) abandoned session, same project path
…/claude-1001/pnpm-store 1.6 G 2026-07-23 (8 days dead) abandoned; the live store is correctly on $HOME

So ~5.2 G — the bulk of the pressure — is dead session scratch that nothing will ever read again. This is not a quota problem; it is P-FLEET-001's stale-session GC, applied to disk instead of tmux sessions. The same missing capability (nothing owns reaping dead ephemeral state) produces both the orphaned-session failure and this one. Reaping dead-session scratch belongs in RM-50 alongside stale tmux-session GC.

Added to RM-01 as acceptance criteria: heavy build artifacts (node_modules, package stores, build output) must land on the main disk in the worktree, never on the shared 30 G /tmp.

Resolution, and the part that is actually the finding. Mos verified the attribution independently (mtimes, no process or lsof holding either path, no live session maps) and reaped both as lead coordinator: /tmp went to 79%, 6.0 G free. But note how it was resolved — a human-authority seat did it by hand, because the authority exists and the reaper does not. That gap is the finding, not the disk usage.

Two doctrine points fall out, both binding on RM-50:

  1. The fix is not "agents should tidy up." Asking each seat to clean its own scratch is instructions are not enforcement (D-4) wearing a different hat. A deterministic reaper must own it — same conclusion the north star reaches for every other class in this mission.
  2. Refusing to unilaterally delete another session's scratch was correct, and the resolution is not "be braver about deleting." An agent guessing that someone else's state is garbage is exactly the unreviewed destructive act the Constitution forbids. The resolution is that ownership and liveness become mechanically decidable, so reaping is a determination rather than a judgement call.

Reaper requirements for RM-50: liveness determined mechanically (process/lsof/session-map, not mtime alone); an age threshold; a dry-run that reports what it would reap and why; and an audit event per reap. Never a heuristic sweep — that would reintroduce the P-WORKFLOW-001 auto-sync failure in a more destructive form.

The self-erasure is the important part. An inert gate that is masked by unrelated downstream commits produces no lasting artifact, which is precisely why this class survives for months. Detection cannot rely on "is main currently red" — it must be per-merge-commit.

This matters more than the one-line fix:

  • It is the P-QUEUE-001 / P-CONFORMANCE-001 class ("gate-6 was inert fleet-wide"), reproduced in the repository this mission is remediating, discovered incidentally.
  • It independently validates OPUS premise A1 ("every gate is inert until proven otherwise") with live evidence rather than argument — which is why RM-02 is adopted as the keystone (§2, X2).
  • The file fix rides in its own hygiene PR. The inert gate itself is NOT quiet-patched. Per Mos: it stays a first-class backlog item, because patching the symptom would destroy the signal.

Binding requirement on RM-02 and RM-55: the gate registry and the conformance harness must assert "every merged commit passed every required gate" — evaluated per merge commit, against that commit's own tree, not against current main. As the table above proves, a "is main green today" check would have reported all-clear. A merged-commit-that-fails-a-required-gate is the exact detection signal, and it must be a registered must-fail case. A gate that cannot prove it blocked something has not been shown to work.


2. Where they genuinely disagree (not averaged — adjudicated)

# Axis OPUS SOL My ruling
X1 Total cost 38 tasks, ~5.3M tok 25 tasks, ~294K tok ~18× apart. Not reconcilable by splitting. They measure different things: SOL explicitly excludes orchestration/review/iteration overhead and assumes one remediation pass; OPUS prices the full loop. Adopt SOL's scope with OPUS's rigor, and treat SOL's G1 as a hard budget checkpoint (§4). Re-estimate empirically after the first three merged PRs rather than trusting either number.
X2 Gate registry (OPUS R-002) Keystone; blocks all P2 Absent; only a queue-guard fix ADOPT OPUS. Empirically validated in this very session: I found pnpm format:check red on main via merged PR #868 — a required gate that did not block. OPUS's premise A1 ("every gate is inert until proven otherwise") is not theoretical; it reproduced today, unprompted. Scope it tighter than 120K.
X3 Drizzle PG first-install defect (R-010) Hidden blocker; everything downstream depends on it Not mentioned ADOPT OPUS. packages/db/src/migrate.ts:30-38 carries a TODO admitting postgres-tier first-install fails today. The spine has only ever been proven on PGlite. Every later migration silently depends on this. SOL missed it.
X4 Rollback artifact for the hard cutover D3: hard cutover needs a rehearsed rollback snapshot SOL-07: import-only, explicitly no dual-write Both obey "no flat-file interim." OPUS wants a one-directional snapshot nothing reads as authority. I read that as compatible with the directive, but it is Jason's call → DECISION-2 (§5).
X5 Where the queue guard sits P0, independent of spine SOL-02, also early Agree it is P0. But ownership collides with parked PR #1023DECISION-3 (§5).
X6 Report-only rollout D5: only with a hard expiry, else withdraw not raised ADOPT OPUS. A report-only gate is by definition inert; expiry is the mechanism that stops it becoming the new fail-open.
X7 Availability trade (FC-7/FC-11) D8: "no DB ⇒ fleet stops" must be pre-committed in writing not raised Genuine availability regression, correctly identified. Needs Jason → folded into DECISION-2.

3. Reconciled DAG

Phases run in order; marks a hard barrier. src shows lineage (O=opus, S=sol, O+S=both). Estimates are given as a range (SOL low / OPUS high) rather than a fabricated midpoint — the spread is itself information, and X1 says we calibrate on real merged PRs.

P0 — Make gates provable, and stop the fleet re-bricking

No gate-introducing task in any later phase may merge before RM-02.

id task src depends_on est (S/O) tier
RM-01 Reproducible non-root checkout; gate fails on code, not env; heavy artifacts OFF shared /tmp (banks D-1/D-2/D-5/D-7) O+S+live 6K / 60K codex
RM-02 Gate registry + negative-control CI check (anti-inert-gate harness) ★keystone O RM-01 — / 120K opus
RM-03 ⏸HOLD Queue-guard: two defects — (a) unknown/no-status/malformed ⇒ ≠0, (b) guard evaluates branch=main instead of the branch being pushed O+S+live RM-02 8K / 100K sonnet
RM-04 Activation/version coherence; block launch on skew, fail SAFE; honest doctor labels O+S RM-01 (in S-01) / 140K sonnet
RM-05 Break-glass replaces the three silent MOSAIC BYPASS fail-opens O RM-04, RM-02 — / 120K opus

RM-05 must not merge before RM-04. The bypasses exist because the lease-broker daemon was never deployed on this host — removing the fail-open before deployment coherence is real re-creates the 2026-07-22 bricking incident. Hard edge, from OPUS.

P1 — Durable spine (PG)

No migration may merge before RM-10.

id task src depends_on est (S/O) tier
RM-10 Fix the Drizzle postgres-tier first-install defect ★hidden blocker O RM-01 — / 90K sonnet
RM-11 Orchestration spine schema (tasks, attempts, gate_results, hash-chained ledger, typed claims) O+S RM-10 12K / 160K opus
RM-12 Spine client, fail-closed connection (no silent PGlite in prod) O RM-11 — / 80K sonnet
RM-13 Atomic claims/transitions + transactional outbox + reconciliation sweeper O+S RM-12 12K / 140K opus

P2 — The single choke point

RM-25 (no-second-path) lands in the same milestone as RM-20, or the choke point is optional.

id task src depends_on est (S/O) tier
RM-20 Canonical MACP contract completion (Task/Result/Event/Claim/tri-state outcome) S 8K / (in R-020) codex
RM-21 Production TaskExecutor backed by @mosaicstack/macp ★keystone O+S RM-12, RM-02, RM-20 16K / 220K opus
RM-22 Gate-runner hardening: fail_on, timeouts, empty gate set = failure O RM-21 — / 120K sonnet
RM-23 Hash-chained MACPEvent ledger in PG + lifecycle EventType extension O+S RM-21, RM-11 — / 160K opus
RM-24 Seat identity from MOSAIC_AGENT_NAME + mandatory tri-state write outcomes O+S RM-21 (in S-03) / 150K opus
RM-25 No-second-path gate: terminal status writable only by the executor O RM-21, RM-23 — / 140K opus
RM-26 packages/coord submits through the executor (retire direct spawn) O+S RM-21 16K / 140K sonnet
RM-27 mosaic yolo/claude/codex/pi launch path records typed Task + events O RM-21, RM-23 — / 160K sonnet
RM-28 Delete the Forge stub executor (empty-gate-list "success"); Forge submits through the real one O+S RM-21 10K / 90K codex
RM-29 One-shot flat-file import + cutover readiness audit (dry-run, idempotent, no dual-write) S RM-13 8K / (in R-062) codex

★ G1 — FIRST DOGFOOD. Stop here and prove it. One live fleet task travels PG claim → TaskExecutor → worker → gates → terminal PG result/event, with no flat-file state. Adopted from SOL wholesale. If G1 cannot carry a real task, do not build Redis, rotation, comms, or conformance — remediate instead. This is the budget escape hatch (§4).

P3 — Rotation lifecycle (finish the Mission Control Plane)

id task src depends_on est (S/O) tier
RM-30 Typed state claims (source/confidence/TTL) with HMAC integrity, fail-closed O+S RM-11, RM-21 (in S-03) / 170K opus
RM-31 Contract-hash binding; stale generation loses mutation authority mechanically O+S RM-21, RM-30 12K / 180K opus
RM-32 Durable compaction/token sensor (per-runtime thresholds, PreCompact event) O RM-23, RM-31 — / 130K sonnet
RM-33 Typed checkpoint writer (structured claims, never transcript) + digest O+S RM-30, RM-32 12K / 150K opus
RM-34 Rotation daemon: watch → checkpoint → revoke → kill → relaunch → rehydrate O+S RM-33, RM-26 16K / 240K opus
RM-35 Rehydration attestation gate: refuse to act on an incomplete claim set O RM-33 — / 130K opus
RM-36 Broker-independent recovery; remove silent bypass; honest capability labels S RM-34 12K / (in R-004) sonnet
RM-37 Delete /compact and continue from the persistent-seat path (substitution, not removal) O+S RM-34, RM-44 (in S-16) / 60K codex

P4 — Comms service

RM-50 (one roster-owned socket per host) precedes identity-addressed delivery.

id task src depends_on est (S/O) tier
RM-40 comms/v1 envelope + protocol-version negotiation, LOUD reject O+S RM-11, RM-31 8K / 140K opus
RM-41 Comms service: PG state machine PENDING→RECEIVED→CONSUMED→DEAD-LETTER O+S RM-40, RM-13 16K / 200K opus
RM-42 tmux transport as a dumb adapter; durable retry before cursor advance O+S RM-41, RM-50 (in S-19) / 160K sonnet
RM-43 Per-class coalescing + supersede (the stale-consumed-as-live fix) O+S RM-41 12K / 130K sonnet
RM-44 Redis Streams hot delivery + provenance guard (Redis is never authority) O+S RM-41, RM-13 12K / 170K opus
RM-45 Retire direct tmux sends; only the service may write a pane O+S RM-42, RM-43 (in S-20) / 100K codex

P5 — Retirements, hygiene, conformance

id task src depends_on est (S/O) tier
RM-50 One roster-owned socket/host; quarantine unmanaged; deterministic reaper for stale sessions AND dead-session disk scratch (D-7) O+S+live RM-04 14K / 150K sonnet
RM-51 Auto-sync allowlist (never auto-stage unknown paths) + worktree/lease isolation O+S RM-02 8K / 110K sonnet
RM-52 Retire the Python controller + duplicate MACP islands (3 → 1) O+S RM-26, RM-27, RM-25, RM-28 14K / 110K codex
RM-53 Flat-file orchestration → DB hard cutover, with rehearsed rollback artifact O+S RM-27, RM-30, RM-34, RM-29 (in S-10) / 200K opus
RM-54 Fleet-wide inert-gate audit against the RM-02 registry O RM-02 — / 120K sonnet
RM-55 Conformance harness: fault-inject the live failure classes on real artifacts O+S RM-35, RM-41, RM-53 18K / 260K opus
RM-56 Retirement proof: CI asserts all three retirements are complete and stay complete O RM-52, RM-45, RM-53 — / 90K codex
RM-57 Operator cutover docs + activation proof; map all 15 decisions to evidence S RM-04, RM-36, RM-45, RM-55 6K / — codex
RM-58 Mechanical pre-dispatch context reset — the orchestrator resets a seat out-of-band and verifies it, rather than asking the agent to reset itself mos-remediation (D-4) RM-31, RM-50 8K sonnet

Critical path: RM-01 → RM-02 → RM-10 → RM-11 → RM-12 → RM-21 → RM-23 → RM-31 → RM-33 → RM-34 → RM-53 → RM-55.


4. Execution discipline

  • Every row is one PR. Author ≠ reviewer; rev-974 is the mosaicstack reviewer identity.
  • Pre-registered, diff-blind acceptance checks are committed BEFORE the reviewer reads the diff. Both decomps wrote their ACs in runnable ⇒0 / ⇒≠0 form specifically to make this possible.
  • Every gate-introducing task carries at least one registered must-fail negative control. This is RM-02's whole purpose; a gate with no proven failure path manufactures evidence.
  • Cost tiers: codex for mechanical/unambiguous, sonnet for normal feature work, opus reserved for security/integrity/cross-cutting-invariant tasks. SOL priced 0 opus tokens; OPUS priced 14 opus tasks. I am keeping opus only where the failure is integrity, not merely complexity.
  • G1 is the budget checkpoint. If the first-dogfood slice overruns SOL's estimate by >3×, stop and re-plan rather than spending the remainder. X1 says neither estimate is trustworthy until calibrated.
  • Defer list adopted from SOL (10 items): mission dashboard/TUI, PRD-to-board auto-decomposition, heuristic churn scoring, Discord/Slack/Telegram adapters, public MCP comms surface, protocol-v2 negotiation, multi-region PG/Redis, event analytics UI.

5. Decisions — all three ruled by Mos, 2026-07-31

DECISION-1 — the wire-in point. RULED: accept the planners (Mos, 2026-07-31). The charter's mosaic_orchestrator.py::run_single_task target is the disabled Python controller this mission retires; wiring the new choke point into the rail we are deleting is wrong.

Corrected target (authoritative): a new production Node TaskExecutor sitting on the live dispatch pathpackages/mosaic launch + packages/coord — which Coord, Forge, and live dispatch all submit through. This is the MACP scout's full recommendation ("replace the block with a Node executor and make Coord/Forge submit through it"), not a resurrection of the Python controller. RM-52 is therefore a deletion task, and Build 1's acceptance is measured on a live mosaic yolo invocation.

Mos ruled this resolvable from the already-accepted retire-the-Python-rail decision — his authority, not a Jason escalation. RM-21/RM-26/RM-27/RM-52 all take the corrected target.

DECISION-2 — rollback artifact + availability trade. ⏸ JASON-PENDING — NOT BLOCKING. The DB build is phases away, so this is queued for Jason's next session rather than escalated now. Binding requirement in the meantime (Mos, from P-RECOVERY-001): the DB spine must NOT be a single-point hard-stop. Design for a broker-independent / degraded mode plus a rollback artifact. Jason finalises only the specific availability target. This reverses my earlier reading of OPUS D8 ("the fallback is: the fleet stops") — that answer is not pre-committed; a degraded mode is now a design requirement on RM-12, RM-13, RM-23, RM-36 and RM-53.

DECISION-3 — RM-03 vs. parked PR #1023. RULED: HOLD RM-03 (Mos, 2026-07-31). Do not open a third gate-6 lane — that is the postmortem's own anti-pattern performed by the remediation. PR #1023 sits in Jason's parked delivery stack; its disposition (close, or supersede by RM-03) is Jason's at his next session.

  • PR #1023 → SUPERSEDED-PENDING-JASON. RM-03 stays HOLD; when Jason rules, RM-03 proceeds as the single correct lane.
  • RM-02 and RM-55 proceed independently and are NOT held. The per-merge-commit gate-assertion requirement is the conformance capability, not the gate-6 fix itself — different scope, no ownership collision.

6. Status

phase state
Decomposition DONE — both planners delivered independently
Reconciliation DONE — this document
Blocking decisions RULED — all 3 closed by Mos 2026-07-31 (§5); D-2's availability target is Jason-pending but non-blocking
Dispatch RM-01 IN FLIGHT — f10-coder (codex), worktree-isolated, AC1AC8 pre-registered
Review PR #1025 with rev-974; ACs pre-registered 22:12:26Z before diff exposure