Compare commits

..
Author SHA1 Message Date
mos-dt-0andClaude Opus 5 b705a4dae9 docs(remediation): sharpen D-8 — the conformance harness must itself be conformance-tested
ci/woodpecker/pr/ci Pipeline was successful
Per Mos. Two requirements, both load-bearing:

RM-02: the gate registry must SELF-VERIFY that every registered case demonstrably
runs and demonstrably fails on a known-bad input. Presence in the registry is not
evidence. A check is not trusted until it has been shown to fail. This is mutation
testing / negative control at the registry level — the conformance harness must
itself be conformance-tested. RM-55 inherits the same recursion: the harness must be
observed red before its green means anything.

Reviewer disclosure: rev-974 could have silently substituted a working fixture and
reported AC2 PASS, and nothing would have caught it. The 'if a check is unrunnable,
say so rather than silently substituting' clause is what makes a PASS mean anything,
and must appear in every reviewer brief this mission issues.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:28:35 -05:00
mos-dt-0andClaude Opus 5 27350d9f8d docs(remediation): board — #1025 merged, D-8 banked, phase EXECUTING
ci/woodpecker/pr/ci Pipeline was canceled
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:27:10 -05:00
mos-dt-0andClaude Opus 5 f8b65463ab docs(remediation): bank D-8 — a pre-registered check that was not runnable as written
ci/woodpecker/pr/ci Pipeline was canceled
PR #1025's AC2 fixture (mkdir -p apps/*/venv/lib) does not do what it claims when the
glob is unmatched. rev-974 ran it exactly as written, caught the defect, re-ran the
intended assertion at an explicit path, and DISCLOSED the substitution rather than
silently reporting PASS.

Pre-registration protects a check from being retrofitted to the implementation. It
does not make the check correct. An unverified gate appeared inside the mechanism
built to catch unverified gates.

Requirement added to RM-02: the registry must self-verify that every registered case
actually runs and can actually fail — presence in the registry is not evidence.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:25:56 -05:00
fb3221af27 docs(remediation): D-7 reaper doctrine into RM-50; adopt board-roll for the LIVE board
ci/woodpecker/pr/ci Pipeline was canceled
D-7 resolution recorded. Mos verified attribution independently and reaped ~5.2G of
dead-session scratch as lead coordinator. The finding is the gap it exposed: the
authority to reap exists, the deterministic reaper does not.

Two doctrine points now binding on RM-50: (1) 'agents should tidy up' is
instructions-are-not-enforcement wearing another hat — a reaper must own it;
(2) refusing to unilaterally delete another session's scratch was correct, and the
fix is mechanically-decidable ownership/liveness, not braver deletion. Reaper
requirements specified: mechanical liveness (not mtime alone), age threshold,
dry-run, audit event per reap — never a heuristic sweep, which would reintroduce
the P-WORKFLOW-001 auto-sync failure in a more destructive form.

Board restructured into an explicit BOARD-ROLL zone and rolled with the framework's
own board-roll.sh (8432B -> 8014B, D-1 archived to BOARD-LEDGER.md) rather than
hand-trimmed. The tool exists precisely to stop coordinators hand-trimming; using it
is the dogfood.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:25:04 -05:00
910b6f4166 docs(remediation): status — decisions ruled, RM-01 in flight
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:25:04 -05:00
1324385bba docs(remediation): apply Mos rulings on D-1/D-2/D-3; bank D-7 tmpfs ENOSPC
D-1 RULED — corrected wire-in target is a new production Node TaskExecutor on the
LIVE dispatch path (packages/mosaic launch + packages/coord) that Coord/Forge/live
dispatch submit through. The Python rail is deleted, not ported. RM-21/26/27/52 updated.

D-2 — Jason-pending and explicitly NON-blocking. Binding meanwhile: the DB spine must
not be a single-point hard-stop; degraded mode + rollback artifact become design
requirements on RM-12/13/23/36/53. This reverses the earlier recommendation to
pre-commit 'no DB means the fleet stops'.

D-3 — RM-03 HOLD; PR #1023 marked SUPERSEDED-PENDING-JASON; RM-02/RM-55 proceed
independently since conformance assertion is not the gate-6 fix itself.

RM-03 now carries BOTH queue-guard defects: wrong exit on unknown, and evaluating
branch=main rather than the branch being pushed.

D-7 — shared 30G /tmp hit 100% ENOSPC, degrading two seats. ~5.2G of it is session
scratch dead for 8-9 days; this session's own footprint is 88K. Same missing
capability as P-FLEET-001 stale-session GC, applied to disk. Folded into RM-50, and
RM-01 gains artifacts-off-shared-tmp acceptance criteria.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:25:04 -05:00
d4f68040ab docs(remediation): board — PRs #1025/#1026 open, review dispatched, D-6 queue-guard finding
Third independent live instance of the inert-gate class this session: the mandated
ci-queue-wait.sh returned exit 0 on state=unknown while guarding a push.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:25:04 -05:00
ba98f88891 docs(remediation): reconciled execution backlog from two independent decompositions
TASK-1 complete. planner-opus (robustness, 38 tasks/8 dissents) and planner-sol
(pragmatic, 25 tasks/10 defers/7 dissents) each decomposed the 4-build plan without
seeing the other's work. Both decomps are committed alongside the reconciliation so
the disagreements stay auditable rather than being flattened into a consensus.

Reconciled into 58 tasks across P0-P5 (docs/remediation/TASKS.md):
- 7 independent convergences, treated as settled because neither planner could see
  the other. The headline: BOTH reject the charter's wire-in point
  (mosaic_orchestrator.py::run_single_task) because that controller is disabled and
  references a dispatcher absent from this checkout — wiring it would produce a
  stranded executor, the same built-but-unwired disease one layer up.
- 7 genuine disagreements ADJUDICATED, not averaged. The cost estimates are ~18x
  apart; rather than split the difference, the plan adopts sol's scope with opus's
  rigor and treats the first-dogfood gate as a hard budget checkpoint.
- 3 decisions escalated (wire-in point, rollback artifact + availability trade,
  queue-guard ownership vs parked PR #1023). No task blocked on them is dispatched.

Keystone dogfood case recorded (TASKS.md 1a): merged PR #868 shipped a file failing
pnpm format:check, then an unrelated PR reformatted it as a side effect, so main went
green and the gate's failure to fire left no trace. Detection must therefore be
per-merge-commit against that commit's own tree.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:25:04 -05:00
4fcbf7c07e docs(remediation): board reflects actual state + 4 dogfood findings from startup
Corrects the premature "DISPATCHED" planner line (per Mos): both planners are now
dispatched for real on GUARANTEED-clean context, not requested-clean.

Records Mos rulings: mission.json is retired-rail residue (do not invest); gate-16
holds on interim mos-dt-0 since rev-974 reviews; remote-control path is Mos-relay.

Captures 4 live failure classes observed while standing this seat up (D-1..D-4),
including an agent that silently ignored an in-message context reset — direct
evidence for the postmortem thesis that instructions are not enforcement.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:25:04 -05:00
28f1e272fe docs(remediation): mission charter + kickstart + live board for the postmortem remediation
15/15 proposals decided (13 accept, 2 modify). Collapses to 4 builds + hygiene on one
PG spine + choke-point service. Durable mission record for the mos-remediation project
orchestrator; compaction-survival resume in KICKSTART.md.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_018XCrpMFmCAXbtSQtQAhDth
2026-07-31 17:25:04 -05:00
mos-dt-0andMos 524146055d fix(hygiene): .prettierignore must exclude Python build/test artifacts (#1025)
ci/woodpecker/push/publish Pipeline failed
ci/woodpecker/push/ci Pipeline was successful
Co-authored-by: mos-dt-0 <[email protected]>
2026-07-31 22:24:49 +00:00
3 changed files with 210 additions and 77 deletions
+15
View File
@@ -0,0 +1,15 @@
<!-- board-roll: 1 entry rolled from BOARD.md -->
### **D-1 / P-ACTIVATION + hygiene — committed `.npmrc` hard-pins `store-dir=/root/.local/share/pnpm/store`.**
Correct for the CI container (runs as root), fatal for EVERY non-root local checkout: `EACCES` on `/root/.local/share/pnpm/store/v10/server/server.json`. A committed config that only works on one runtime is exactly the activation-skew class. Fix candidate: make store-dir env-overridable, not hardcoded.
<!-- board-roll: 2 entries rolled from BOARD.md -->
### **D-3 / P-FLEET-001 — the seats running this mission are UNMANAGED.** `mos-remediation`, `rev-974`,
`planner-opus`, `planner-sol` appear in NO roster (`~/.config/mosaic/fleet/roster.yaml`, `agents/`). Planners run on socket `default`; the roster declares `mosaic-fleet`. This is the exact "one roster-owned socket/host + quarantine unmanaged + stale GC" failure P-FLEET-001 indicts — observed on the remediation mission's own fleet. Prerequisite for INBOX identity-addressing.
### **D-2 / hygiene — husky `prepare` fails `EPERM` copying into root-owned `.husky/_/`.** Repo working
tree has root-owned dirs (`.husky/`, repo root) under a non-root agent. Worked around with the intended `HUSKY=0` escape hatch (does NOT disable the existing pre-commit/pre-push hooks).
+35 -38
View File
@@ -1,6 +1,6 @@
# mos-remediation — LIVE BOARD (keep < 8 KB) # mos-remediation — LIVE BOARD (keep < 8 KB)
**Phase:** PLANNING COMPLETE — reconciled backlog landed; 3 decisions block dispatch. **Phase:** EXECUTING — first PR merged; RM-01 in flight; all 3 decisions ruled.
**Updated:** 2026-07-31 (mos-remediation orchestrator; seat active on `mosaic-fleet`). **Updated:** 2026-07-31 (mos-remediation orchestrator; seat active on `mosaic-fleet`).
## Head ## Head
@@ -16,14 +16,14 @@
## In-flight ## In-flight
| Task | Owner | State | | Task | Owner | State |
| ------------------------------------------------- | --------------- | ----------------------------------------------------- | | ------------------------------------------------- | --------------- | ------------------------------------------------------------------ |
| PR #1026 `remediation/mission-setup` (docs-only) | mos-remediation | OPEN, stacked on #1025; retarget to main after merge | | PR #1026 docs (mission record + backlog) | mos-remediation | OPEN, retargeted to main, rebased; diff verified docs-only |
| PR #1025 `fix/hygiene-inert-format-gate` | rev-974 | OPEN; ACs pre-registered 22:12:26Z; review DISPATCHED | | PR #1025 hygiene | — | **MERGED** 52414605; rev-974 APPROVE + CI #2158 8/8 terminal-green |
| DECISION-1 wire-in point (charter change) | Mos / Jason | ESCALATED — both planners reject the charter's target | | DECISION-1 wire-in point (charter change) | Mos / Jason | ESCALATED — both planners reject the charter's target |
| DECISION-2 rollback artifact + availability trade | Jason | ESCALATED | | DECISION-2 rollback artifact + availability trade | Jason | ESCALATED |
| DECISION-3 RM-03 vs parked PR #1023 ownership | Mos | ESCALATED | | DECISION-3 RM-03 vs parked PR #1023 ownership | Mos | ESCALATED |
| RM-01 reproducible checkout (unblocks everything) | unassigned | READY TO DISPATCH | | RM-01 reproducible checkout (unblocks everything) | unassigned | READY TO DISPATCH |
## Fleet seats ## Fleet seats
@@ -53,37 +53,34 @@
4. Hygiene + conformance harness 4. Hygiene + conformance harness
Cross-cutting retirements: flat-file tracking, 3 MACP islands, silent MOSAIC BYPASS. Cross-cutting retirements: flat-file tracking, 3 MACP islands, silent MOSAIC BYPASS.
## Dogfood evidence captured this session (live failure classes, not hypotheticals) ## Dogfood evidence live failure classes, not hypotheticals
- **D-1 / P-ACTIVATION + hygiene — committed `.npmrc` hard-pins `store-dir=/root/.local/share/pnpm/store`.** > Newest first. Oldest entries roll to `BOARD-LEDGER.md` via `board-roll.sh` when this file
Correct for the CI container (runs as root), fatal for EVERY non-root local checkout: `EACCES` on > exceeds its 8 KB cap. Keystone detail is duplicated in `TASKS.md` §1a, so rolling loses nothing.
`/root/.local/share/pnpm/store/v10/server/server.json`. A committed config that only works on one
runtime is exactly the activation-skew class. Fix candidate: make store-dir env-overridable, not hardcoded.
- **D-2 / hygiene — husky `prepare` fails `EPERM` copying into root-owned `.husky/_/`.** Repo working
tree has root-owned dirs (`.husky/`, repo root) under a non-root agent. Worked around with the
intended `HUSKY=0` escape hatch (does NOT disable the existing pre-commit/pre-push hooks).
- **D-3 / P-FLEET-001 — the seats running this mission are UNMANAGED.** `mos-remediation`, `rev-974`,
`planner-opus`, `planner-sol` appear in NO roster (`~/.config/mosaic/fleet/roster.yaml`, `agents/`).
Planners run on socket `default`; the roster declares `mosaic-fleet`. This is the exact
"one roster-owned socket/host + quarantine unmanaged + stale GC" failure P-FLEET-001 indicts —
observed on the remediation mission's own fleet. Prerequisite for INBOX identity-addressing.
- **D-4 / P-LIFECYCLE + hygiene — a dispatched agent silently IGNORED an in-message context reset.**
planner-sol was at 64.3%/372k; the brief asked it to reset first; it began work on dirty context anyway.
Only an out-of-band `/new` driven by the orchestrator guaranteed clean state. Confirms the postmortem
thesis: **instructions are not enforcement.** Reset must be a mechanical pre-dispatch step, not a request.
- **D-6 / P-QUEUE-001 — the mandated queue guard returned PASS on an UNKNOWN state, live, today.** <!-- BOARD-ROLL:START -->
Running the required `ci-queue-wait.sh --purpose push` before pushing produced
`state=unknown ... exit 0` — the exact defect at `ci-queue-wait.sh:282-288` that PR #1023 is parked ### **D-8 / P-CONFORMANCE-001 — a PRE-REGISTERED acceptance check that was not runnable as written.**
on. It also evaluated `branch=main` rather than the branch being pushed. The mission's own required
pre-push gate passed me on an indeterminate result. Third independent live instance of the class. PR #1025 AC2's fixture `mkdir -p apps/*/venv/lib` creates a literal `apps/*/venv/lib` dir when the glob is unmatched — it did not test what it claimed. rev-974 ran it exactly as written, caught it, re-ran the intended assertion at an explicit path, and **disclosed** rather than silently substituting a working fixture and reporting PASS. **Pre-registration protects a check from being retrofitted to the implementation; it does not make the check correct.** An unverified gate appeared inside the mechanism built to catch unverified gates. Hard requirement on RM-02: the registry must self-verify that every registered case runs AND can fail — presence is not evidence.
- **D-5 / P-QUEUE-001 + P-CONFORMANCE-001 — KEYSTONE: an inert gate that erased its own evidence.**
Merged PR #868 (`b79336a8`) shipped a file that FAILS `pnpm format:check` ⇒ the CI format gate did ### **D-7 / P-FLEET-001 — stale-GC-on-disk: shared 30G /tmp hit 100% ENOSPC, degrading two seats.**
not block. An unrelated later PR (#872) then reformatted that file via its own `lint-staged`, so
`main` went green again and nobody learned the gate had failed to fire. Verified blob-level under the ~5.2G was session scratch dead 8-9 days (this session's own footprint: 88K). Same missing capability as orphaned-tmux-session GC, applied to disk — not a quota or discipline problem. Resolved manually by Mos (lead coordinator) after independent verification; `/tmp` now 79%. **The gap IS the finding:** the authority to reap exists, the deterministic reaper does not. Folded into RM-50 with explicit requirements (mechanical liveness, age threshold, dry-run, audit event per reap — never a heuristic sweep). Refusing to unilaterally delete another session's scratch was correct doctrine; the fix is a reaper, not braver agents.
repo's own config. **Detection must be per-merge-commit against that commit's own tree** — a "is main
green today" check reports all-clear on this exact defect. Binding on RM-02/RM-55. Full chain in ### **D-6 / P-QUEUE-001 — the mandated queue guard returned PASS on an UNKNOWN state, live, today.**
`TASKS.md` §1a. NOT quiet-patched, by Mos's ruling: patching the symptom destroys the signal.
Running the required `ci-queue-wait.sh --purpose push` before pushing produced `state=unknown ... exit 0` — the exact defect at `ci-queue-wait.sh:282-288` that PR #1023 is parked on. It also evaluated `branch=main` rather than the branch being pushed. The mission's own required pre-push gate passed me on an indeterminate result. Third independent live instance of the class.
### **D-5 / P-QUEUE-001 + P-CONFORMANCE-001 — KEYSTONE: an inert gate that erased its own evidence.**
Merged PR #868 (`b79336a8`) shipped a file that FAILS `pnpm format:check` ⇒ the CI format gate did not block. An unrelated later PR (#872) then reformatted that file via its own `lint-staged`, so `main` went green again and nobody learned the gate had failed to fire. Verified blob-level under the repo's own config. **Detection must be per-merge-commit against that commit's own tree** — a "is main green today" check reports all-clear on this exact defect. Binding on RM-02/RM-55. Full chain in `TASKS.md` §1a. NOT quiet-patched, by Mos's ruling: patching the symptom destroys the signal.
### **D-4 / P-LIFECYCLE + hygiene — a dispatched agent silently IGNORED an in-message context reset.**
planner-sol was at 64.3%/372k; the brief asked it to reset first; it began work on dirty context anyway. Only an out-of-band `/new` driven by the orchestrator guaranteed clean state. Confirms the postmortem thesis: **instructions are not enforcement.** Reset must be a mechanical pre-dispatch step, not a request.
<!-- BOARD-ROLL:END -->
## Decisions log ## Decisions log
+160 -39
View File
@@ -4,8 +4,8 @@
**Sources:** [`DECOMP-OPUS.md`](./DECOMP-OPUS.md) (robustness, 38 tasks / 8 dissents) and **Sources:** [`DECOMP-OPUS.md`](./DECOMP-OPUS.md) (robustness, 38 tasks / 8 dissents) and
[`DECOMP-SOL.md`](./DECOMP-SOL.md) (pragmatic, 25 tasks / 10 defers / 7 dissents), produced [`DECOMP-SOL.md`](./DECOMP-SOL.md) (pragmatic, 25 tasks / 10 defers / 7 dissents), produced
**independently** — neither planner read the other. Charter: [`MISSION.md`](./MISSION.md). **independently** — neither planner read the other. Charter: [`MISSION.md`](./MISSION.md).
**Status:** PLANNING — this backlog is proposed, not yet dispatched. Three items need a Mos/Jason **Status:** EXECUTING — all three blocking decisions RULED by Mos on 2026-07-31 (§5). **RM-01 is
ruling before the affected tasks dispatch (§5). dispatched.** RM-03 is held pending Jason's disposition of PR #1023; nothing else is blocked.
> **Provenance of the inputs (both clean).** `planner-opus` ran in a fresh session throughout. > **Provenance of the inputs (both clean).** `planner-opus` ran in a fresh session throughout.
> `planner-sol` initially began work at 64.3% dirty context despite a brief instructing it to reset; > `planner-sol` initially began work at 64.3% dirty context despite a brief instructing it to reset;
@@ -40,7 +40,10 @@ point. Both planners independently rejected it on the same evidence: that contro
point into a disabled rail produces **a stranded executor — the identical built-but-unwired disease, point into a disabled rail produces **a stranded executor — the identical built-but-unwired disease,
one layer up, that would look "done" in a PR.** The live paths are one layer up, that would look "done" in a PR.** The live paths are
`packages/mosaic/src/commands/launch.ts` and `packages/coord/src/runner.ts`. `packages/mosaic/src/commands/launch.ts` and `packages/coord/src/runner.ts`.
This contradicts the charter and is escalated as **DECISION-1** (§5). This contradicted the charter and was escalated as DECISION-1 — **now RULED in the planners' favour by
Mos (§5)**. The corrected target is a new production Node `TaskExecutor` on the live dispatch path
(`packages/mosaic` launch + `packages/coord`) that Coord/Forge/live dispatch submit through; the
Python rail is deleted, not ported.
--- ---
@@ -66,6 +69,115 @@ side effect, so `main` went green again **without anyone ever learning the gate
> present-tense framing was wrong. The hygiene PR therefore carries the `.prettierignore` fix only — > present-tense framing was wrong. The hygiene PR therefore carries the `.prettierignore` fix only —
> the README needs no fix today. > the README needs no fix today.
### Third live instance, same class — the queue guard, hit by this orchestrator
Running the **mandated** pre-push guard during TASK-0:
```
$ ~/.config/mosaic/tools/git/ci-queue-wait.sh --purpose push
[ci-queue-wait] platform=gitea purpose=push branch=main sha=06e0d403…
[ci-queue-wait] state=unknown purpose=push branch=main
$ echo $? → 0
```
**Two distinct defects in one tool**, both feeding RM-03:
1. **Wrong exit**`state=unknown``exit 0`. The defect at `ci-queue-wait.sh:282-288` that OPUS
documented and that PR #1023 is parked on. A required gate returned PASS on an indeterminate result.
2. **Wrong branch** — it evaluated `branch=main`, not the branch actually being pushed. Even a
correctly-exiting guard would have been answering the wrong question.
**Standing doctrine (Mos):** until RM-03 lands, a green from this guard carries **zero information**
and must not be cited as merge evidence. Rely on reviewer clearance + real CI.
Three independent live instances in a single session — format gate, agent context reset, queue guard —
is the class confirmed, not anecdote.
### D-8 — a PRE-REGISTERED acceptance check that was not runnable as written
On PR #1025 the author (me) pre-registered AC2 with the fixture snippet `mkdir -p apps/*/venv/lib`.
In bash, when no `venv` exists the glob is unmatched and passes through literally, creating a
directory named `apps/*/venv/lib` rather than one per workspace. The check as written did not test
what it claimed to test.
`rev-974` ran it **exactly as written**, observed the wrong behaviour, then re-ran the intended
assertion at an explicit path — **and said so in the review** rather than silently substituting a
working fixture and reporting PASS.
Two things this establishes:
1. **The instruction "do not adjust a check to fit the diff; if it is unrunnable, say so explicitly"
worked.** A silent substitution here would have produced a green AC2 that proved nothing, on the
exact task whose subject is gates that appear to work. The disclosure is what made the PASS
meaningful.
2. **Pre-registration does not confer correctness.** A pre-registered check is protected from being
retrofitted to the implementation; it is not protected from being _wrong when written_. This is a
small instance of the mission's own class — an unverified gate — occurring inside the mechanism
built to catch unverified gates.
**Requirement on RM-02 (non-negotiable, sharpened by Mos).** The registry must **self-verify** that
every registered case demonstrably **runs** and demonstrably **fails on a known-bad input**.
Presence in the registry is **not** evidence. **A check is not trusted until it has been shown to
fail.** This is mutation testing / negative control applied _at the registry level_ — meaning
**the conformance harness must itself be conformance-tested.** A registered case that cannot fail, or
cannot run, is exactly as inert as an unregistered one, and the registry check must detect that
itself rather than assume it.
**Requirement on RM-55.** The same recursion applies to the harness: it must be observed red before
its green is worth anything (OPUS R-063 AC1 already states this; D-8 is the empirical case for it).
**Second, equally load-bearing lesson — reviewer disclosure is what makes a review trustworthy.**
rev-974 could have silently swapped in a working fixture and reported `AC2 PASS`. Nothing in the
process would have caught it, and the resulting green would have certified nothing — on the very task
whose subject is gates that only appear to work. The brief's instruction — _"do not adjust a check to
fit the diff; if it is genuinely unrunnable as specified, say so explicitly and explain why rather
than silently substituting your own"_ — is therefore not boilerplate. It is the clause that makes a
PASS mean something, and it must appear in **every** reviewer brief this mission issues.
### D-7 — shared-tmpfs contention → cascading ENOSPC (live incident, 2026-07-31)
The shared 30 G `/tmp` hit **100% ENOSPC** mid-session. It broke tool calls in **two different seats**
(mine and Mos's) — a single full disk degrades every agent on the host at once. Recurring: prior
incidents 2026-06-18 and 2026-07-17.
Attribution matters, because the wrong owner cleans the wrong thing. Measured:
| path | size | last modified | owner |
| ---------------------------------------------- | --------- | ---------------------------- | ------------------------------------------------- |
| `…/-src-mosaic-stack/6d2faee6…` (this session) | **88 K** | live | mos-remediation |
| `…/-src-mosaic-stack/c743185d…` | **3.6 G** | **2026-07-22** (9 days dead) | abandoned session, same project path |
| `…/claude-1001/pnpm-store` | **1.6 G** | **2026-07-23** (8 days dead) | abandoned; the live store is correctly on `$HOME` |
So ~5.2 G — the bulk of the pressure — is **dead session scratch that nothing will ever read again**.
This is not a quota problem; it is **P-FLEET-001's stale-session GC, applied to disk instead of tmux
sessions.** The same missing capability (nothing owns reaping dead ephemeral state) produces both the
orphaned-session failure and this one. Reaping dead-session scratch belongs in RM-50 alongside stale
tmux-session GC.
**Added to RM-01 as acceptance criteria:** heavy build artifacts (node_modules, package stores, build
output) must land on the main disk in the worktree, never on the shared 30 G `/tmp`.
**Resolution, and the part that is actually the finding.** Mos verified the attribution independently
(mtimes, no process or `lsof` holding either path, no live session maps) and reaped both as lead
coordinator: `/tmp` went to 79%, 6.0 G free. But note _how_ it was resolved — **a human-authority seat
did it by hand, because the authority exists and the reaper does not.** That gap is the finding, not
the disk usage.
Two doctrine points fall out, both binding on RM-50:
1. **The fix is not "agents should tidy up."** Asking each seat to clean its own scratch is
`instructions are not enforcement` (D-4) wearing a different hat. A deterministic reaper must own
it — same conclusion the north star reaches for every other class in this mission.
2. **Refusing to unilaterally delete another session's scratch was correct, and the resolution is not
"be braver about deleting."** An agent guessing that someone else's state is garbage is exactly the
unreviewed destructive act the Constitution forbids. The resolution is that _ownership and liveness
become mechanically decidable_, so reaping is a determination rather than a judgement call.
**Reaper requirements for RM-50:** liveness determined mechanically (process/`lsof`/session-map, not
mtime alone); an age threshold; a dry-run that reports what it would reap and why; and an audit event
per reap. Never a heuristic sweep — that would reintroduce the P-WORKFLOW-001 auto-sync failure in a
more destructive form.
**The self-erasure is the important part.** An inert gate that is masked by unrelated downstream **The self-erasure is the important part.** An inert gate that is masked by unrelated downstream
commits produces no lasting artifact, which is precisely why this class survives for months. Detection commits produces no lasting artifact, which is precisely why this class survives for months. Detection
cannot rely on "is `main` currently red" — it must be per-merge-commit. cannot rely on "is `main` currently red" — it must be per-merge-commit.
@@ -112,13 +224,13 @@ spread is itself information, and X1 says we calibrate on real merged PRs.
_No gate-introducing task in any later phase may merge before RM-02._ _No gate-introducing task in any later phase may merge before RM-02._
| id | task | src | depends_on | est (S/O) | tier | | id | task | src | depends_on | est (S/O) | tier |
| ----- | -------------------------------------------------------------------------------------------- | --- | ------------ | ---------------- | ------ | | --------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ------------ | ---------------- | ------ |
| RM-01 | Reproducible non-root checkout; pre-push gate fails on **code, not env** (banks D-1/D-2/D-5) | O+S | — | 6K / 60K | codex | | RM-01 | Reproducible non-root checkout; gate fails on **code, not env**; heavy artifacts OFF shared `/tmp` (banks D-1/D-2/D-5/D-7) | O+S+live | — | 6K / 60K | codex |
| RM-02 | **Gate registry + negative-control CI check** (anti-inert-gate harness) ★keystone | O | RM-01 | — / 120K | opus | | RM-02 | **Gate registry + negative-control CI check** (anti-inert-gate harness) ★keystone | O | RM-01 | — / 120K | opus |
| RM-03 | Queue-guard fail-closed rework (`unknown`/`no-status`/malformed ⇒ ≠0) | O+S | RM-02 | 8K / 100K | sonnet | | RM-03 ⏸**HOLD** | Queue-guard: **two** defects — (a) `unknown`/`no-status`/malformed ⇒ ≠0, (b) guard evaluates `branch=main` instead of the branch being pushed | O+S+live | RM-02 | 8K / 100K | sonnet |
| RM-04 | Activation/version coherence; block launch on skew, fail SAFE; honest `doctor` labels | O+S | RM-01 | (in S-01) / 140K | sonnet | | RM-04 | Activation/version coherence; block launch on skew, fail SAFE; honest `doctor` labels | O+S | RM-01 | (in S-01) / 140K | sonnet |
| RM-05 | Break-glass replaces the three silent `MOSAIC BYPASS` fail-opens | O | RM-04, RM-02 | — / 120K | opus | | RM-05 | Break-glass replaces the three silent `MOSAIC BYPASS` fail-opens | O | RM-04, RM-02 | — / 120K | opus |
> ⚠ **RM-05 must not merge before RM-04.** The bypasses exist because the lease-broker daemon was > ⚠ **RM-05 must not merge before RM-04.** The bypasses exist because the lease-broker daemon was
> never _deployed_ on this host — removing the fail-open before deployment coherence is real > never _deployed_ on this host — removing the fail-open before deployment coherence is real
@@ -187,7 +299,7 @@ spread is itself information, and X1 says we calibrate on real merged PRs.
| id | task | src | depends_on | est (S/O) | tier | | id | task | src | depends_on | est (S/O) | tier |
| ----- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------- | -------------------------- | ---------------- | ------ | | ----- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------- | -------------------------- | ---------------- | ------ |
| RM-50 | One roster-owned socket/host; quarantine unmanaged; stale-session GC | O+S | RM-04 | 14K / 150K | sonnet | | RM-50 | One roster-owned socket/host; quarantine unmanaged; **deterministic reaper for stale sessions AND dead-session disk scratch** (D-7) | O+S+live | RM-04 | 14K / 150K | sonnet |
| RM-51 | Auto-sync **allowlist** (never auto-stage unknown paths) + worktree/lease isolation | O+S | RM-02 | 8K / 110K | sonnet | | RM-51 | Auto-sync **allowlist** (never auto-stage unknown paths) + worktree/lease isolation | O+S | RM-02 | 8K / 110K | sonnet |
| RM-52 | Retire the Python controller + duplicate MACP islands (3 → 1) | O+S | RM-26, RM-27, RM-25, RM-28 | 14K / 110K | codex | | RM-52 | Retire the Python controller + duplicate MACP islands (3 → 1) | O+S | RM-26, RM-27, RM-25, RM-28 | 14K / 110K | codex |
| RM-53 | Flat-file orchestration → DB hard cutover, with rehearsed rollback artifact | O+S | RM-27, RM-30, RM-34, RM-29 | (in S-10) / 200K | opus | | RM-53 | Flat-file orchestration → DB hard cutover, with rehearsed rollback artifact | O+S | RM-27, RM-30, RM-34, RM-29 | (in S-10) / 200K | opus |
@@ -219,40 +331,49 @@ spread is itself information, and X1 says we calibrate on real merged PRs.
--- ---
## 5. Blocking decisions (need Mos, or Jason via Mos) ## 5. Decisions — all three ruled by Mos, 2026-07-31
These are escalated because they change the charter, the accepted directive, or another lane's **DECISION-1 — the wire-in point. ✅ RULED: accept the planners (Mos, 2026-07-31).**
ownership — none is a question I should answer unilaterally. The charter's `mosaic_orchestrator.py::run_single_task` target is the **disabled Python controller
this mission retires**; wiring the new choke point into the rail we are deleting is wrong.
**DECISION-1 — the wire-in point (changes `MISSION.md`).** > **Corrected target (authoritative):** a **new production Node `TaskExecutor`** sitting on the
Both planners independently reject wiring the choke point into > **live dispatch path** — `packages/mosaic` launch + `packages/coord` — which Coord, Forge, and live
`mosaic_orchestrator.py::run_single_task`, because that controller is disabled and references a > dispatch all **submit through**. This is the MACP scout's _full_ recommendation ("replace the block
non-existent dispatcher. They propose wiring `packages/coord/src/runner.ts` + > **with** a Node executor **and** make Coord/Forge submit through it"), not a resurrection of the
`packages/mosaic/src/commands/launch.ts` and **deleting** the Python rail instead. This makes RM-52 a > Python controller. RM-52 is therefore a **deletion** task, and Build 1's acceptance is measured on a
deletion task rather than an integration task, and moves Build 1's acceptance onto a live > live `mosaic yolo` invocation.
`mosaic yolo` invocation. _My recommendation: accept — wiring a disabled rail reproduces the exact
disease this mission exists to cure._
**DECISION-2 — rollback artifact + the availability trade (needs Jason).** Mos ruled this resolvable from the already-accepted retire-the-Python-rail decision — his authority,
(a) Does "hard cutover, no flat-file interim" permit a **rehearsed, one-directional, read-as-authority- not a Jason escalation. RM-21/RM-26/RM-27/RM-52 all take the corrected target.
by-nobody** rollback snapshot (OPUS D3)? (b) Do we pre-commit in writing that "no DB ⇒ the fleet
stops" and "unpersistable event ⇒ the operation fails" (OPUS D8)? That is a real availability
regression versus today's limping flat-file fleet — correct for a system whose defining failure is
_silent continuation_, but it should be a decision, not a 2am discovery.
_My recommendation: yes to both; a rollback snapshot nothing reads is not an interim tracking system._
**DECISION-3RM-03 ownership vs. parked PR #1023 (needs Mos).** **DECISION-2rollback artifact + availability trade. ⏸ JASON-PENDING — NOT BLOCKING.**
P-QUEUE-001 is absorbed into this mission, but PR #1023 is explicitly parked under Mos. Two lanes can The DB build is phases away, so this is queued for Jason's next session rather than escalated now.
legitimately claim it. A third recursion on the gate-6 defect — performed by the remediation itself — **Binding requirement in the meantime (Mos, from P-RECOVERY-001):** the DB spine **must NOT be a
would be the postmortem's own anti-pattern. _Not dispatching RM-03 until Mos rules._ single-point hard-stop.** Design for a broker-independent / degraded mode **plus** a rollback
artifact. Jason finalises only the specific availability target. This reverses my earlier reading of
OPUS D8 ("the fallback is: the fleet stops") — that answer is **not** pre-committed; a degraded mode
is now a design requirement on RM-12, RM-13, RM-23, RM-36 and RM-53.
**DECISION-3 — RM-03 vs. parked PR #1023. ✅ RULED: HOLD RM-03 (Mos, 2026-07-31).**
Do **not** open a third gate-6 lane — that is the postmortem's own anti-pattern performed by the
remediation. PR #1023 sits in Jason's **parked delivery stack**; its disposition (close, or supersede
by RM-03) is Jason's at his next session.
- **PR #1023`SUPERSEDED-PENDING-JASON`.** RM-03 stays `HOLD`; when Jason rules, RM-03 proceeds as
the single correct lane.
- **RM-02 and RM-55 proceed independently and are NOT held.** The per-merge-commit gate-assertion
requirement is the _conformance_ capability, not the gate-6 fix itself — different scope, no
ownership collision.
--- ---
## 6. Status ## 6. Status
| phase | state | | phase | state |
| ------------------ | ------------------------------------------------------------------------------------ | | ------------------ | ------------------------------------------------------------------------------------------------------------ |
| Decomposition | DONE — both planners delivered independently | | Decomposition | DONE — both planners delivered independently |
| Reconciliation | DONE — this document | | Reconciliation | DONE — this document |
| Blocking decisions | **OPEN — 3 escalated to Mos (§5)** | | Blocking decisions | **RULED** — all 3 closed by Mos 2026-07-31 (§5); D-2's availability target is Jason-pending but non-blocking |
| Dispatch | NOT STARTED — RM-01 is dispatchable now; it depends on nothing and blocks everything | | Dispatch | **RM-01 IN FLIGHT** — f10-coder (codex), worktree-isolated, AC1AC8 pre-registered |
| Review | PR #1025 with rev-974; ACs pre-registered 22:12:26Z before diff exposure |