Compare commits

..
Author SHA1 Message Date
mos-dt-0andClaude Opus 5 bb962dc25d docs(remediation): status — decisions ruled, RM-01 in flight
ci/woodpecker/pr/ci Pipeline was canceled
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:19:48 -05:00
mos-dt-0andClaude Opus 5 cee54db527 docs(remediation): apply Mos rulings on D-1/D-2/D-3; bank D-7 tmpfs ENOSPC
D-1 RULED — corrected wire-in target is a new production Node TaskExecutor on the
LIVE dispatch path (packages/mosaic launch + packages/coord) that Coord/Forge/live
dispatch submit through. The Python rail is deleted, not ported. RM-21/26/27/52 updated.

D-2 — Jason-pending and explicitly NON-blocking. Binding meanwhile: the DB spine must
not be a single-point hard-stop; degraded mode + rollback artifact become design
requirements on RM-12/13/23/36/53. This reverses the earlier recommendation to
pre-commit 'no DB means the fleet stops'.

D-3 — RM-03 HOLD; PR #1023 marked SUPERSEDED-PENDING-JASON; RM-02/RM-55 proceed
independently since conformance assertion is not the gate-6 fix itself.

RM-03 now carries BOTH queue-guard defects: wrong exit on unknown, and evaluating
branch=main rather than the branch being pushed.

D-7 — shared 30G /tmp hit 100% ENOSPC, degrading two seats. ~5.2G of it is session
scratch dead for 8-9 days; this session's own footprint is 88K. Same missing
capability as P-FLEET-001 stale-session GC, applied to disk. Folded into RM-50, and
RM-01 gains artifacts-off-shared-tmp acceptance criteria.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:19:30 -05:00
mos-dt-0andClaude Opus 5 b6d6abf125 docs(remediation): board — PRs #1025/#1026 open, review dispatched, D-6 queue-guard finding
ci/woodpecker/pr/ci Pipeline was canceled
Third independent live instance of the inert-gate class this session: the mandated
ci-queue-wait.sh returned exit 0 on state=unknown while guarding a push.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:13:39 -05:00
0e9e9246bf docs(remediation): reconciled execution backlog from two independent decompositions
ci/woodpecker/pr/ci Pipeline was canceled
TASK-1 complete. planner-opus (robustness, 38 tasks/8 dissents) and planner-sol
(pragmatic, 25 tasks/10 defers/7 dissents) each decomposed the 4-build plan without
seeing the other's work. Both decomps are committed alongside the reconciliation so
the disagreements stay auditable rather than being flattened into a consensus.

Reconciled into 58 tasks across P0-P5 (docs/remediation/TASKS.md):
- 7 independent convergences, treated as settled because neither planner could see
  the other. The headline: BOTH reject the charter's wire-in point
  (mosaic_orchestrator.py::run_single_task) because that controller is disabled and
  references a dispatcher absent from this checkout — wiring it would produce a
  stranded executor, the same built-but-unwired disease one layer up.
- 7 genuine disagreements ADJUDICATED, not averaged. The cost estimates are ~18x
  apart; rather than split the difference, the plan adopts sol's scope with opus's
  rigor and treats the first-dogfood gate as a hard budget checkpoint.
- 3 decisions escalated (wire-in point, rollback artifact + availability trade,
  queue-guard ownership vs parked PR #1023). No task blocked on them is dispatched.

Keystone dogfood case recorded (TASKS.md 1a): merged PR #868 shipped a file failing
pnpm format:check, then an unrelated PR reformatted it as a side effect, so main went
green and the gate's failure to fire left no trace. Detection must therefore be
per-merge-commit against that commit's own tree.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:09:58 -05:00
e80cebe7f0 docs(remediation): board reflects actual state + 4 dogfood findings from startup
Corrects the premature "DISPATCHED" planner line (per Mos): both planners are now
dispatched for real on GUARANTEED-clean context, not requested-clean.

Records Mos rulings: mission.json is retired-rail residue (do not invest); gate-16
holds on interim mos-dt-0 since rev-974 reviews; remote-control path is Mos-relay.

Captures 4 live failure classes observed while standing this seat up (D-1..D-4),
including an agent that silently ignored an in-message context reset — direct
evidence for the postmortem thesis that instructions are not enforcement.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:09:58 -05:00
21d5878515 docs(remediation): mission charter + kickstart + live board for the postmortem remediation
15/15 proposals decided (13 accept, 2 modify). Collapses to 4 builds + hygiene on one
PG spine + choke-point service. Durable mission record for the mos-remediation project
orchestrator; compaction-survival resume in KICKSTART.md.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Claude-Session: https://claude.ai/code/session_018XCrpMFmCAXbtSQtQAhDth
2026-07-31 17:09:58 -05:00
mos-dt-0andClaude Opus 5 1f4c7c0d37 fix(hygiene): .prettierignore must exclude Python build/test artifacts
ci/woodpecker/pr/ci Pipeline was successful
Prettier had no excludes for venv/__pycache__/.mypy_cache/.pytest_cache/htmlcov,
so any local Python virtualenv in the tree drops thousands of third-party files
into `pnpm format:check` and makes the gate unpassable in a working checkout
(observed: ~2400 files from an untracked apps/coordinator venv).

Same category as the existing node_modules/dist/.next entries. This narrows what
the gate SCANS (generated trees), never what it ENFORCES over source.

Found while making the three gates pass honestly for the remediation mission's
TASK-0, rather than reaching for --no-verify.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-07-31 17:04:41 -05:00
3 changed files with 39 additions and 113 deletions
-15
View File
@@ -1,15 +0,0 @@
<!-- board-roll: 1 entry rolled from BOARD.md -->
### **D-1 / P-ACTIVATION + hygiene — committed `.npmrc` hard-pins `store-dir=/root/.local/share/pnpm/store`.**
Correct for the CI container (runs as root), fatal for EVERY non-root local checkout: `EACCES` on `/root/.local/share/pnpm/store/v10/server/server.json`. A committed config that only works on one runtime is exactly the activation-skew class. Fix candidate: make store-dir env-overridable, not hardcoded.
<!-- board-roll: 2 entries rolled from BOARD.md -->
### **D-3 / P-FLEET-001 — the seats running this mission are UNMANAGED.** `mos-remediation`, `rev-974`,
`planner-opus`, `planner-sol` appear in NO roster (`~/.config/mosaic/fleet/roster.yaml`, `agents/`). Planners run on socket `default`; the roster declares `mosaic-fleet`. This is the exact "one roster-owned socket/host + quarantine unmanaged + stale GC" failure P-FLEET-001 indicts — observed on the remediation mission's own fleet. Prerequisite for INBOX identity-addressing.
### **D-2 / hygiene — husky `prepare` fails `EPERM` copying into root-owned `.husky/_/`.** Repo working
tree has root-owned dirs (`.husky/`, repo root) under a non-root agent. Worked around with the intended `HUSKY=0` escape hatch (does NOT disable the existing pre-commit/pre-push hooks).
+38 -35
View File
@@ -1,6 +1,6 @@
# mos-remediation — LIVE BOARD (keep < 8 KB) # mos-remediation — LIVE BOARD (keep < 8 KB)
**Phase:** EXECUTING — first PR merged; RM-01 in flight; all 3 decisions ruled. **Phase:** PLANNING COMPLETE — reconciled backlog landed; 3 decisions block dispatch.
**Updated:** 2026-07-31 (mos-remediation orchestrator; seat active on `mosaic-fleet`). **Updated:** 2026-07-31 (mos-remediation orchestrator; seat active on `mosaic-fleet`).
## Head ## Head
@@ -16,14 +16,14 @@
## In-flight ## In-flight
| Task | Owner | State | | Task | Owner | State |
| ------------------------------------------------- | --------------- | ------------------------------------------------------------------ | | ------------------------------------------------- | --------------- | ----------------------------------------------------- |
| PR #1026 docs (mission record + backlog) | mos-remediation | OPEN, retargeted to main, rebased; diff verified docs-only | | PR #1026 `remediation/mission-setup` (docs-only) | mos-remediation | OPEN, stacked on #1025; retarget to main after merge |
| PR #1025 hygiene | | **MERGED** 52414605; rev-974 APPROVE + CI #2158 8/8 terminal-green | | PR #1025 `fix/hygiene-inert-format-gate` | rev-974 | OPEN; ACs pre-registered 22:12:26Z; review DISPATCHED |
| DECISION-1 wire-in point (charter change) | Mos / Jason | ESCALATED — both planners reject the charter's target | | DECISION-1 wire-in point (charter change) | Mos / Jason | ESCALATED — both planners reject the charter's target |
| DECISION-2 rollback artifact + availability trade | Jason | ESCALATED | | DECISION-2 rollback artifact + availability trade | Jason | ESCALATED |
| DECISION-3 RM-03 vs parked PR #1023 ownership | Mos | ESCALATED | | DECISION-3 RM-03 vs parked PR #1023 ownership | Mos | ESCALATED |
| RM-01 reproducible checkout (unblocks everything) | unassigned | READY TO DISPATCH | | RM-01 reproducible checkout (unblocks everything) | unassigned | READY TO DISPATCH |
## Fleet seats ## Fleet seats
@@ -53,34 +53,37 @@
4. Hygiene + conformance harness 4. Hygiene + conformance harness
Cross-cutting retirements: flat-file tracking, 3 MACP islands, silent MOSAIC BYPASS. Cross-cutting retirements: flat-file tracking, 3 MACP islands, silent MOSAIC BYPASS.
## Dogfood evidence live failure classes, not hypotheticals ## Dogfood evidence captured this session (live failure classes, not hypotheticals)
> Newest first. Oldest entries roll to `BOARD-LEDGER.md` via `board-roll.sh` when this file - **D-1 / P-ACTIVATION + hygiene — committed `.npmrc` hard-pins `store-dir=/root/.local/share/pnpm/store`.**
> exceeds its 8 KB cap. Keystone detail is duplicated in `TASKS.md` §1a, so rolling loses nothing. Correct for the CI container (runs as root), fatal for EVERY non-root local checkout: `EACCES` on
`/root/.local/share/pnpm/store/v10/server/server.json`. A committed config that only works on one
runtime is exactly the activation-skew class. Fix candidate: make store-dir env-overridable, not hardcoded.
- **D-2 / hygiene — husky `prepare` fails `EPERM` copying into root-owned `.husky/_/`.** Repo working
tree has root-owned dirs (`.husky/`, repo root) under a non-root agent. Worked around with the
intended `HUSKY=0` escape hatch (does NOT disable the existing pre-commit/pre-push hooks).
- **D-3 / P-FLEET-001 — the seats running this mission are UNMANAGED.** `mos-remediation`, `rev-974`,
`planner-opus`, `planner-sol` appear in NO roster (`~/.config/mosaic/fleet/roster.yaml`, `agents/`).
Planners run on socket `default`; the roster declares `mosaic-fleet`. This is the exact
"one roster-owned socket/host + quarantine unmanaged + stale GC" failure P-FLEET-001 indicts —
observed on the remediation mission's own fleet. Prerequisite for INBOX identity-addressing.
- **D-4 / P-LIFECYCLE + hygiene — a dispatched agent silently IGNORED an in-message context reset.**
planner-sol was at 64.3%/372k; the brief asked it to reset first; it began work on dirty context anyway.
Only an out-of-band `/new` driven by the orchestrator guaranteed clean state. Confirms the postmortem
thesis: **instructions are not enforcement.** Reset must be a mechanical pre-dispatch step, not a request.
<!-- BOARD-ROLL:START --> - **D-6 / P-QUEUE-001 — the mandated queue guard returned PASS on an UNKNOWN state, live, today.**
Running the required `ci-queue-wait.sh --purpose push` before pushing produced
### **D-8 / P-CONFORMANCE-001 — a PRE-REGISTERED acceptance check that was not runnable as written.** `state=unknown ... exit 0` — the exact defect at `ci-queue-wait.sh:282-288` that PR #1023 is parked
on. It also evaluated `branch=main` rather than the branch being pushed. The mission's own required
PR #1025 AC2's fixture `mkdir -p apps/*/venv/lib` creates a literal `apps/*/venv/lib` dir when the glob is unmatched — it did not test what it claimed. rev-974 ran it exactly as written, caught it, re-ran the intended assertion at an explicit path, and **disclosed** rather than silently substituting a working fixture and reporting PASS. **Pre-registration protects a check from being retrofitted to the implementation; it does not make the check correct.** An unverified gate appeared inside the mechanism built to catch unverified gates. Hard requirement on RM-02: the registry must self-verify that every registered case runs AND can fail — presence is not evidence. pre-push gate passed me on an indeterminate result. Third independent live instance of the class.
- **D-5 / P-QUEUE-001 + P-CONFORMANCE-001 — KEYSTONE: an inert gate that erased its own evidence.**
### **D-7 / P-FLEET-001 — stale-GC-on-disk: shared 30G /tmp hit 100% ENOSPC, degrading two seats.** Merged PR #868 (`b79336a8`) shipped a file that FAILS `pnpm format:check` ⇒ the CI format gate did
not block. An unrelated later PR (#872) then reformatted that file via its own `lint-staged`, so
~5.2G was session scratch dead 8-9 days (this session's own footprint: 88K). Same missing capability as orphaned-tmux-session GC, applied to disk — not a quota or discipline problem. Resolved manually by Mos (lead coordinator) after independent verification; `/tmp` now 79%. **The gap IS the finding:** the authority to reap exists, the deterministic reaper does not. Folded into RM-50 with explicit requirements (mechanical liveness, age threshold, dry-run, audit event per reap — never a heuristic sweep). Refusing to unilaterally delete another session's scratch was correct doctrine; the fix is a reaper, not braver agents. `main` went green again and nobody learned the gate had failed to fire. Verified blob-level under the
repo's own config. **Detection must be per-merge-commit against that commit's own tree** — a "is main
### **D-6 / P-QUEUE-001 — the mandated queue guard returned PASS on an UNKNOWN state, live, today.** green today" check reports all-clear on this exact defect. Binding on RM-02/RM-55. Full chain in
`TASKS.md` §1a. NOT quiet-patched, by Mos's ruling: patching the symptom destroys the signal.
Running the required `ci-queue-wait.sh --purpose push` before pushing produced `state=unknown ... exit 0` — the exact defect at `ci-queue-wait.sh:282-288` that PR #1023 is parked on. It also evaluated `branch=main` rather than the branch being pushed. The mission's own required pre-push gate passed me on an indeterminate result. Third independent live instance of the class.
### **D-5 / P-QUEUE-001 + P-CONFORMANCE-001 — KEYSTONE: an inert gate that erased its own evidence.**
Merged PR #868 (`b79336a8`) shipped a file that FAILS `pnpm format:check` ⇒ the CI format gate did not block. An unrelated later PR (#872) then reformatted that file via its own `lint-staged`, so `main` went green again and nobody learned the gate had failed to fire. Verified blob-level under the repo's own config. **Detection must be per-merge-commit against that commit's own tree** — a "is main green today" check reports all-clear on this exact defect. Binding on RM-02/RM-55. Full chain in `TASKS.md` §1a. NOT quiet-patched, by Mos's ruling: patching the symptom destroys the signal.
### **D-4 / P-LIFECYCLE + hygiene — a dispatched agent silently IGNORED an in-message context reset.**
planner-sol was at 64.3%/372k; the brief asked it to reset first; it began work on dirty context anyway. Only an out-of-band `/new` driven by the orchestrator guaranteed clean state. Confirms the postmortem thesis: **instructions are not enforcement.** Reset must be a mechanical pre-dispatch step, not a request.
<!-- BOARD-ROLL:END -->
## Decisions log ## Decisions log
+1 -63
View File
@@ -93,47 +93,6 @@ and must not be cited as merge evidence. Rely on reviewer clearance + real CI.
Three independent live instances in a single session — format gate, agent context reset, queue guard — Three independent live instances in a single session — format gate, agent context reset, queue guard —
is the class confirmed, not anecdote. is the class confirmed, not anecdote.
### D-8 — a PRE-REGISTERED acceptance check that was not runnable as written
On PR #1025 the author (me) pre-registered AC2 with the fixture snippet `mkdir -p apps/*/venv/lib`.
In bash, when no `venv` exists the glob is unmatched and passes through literally, creating a
directory named `apps/*/venv/lib` rather than one per workspace. The check as written did not test
what it claimed to test.
`rev-974` ran it **exactly as written**, observed the wrong behaviour, then re-ran the intended
assertion at an explicit path — **and said so in the review** rather than silently substituting a
working fixture and reporting PASS.
Two things this establishes:
1. **The instruction "do not adjust a check to fit the diff; if it is unrunnable, say so explicitly"
worked.** A silent substitution here would have produced a green AC2 that proved nothing, on the
exact task whose subject is gates that appear to work. The disclosure is what made the PASS
meaningful.
2. **Pre-registration does not confer correctness.** A pre-registered check is protected from being
retrofitted to the implementation; it is not protected from being _wrong when written_. This is a
small instance of the mission's own class — an unverified gate — occurring inside the mechanism
built to catch unverified gates.
**Requirement on RM-02 (non-negotiable, sharpened by Mos).** The registry must **self-verify** that
every registered case demonstrably **runs** and demonstrably **fails on a known-bad input**.
Presence in the registry is **not** evidence. **A check is not trusted until it has been shown to
fail.** This is mutation testing / negative control applied _at the registry level_ — meaning
**the conformance harness must itself be conformance-tested.** A registered case that cannot fail, or
cannot run, is exactly as inert as an unregistered one, and the registry check must detect that
itself rather than assume it.
**Requirement on RM-55.** The same recursion applies to the harness: it must be observed red before
its green is worth anything (OPUS R-063 AC1 already states this; D-8 is the empirical case for it).
**Second, equally load-bearing lesson — reviewer disclosure is what makes a review trustworthy.**
rev-974 could have silently swapped in a working fixture and reported `AC2 PASS`. Nothing in the
process would have caught it, and the resulting green would have certified nothing — on the very task
whose subject is gates that only appear to work. The brief's instruction — _"do not adjust a check to
fit the diff; if it is genuinely unrunnable as specified, say so explicitly and explain why rather
than silently substituting your own"_ — is therefore not boilerplate. It is the clause that makes a
PASS mean something, and it must appear in **every** reviewer brief this mission issues.
### D-7 — shared-tmpfs contention → cascading ENOSPC (live incident, 2026-07-31) ### D-7 — shared-tmpfs contention → cascading ENOSPC (live incident, 2026-07-31)
The shared 30 G `/tmp` hit **100% ENOSPC** mid-session. It broke tool calls in **two different seats** The shared 30 G `/tmp` hit **100% ENOSPC** mid-session. It broke tool calls in **two different seats**
@@ -157,27 +116,6 @@ tmux-session GC.
**Added to RM-01 as acceptance criteria:** heavy build artifacts (node_modules, package stores, build **Added to RM-01 as acceptance criteria:** heavy build artifacts (node_modules, package stores, build
output) must land on the main disk in the worktree, never on the shared 30 G `/tmp`. output) must land on the main disk in the worktree, never on the shared 30 G `/tmp`.
**Resolution, and the part that is actually the finding.** Mos verified the attribution independently
(mtimes, no process or `lsof` holding either path, no live session maps) and reaped both as lead
coordinator: `/tmp` went to 79%, 6.0 G free. But note _how_ it was resolved — **a human-authority seat
did it by hand, because the authority exists and the reaper does not.** That gap is the finding, not
the disk usage.
Two doctrine points fall out, both binding on RM-50:
1. **The fix is not "agents should tidy up."** Asking each seat to clean its own scratch is
`instructions are not enforcement` (D-4) wearing a different hat. A deterministic reaper must own
it — same conclusion the north star reaches for every other class in this mission.
2. **Refusing to unilaterally delete another session's scratch was correct, and the resolution is not
"be braver about deleting."** An agent guessing that someone else's state is garbage is exactly the
unreviewed destructive act the Constitution forbids. The resolution is that _ownership and liveness
become mechanically decidable_, so reaping is a determination rather than a judgement call.
**Reaper requirements for RM-50:** liveness determined mechanically (process/`lsof`/session-map, not
mtime alone); an age threshold; a dry-run that reports what it would reap and why; and an audit event
per reap. Never a heuristic sweep — that would reintroduce the P-WORKFLOW-001 auto-sync failure in a
more destructive form.
**The self-erasure is the important part.** An inert gate that is masked by unrelated downstream **The self-erasure is the important part.** An inert gate that is masked by unrelated downstream
commits produces no lasting artifact, which is precisely why this class survives for months. Detection commits produces no lasting artifact, which is precisely why this class survives for months. Detection
cannot rely on "is `main` currently red" — it must be per-merge-commit. cannot rely on "is `main` currently red" — it must be per-merge-commit.
@@ -299,7 +237,7 @@ spread is itself information, and X1 says we calibrate on real merged PRs.
| id | task | src | depends_on | est (S/O) | tier | | id | task | src | depends_on | est (S/O) | tier |
| ----- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------- | -------------------------- | ---------------- | ------ | | ----- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------- | -------------------------- | ---------------- | ------ |
| RM-50 | One roster-owned socket/host; quarantine unmanaged; **deterministic reaper for stale sessions AND dead-session disk scratch** (D-7) | O+S+live | RM-04 | 14K / 150K | sonnet | | RM-50 | One roster-owned socket/host; quarantine unmanaged; stale-session GC | O+S | RM-04 | 14K / 150K | sonnet |
| RM-51 | Auto-sync **allowlist** (never auto-stage unknown paths) + worktree/lease isolation | O+S | RM-02 | 8K / 110K | sonnet | | RM-51 | Auto-sync **allowlist** (never auto-stage unknown paths) + worktree/lease isolation | O+S | RM-02 | 8K / 110K | sonnet |
| RM-52 | Retire the Python controller + duplicate MACP islands (3 → 1) | O+S | RM-26, RM-27, RM-25, RM-28 | 14K / 110K | codex | | RM-52 | Retire the Python controller + duplicate MACP islands (3 → 1) | O+S | RM-26, RM-27, RM-25, RM-28 | 14K / 110K | codex |
| RM-53 | Flat-file orchestration → DB hard cutover, with rehearsed rollback artifact | O+S | RM-27, RM-30, RM-34, RM-29 | (in S-10) / 200K | opus | | RM-53 | Flat-file orchestration → DB hard cutover, with rehearsed rollback artifact | O+S | RM-27, RM-30, RM-34, RM-29 | (in S-10) / 200K | opus |