docs(remediation): mission-state snapshot at the RM-01 seam (#1028)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful

Co-authored-by: mos-dt-0 <[email protected]>
This commit was merged in pull request #1028.
This commit is contained in:
2026-08-01 01:08:11 +00:00
committed by Mos
parent f58b3699a6
commit f65e9ea656
5 changed files with 656 additions and 40 deletions
+526 -11
View File
@@ -93,6 +93,520 @@ and must not be cited as merge evidence. Rely on reviewer clearance + real CI.
Three independent live instances in a single session — format gate, agent context reset, queue guard —
is the class confirmed, not anecdote.
### D-20 — the orchestrator's own documentation overclaimed, and a reviewer disproved it empirically
`rev-974` blocked PR #1027 a second time. **The defect was not in the code — it was in this file**, at
D-18's entry, written by the orchestrator.
Two faults, both mine:
1. **D-18's AC2 restatement omitted the scope clause** that D-19 later established as mandatory
("within an accidental/independent-mutation threat model").
2. **D-18 asserted that the tampered-manifest control turns integrity "from a claim into a property."**
It does not, and _cannot_. That sentence was written **before** D-19 proved the property impossible
at this layer, and was never revised when D-19 landed.
**The reviewer did not merely read it — it disproved it.** It performed a **same-UID consistent
manifest + marker rewrite**, and **preflight passed**. My documented claim was falsified by experiment.
Code, README, scratchpad and PR body all stated both threat-model directions correctly; **this file was
the only place still overclaiming.**
**This is two banked findings firing on the orchestrator at once:**
- **The integrity-claim corollary** — I wrote an integrity _claim_ in the voice of an integrity
_property_, in the very document that defines the rule against doing so.
- **D-14 (propagation)** — D-19 superseded D-18's assertion. I propagated the consequence into the
charter and the delivery conditions, but **not back into D-18 itself.** A ruling that fails to
propagate _backwards_ into the finding it supersedes is the same defect as one that fails to
propagate forwards, and I did not audit for that direction.
**Corrected in place**, with the original wording quoted and the empirical disproof recorded, rather
than silently rewritten — the same standard demanded of any restated criterion.
**Requirement on RM-02 (fifth clause).** Documentation asserting a _security or integrity_ property is
itself a claim requiring a negative control. Where a document states "X is guaranteed", the registry
must hold a case that **fails if X is not guaranteed** — and that case must have been observed red.
**Prose is not exempt from the mission's own evidentiary standard**, and prose in the _governing_
document least of all: it is the artifact most likely to be quoted as authority long after the code has
moved on.
**Reviewer credit.** rev-974 was briefed that its highest-priority check was "confirm the PR claims no
more than it can deliver, and a softened or omitted boundary is a finding even though the code works."
It applied that instruction **to the orchestrator's own governing document** and produced an experiment
to settle it. That is the review standard this mission is trying to make ordinary.
### D-19 — an integrity property that cannot exist at the layer it was specified
Implementing D-18's manifest, the seat + a Codex security review reached **CWE-345**: the symlink
manifest and the source-hash marker both live in the **same same-UID writable generated tree**, so an
actor with that UID can plant a rogue link, regenerate _both_, retain the fingerprint, and pass. **No
local cryptographic construction fixes self-authentication** without a key outside that actor's
authority; relocating the marker changes the path, not the authority.
The seat **escalated rather than describing self-authentication as tamper-resistant** — the explicit
failure mode the charter corollary demands. That is the corollary working, on its first real test.
**Ruling — Option A: scope AC2 to accidental / independent / stale mutation; retain the design.**
Rationale, recorded so it can be challenged:
1. **The undefendable boundary is not the weak link.** An actor with same-UID write can already edit the
source, the tests, `scripts/preflight.mjs` itself, and `.husky/*`. If they have that, _nothing_ in the
local checkout is trustworthy — hardening the manifest buys no real security while **implying
protection that does not exist**, which is worse than the gap.
2. **What AC2 is actually for.** These checks exist because a five-month-stale `.next` produced 19
phantom `TS2307` errors indistinguishable from real ones (**D-5**). That is staleness, drift and
foreign residue — and against that class the design demonstrably works.
3. **A real trust anchor arrives later, from this mission's own architecture.** An anchor must live
outside the actor's authority; for a fleet running as one user that means a separate service —
precisely the **choke-point executor + PG spine** of Builds 12, which verify outside the worktree's
authority. Hand-rolling key distribution for a local preflight now would duplicate that work badly.
4. Option C (structural policy, no manifest) is strictly worse — it cannot detect a **removed** expected link.
**Option A is acceptable only with honest labelling**, or it becomes the disease it is meant to cure.
Conditions (last two added/sharpened by Mos):
- Threat model stated verbatim in the code **and** the PR; the words _tamper-proof / tamper-evident /
secure_ **barred** from that context; the scope carried in AC2's restatement; every control kept
RED-first including manifest-only tamper.
- **State the boundary in BOTH directions.** Not only what it does _not_ defend (same-UID write; no
local construction can) but, beside it, what it **does** defend: accidental / independent / stale /
foreign-residue mutation — the **D-5** class it was born from (the five-month `.next` and its 19
phantom `TS2307`s). _A reader who sees only the negative dismisses the check as worthless; one who
sees only the positive over-trusts it. Both together is the honest artifact._
- **The residual risk is a HARD TRACKED DEPENDENCY EDGE, not a comment.** It is **RM-59**, owned by the
choke-point executor + spine work (`depends_on: RM-12, RM-21, RM-25`), and the AC2 scope note must
cite that id. _"Record where the real guarantee comes from" only holds if the record is a live
dependency someone must close._ **A documented gap with no owner becomes a permanent gap that reads
as intentional.**
**The generalizable rule.** When a required property **cannot exist at the layer where it was
specified**, the honest moves are: implement what the layer _can_ guarantee, **state the boundary
precisely**, and record where the real guarantee will come from. **A known gap that is written down is
acceptable; a gap that is implied fixed is not.** Silence here would have shipped a verification
artifact that verifies nothing — with a green to prove it.
### D-18 — two pre-registered criteria were mutually unsatisfiable, discoverable only at implementation
Implementing D-17's fix surfaced a conflict **between** pre-registered criteria:
- **AC2** (as written) — reject symlinked generated state.
- **AC4** — the canonical `pnpm -w build` succeeds and leaves no residue.
Verified independently rather than taken on report: `apps/web/next.config.ts:4` sets
`output: 'standalone'`, and the built tree contains **42 legitimate pnpm dependency symlinks** under
`.next/standalone/node_modules`. A blanket descendant-symlink rejection makes the canonical build fail
its own preflight with exit 43. **AC2 read literally is unsatisfiable alongside AC4 under this
configuration**, and nothing short of building the tree would have revealed it.
**Third distinct failure mode of a pre-registered check set**, completing the chain:
| finding | a pre-registered check set can be… |
| ------- | ------------------------------------------------------------------ |
| D-8 | **wrong** — a check that does not test what it claims |
| D-17 | **incomplete** — green while a criterion's requirement is untested |
| D-18 | **internally inconsistent** — two criteria that cannot both hold |
The implementing seat escalated instead of silently picking a winner. That matters: **quietly resolving
a conflict between pre-registered criteria destroys the point of pre-registering them** — the registration
exists so that changes of meaning are auditable rather than absorbed.
**Resolution (orchestrator ruling).** Approved a **build-certified symlink manifest**: `.next` itself is
still rejected as a symlink; descendants are rejected unless _exactly_ certified by a manifest the build
publishes atomically. Strictly **stronger** than blanket rejection — it also catches a **retargeted**
symlink, which blanket rejection cannot distinguish from a legitimate one.
**AC2 restated (recorded, not absorbed).** _Generated state must reject `.next` itself being a symlink
or non-directory, and must reject any descendant symlink not exactly certified by the build manifest —
added, removed, retargeted, or manifest-only-tampered all fail with exit 43 — **within an accidental /
independent-mutation threat model.**_
> ⚠ **This entry is superseded in part by [D-19](#d-19). Do not read D-18 standalone.** The scope clause
> above is load-bearing: the design **cannot** defend against a same-UID actor, which can rewrite the
> manifest and the marker consistently (CWE-345). D-18 was written **before** that impossibility was
> established.
**Hardening required before this counts.** The manifest is itself generated state, so **a manifest
writable by whoever plants a rogue symlink certifies the attack** — that is the one way this design
fails. It must sit inside the same ownership/fingerprint envelope, published atomically via the existing
marker mechanism, with negative controls **observed red first** for: added, removed, retargeted,
**manifest-only-tampered**, plus a positive control that the canonical build passes.
> ⚠ **CORRECTED (D-20).** This paragraph originally ended: _"without it, integrity is a claim rather
> than a property."_ **That overclaimed**, by implying the control makes integrity a _property_. It does
> not, and cannot. The manifest-only-tamper control detects **independent** mutation of the manifest;
> it confers **no authenticity** against an actor who rewrites manifest _and_ marker together.
> `rev-974` disproved the original wording empirically — a same-UID consistent manifest+marker rewrite
> **passed preflight**. Integrity here remains a scoped **drift-detection** property, never an
> authenticity one. See D-19 and the charter principle on properties that cannot exist at their
> specified layer.
**Requirement on RM-02 (fourth clause).** The registry must detect **conflicts between registered
criteria**, not only wrongness and coverage. Two criteria that cannot simultaneously hold is a registry
defect discoverable by construction — and when a criterion is restated, the registry must retain the
original text, the restatement, and the reason, so the evolution stays auditable.
### D-17 — a pre-registered criterion passed a green suite without being satisfied
`rev-974` returned **CHANGES REQUESTED** on PR #1027 with one blocking finding, and it is the sharpest
instance of the session's theme because it occurred **inside our own verification machinery**.
**AC2** was pre-registered before any code was written, and explicitly required that **symlinked
generated state be rejected**. The implementation does not do it:
```sh
ln -s /etc/hosts apps/web/.next/reviewer-symlink
pnpm preflight # → "checkout preflight passed", exit 0
# → required: generated-state exit 43
```
The acceptance suite was **21/21 green** throughout. Confirmed independently rather than relayed:
`scripts/preflight.mjs:82-92` rejects symlinks on the **source** path; `:28` merely _skips_ symlinked
directories rather than rejecting them; and the **generated-state** path at `:141-163` `lstat`s and
checks `uid` (ownership) but **never** calls `isSymbolicLink()`. The suite's only symlink cases
(`preflight.test.mjs:59`, `:115`) cover the turbo binary and a _source_ file. No generated-state case
exists anywhere.
**So: criterion pre-registered, suite green, requirement unmet.** Nobody was careless — the coverage gap
is _invisible from a green_, which is the entire problem.
**This sharpens D-8 rather than repeating it.** D-8 established that pre-registration does not confer
_correctness_ (a check can be wrong when written). D-17 establishes the adjacent failure:
**pre-registration does not confer _coverage_** — a suite can be green, and every registered criterion
can appear satisfied, while a criterion's actual requirement is untested. The two together mean a
registry of checks needs **two** properties, not one: each check must be _right_, and the set must
_actually exercise_ what it claims.
**Requirement on RM-02 (third clause).** The registry must bind each acceptance criterion to the
**specific case that exercises it**, and prove that case red before trusting its green. A criterion with
no case that can fail for _that criterion's stated reason_ is unregistered in substance however it
appears in the manifest. This is mutation testing pointed at the **criterion-to-case mapping**, not
merely at the gate.
**Credit where due:** the reviewer also declined to re-run AC8, stating plainly that the PR carried it
forward with no runnable command rather than silently substituting a different boundary test. That is
the D-8 clause working a second time, in the same review that produced D-17.
### D-16 — the local test gate and the CI test gate disagree by environment
Mos flagged a shape worth chasing: if `pnpm test` exits non-zero on a _pre-existing_ guard, then either
`main` is red and merges step around it (the #868 shape again), or CI does not run that path. **Both
branches turned out wrong, and the truth is a third thing.** Established by running it, not by asking:
CI runs **exactly** `pnpm test` (`.woodpecker/ci.yml`, `test` step) — the same command. So the path _is_
exercised. Yet:
| environment | result |
| ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| CI container | `test` step **green** (#2158, #2167) |
| this host, clean worktree | **exit 97**`WAKE-ASSERT INIT ABORT: BASH_LINENO convention violated on bash 5.2.15(1)-release … expected [3 4], probe reported [3 5] (#973)`, after `PASS=18 FAIL=0` |
| this host, main checkout | **exit 1** — a _different_, second defect (below) |
The guard is **environment-dependent**: it aborts on this host's bash and not in CI's container. `git
diff origin/main...` confirms PR #1027 touches **zero** files under `packages/mosaic`, so the guard is
genuinely pre-existing and unrelated — **f10-coder's report was accurate in every particular**, and
`main` is equally affected on this host.
**So it is not "merges step around a red" — it is worse in one specific way: the local gate and the CI
gate do not agree about what passing means.** No agent on this host can obtain a green `pnpm test` at
all, on any branch, including `main`. A gate an operator cannot run is a gate that only CI enforces,
and a gate only CI enforces cannot be a pre-push gate. This is the hermeticity/portability class
already live as #1007 (PR #1024).
**Second, independent defect found while establishing the above.** In the main checkout the same
package fails differently — exit 1 — because a test **scans the working tree** and asserts on files it
finds, picking up `apps/coordinator/venv/**` (third-party `site-packages`: `pi = math.pi` in `rich`,
`setuptools`, `mypy`). **A test whose result depends on untracked files present in the tree is not
hermetic.** This is the _same_ contamination source that made `pnpm format:check` unpassable (D-1/D-7
hygiene) — one untracked foreign tree silently breaking two independent gates.
**Requirements.** RM-01/RM-04: a gate must produce the same verdict on a developer host and in CI, or
declare loudly that it cannot run here — never diverge silently. RM-02 registers both as cases: the
environment-divergence guard, and a hermeticity control asserting a suite's verdict is unchanged by the
presence of untracked directories. Coordinate with #1007/#1024 rather than opening a third lane.
**Ownership (Mos, 2026-07-31).** The hermeticity fix **is** PR #1024, which sits in **Jason's parked
delivery stack** — so, like #1023, its disposition is Jason's. Marked `SUPERSEDED-PENDING-JASON`
alongside #1023. D-16 strengthens the urgency but does not transfer ownership: **we do not open a third
lane on a parked PR.** The one-line escalation for Jason: _two independent gates
(`format:check`, `pnpm test`) were broken by a single untracked directory, and a third
(`pnpm test`) disagrees between host and CI — non-hermetic gates make every green host-dependent._
**Sharpened statement of the class (Mos).** A pre-push gate an operator _cannot run locally_ is a gate
only CI enforces — so pointing `.husky/pre-push` at it **misrepresents where the gate lives**. Combined
with the shared root cause across two gates, the finding is: **non-hermetic gates make every green
host-dependent.** A gate that only appears to pass depending on which host runs it is this mission's
exact subject, one meta-level up.
**#1027 disposition (Mos):** proceeds on **CI-green**. CI is the authoritative gate; the local exit-97
is a known host-specific guard abort, irrelevant to the merge decision.
### D-15 — token scope is not repository permission (a THIRD capability layer)
`f10-coder` was provisioned with `gitea-mosaicstack-f10-coder.token`, scopes `write:repository` +
`write:issue`, and the mint was verified by "repo access returns 200". It then failed to push:
```
remote: error: User permission denied for writing.
remote: error: pre-receive hook declined
```
Verified objectively rather than inferred (per the charter principle):
| probe | result |
| ------------------------------------------------------ | ------------------------------------------- |
| `GET /repos/mosaicstack/stack/collaborators/f10-coder` | **404** — not a collaborator |
| repo permissions as seen by **its own token** | `admin: false`, `push: false`, `pull: true` |
**Capability has at least three independent layers, and satisfying two proves nothing about the third:**
1. **Token file exists** → raw-API authentication works (D-11b).
2. **`tea` login exists** → tea-dependent wrapper paths work (D-13).
3. **Repository permission granted** (collaborator/team membership) → _writes_ are actually authorised.
A token can carry `write:repository` scope and still be refused, because **scope bounds what a token
may attempt; repository permission decides what the user may do.** They are different systems.
**This is the charter principle failing on the very check meant to confirm capability.** The mint was
validated by an HTTP 200 on a _read_. A 200 proves reachability; it does not prove the property that
was required, which was **write**. Both the provisioner and I accepted it — the same
`written-unverified` treated as `verified` as D-12, one layer up, on a check whose entire purpose was
verification.
**Requirements.** RM-50's pre-dispatch capability check must probe the **effective permission for the
operation intended** — for push authority, assert `permissions.push == true` as that seat, not token
existence and not a 200 on a read. RM-04's registry reconciliation covers all three layers, with a
must-fail control for each. A capability check that cannot fail on a seat lacking write permission is
itself an inert gate.
### D-14 — a ruled decision did not propagate to the authoritative record
DECISION-1 (the corrected choke-point wire-in target) was ruled by the coordinator and applied to
`TASKS.md`. **`MISSION.md` — the charter, the document a cold-starting seat reads first — kept the
superseded target for hours.** It was flagged `CONTESTED` in a board note, then the ruling landed and
nobody edited the charter. A seat resuming from the charter would have read the _rejected_ target as
authoritative and wired the choke point into a disabled rail — the precise failure the ruling existed
to prevent.
Caught by hand, during an unrelated edit. Nothing would have caught it otherwise.
**This is P-MISSION-001 turned on ourselves.** The mission's own thesis is that convention exists and
_enforcement_ is the gap: a decision that lives in a chat ruling and a board note, but not in the
source of truth, has not actually been made — it has been _agreed_. The two are different, and the
difference is exactly what this mission is about.
> ⚠ **AMENDED by D-20 — propagation is BIDIRECTIONAL.** As first written, this requirement was read by
> both the orchestrator and the coordinator as _forward_ propagation only: a ruling reaches the
> documents that state the new rule. **D-20 proved that insufficient.** When D-19 superseded part of
> D-18, the consequence propagated forward into the charter and the delivery conditions but **never
> backward into D-18 itself**, which went on asserting a withdrawn claim — and a reviewer disproved it
> by experiment. **A supersession must update BOTH the documents that render the new rule AND the
> finding it retires, with the retired wording quoted rather than deleted.** Backward propagation is
> the same defect as forward; neither of us audited that direction until it bit.
**Requirement (not merely a fix).** A ruled decision must propagate **mechanically** to the
authoritative record; it must not depend on someone remembering to edit a second file. Concretely, once
mission state is DB-backed (RM-53 / the P-MISSION cutover):
- a decision is a **record**, not prose duplicated across documents;
- documents _render_ decisions rather than restating them, so there is one place to be wrong;
- and where duplication is unavoidable, a check asserts the authoritative record and the derived
document agree — with a must-fail control proving divergence is detected.
Until then, the interim rule: **the same commit that records a ruling updates every document that
states it — and every finding it supersedes.** Interim rules are exactly what the DB cutover exists to replace.
**The rule found a second instance within minutes of being written.** Auditing the charter against all
rulings to date (rather than waiting to be bitten again) surfaced that **DECISION-2 had also not
propagated**: `MISSION.md`'s standing directives still stated the DB hard-cutover with no mention of
Mos's binding qualification that _the spine must not be a single-point hard-stop_ (degraded mode +
rollback artifact required). A seat reading the charter would have designed toward an availability
posture the coordinator had explicitly rejected — and would have found the superseded
"no DB ⇒ the fleet stops" recommendation nowhere contradicted. Now corrected in place.
**Two un-propagated rulings out of two rulings that touched charter text.** The propagation gap is not
an oversight that happened once; without a mechanism it is the _default outcome_. That is the argument
for making this a requirement rather than a discipline.
### D-13 — two credential registries that can disagree (why the `--draft` fallback fired at all)
Diagnosing D-12's root cause surfaced a distinct defect. There are **two parallel credential
registries**, and capability in one does not imply capability in the other:
| registry | contents for identity `mos-dt-0` on `mosaicstack` |
| ------------------------------------------------------ | -------------------------------------------------------------------------------------------- |
| token files — `~/.config/mosaic/secrets/gitea-tokens/` | `gitea-mosaicstack-mos-dt-0.token` **EXISTS** |
| `tea login list` | **NO** `mosaicstack` login for `mos-dt-0` (only `mosaicstack-mos` and `mosaicstack-rev-974`) |
So `get_gitea_token` succeeds and every raw-API path works, while every **tea-dependent** wrapper path
fails its login validation and silently degrades to the API fallback — which is exactly what dropped
`--draft`. **tea is not "stale"; the login simply does not exist for that identity.**
This matters beyond one flag: capability was declared authoritative by the token-file set (D-11b), but
that registry does not govern the tea path. A seat can be _fully provisioned_ by the authoritative
registry and still lose functionality with no error — only a warning, and only on the degraded path.
**Requirements.** RM-04 (activation coherence): the two registries must be reconciled — one source of
truth, or a startup check asserting they agree, with a must-fail control proving disagreement is
detected. RM-50: the pre-dispatch capability check must verify capability for the **path actually
used**, not merely token-file presence.
**Confirmed working despite the gap** (so this is degradation, not outage): pushes, `pr-merge.sh`,
PR/issue creation via API fallback, comment posting, and all reads. Impact is confined to
tea-only features — `--draft`, `--labels`, `--milestone`.
**Reconciliation run by Mos (the manual form of RM-04's assert-agreement, done once by hand).** For
`git.mosaicstack.dev`, the token-file registry holds **six** seats; `tea` holds logins for **two**:
| state | seats |
| ------------------------------------------------ | -------------------------------------------------------- |
| token file present, **no** mosaicstack tea login | `f10-coder`, `jarvis`, `mos-admin`, `mos-dt-0`, `pepper` |
| token file present **and** tea login present | `rev-974` (only) |
**Five of six provisioned seats are silently degraded on tea-only features.** This is _systemic_, not
a one-off — which is why the fix is registry reconciliation (RM-04) and not a per-seat mint. Minting
one seat would clear a symptom and leave the class live.
Mos deliberately deferred the mint: it is not on RM-01's critical path, and additively editing shared
`tea` config underneath running work is a change he declined to make without cause. Full remediation —
mint the five missing logins **and** wire the startup must-fail assertion that _detects_ disagreement —
lands as RM-04 at a non-critical seam, or immediately if any seat needs a tea-only feature to progress.
**Correction of record:** this supersedes D-11(b)'s claim that the token-file set is _the_ authoritative
capability registry. It is **necessary but not sufficient**. Capability is **per-path**: the token file
governs the raw-API path, the tea login governs the tea path, and the two can disagree silently.
### D-12 — a requested SAFETY flag was silently degraded, and I did not check
I created PR #1027 with `pr-create.sh ... -d` (draft) because it carries **partial, unproven work**.
`tea` authentication was stale, so the wrapper fell back to its raw-API path — which cannot set draft —
and emitted:
```
Warning: API fallback applies title/body/head/base only; labels/milestone/draft require authenticated tea setup.
```
The PR was created **not-draft**. I read the success output, saw the PR number, and moved on. I then
reported to the coordinator that the PR was "opened as draft". **It was open, mergeable, and marked
ready for ~25 minutes**, protected only by the words "DRAFT" and "do not merge" in its title and body —
i.e. by prose a human might read, not by the platform control I asked for. Detected only because a
watcher polled `draft:` and the value disagreed with my belief. Corrected by setting the `WIP:` title
prefix (Gitea's draft mechanism); `draft: True` verified after.
**Three distinct failures, and the third is mine:**
1. **Silent degradation of a safety flag.** The fallback path dropped `--draft` and still exited 0. A
fallback that cannot honour a _safety_ argument must fail, not proceed — degrading `--labels` is a
nuisance; degrading `--draft` publishes unproven work as ready to merge.
2. **The warning went to stderr and nothing consumed it.** It was correct, specific, and ignored — a
warning nobody acts on is indistinguishable from no warning.
3. **I did not verify the flag took effect.** I checked that the PR existed, not that it had the
property I required. This is the mission's own thesis turned on me: **I trusted a success exit code
over an observed state**, on exactly the class of tool this mission exists to distrust.
**Requirements.** RM-02: a wrapper that cannot honour a safety-relevant argument must exit non-zero —
registered with a must-fail control asserting `--draft` on a degraded path fails rather than proceeds.
RM-24 (tri-state write outcomes): this is precisely `written-unverified` being treated as `verified`
the PR write succeeded, the _requested property_ was never confirmed, and no one looked.
### D-11 — seat identity did not survive into git, and seat capability is invisible at dispatch
Two defects, one dispatch (RM-01 → `f10-coder`):
**(a) Identity drift — P-WRAPPER-001, reproduced on our own delivery.** The seat's commits are
authored `mosaic-coder <[email protected]>` — the generic fallback. **You cannot tell from
git history which seat did this work.** Recorded, not rewritten: the drift is the evidence.
> **Mechanism, corrected (Mos).** My original framing here was wrong, and the error was in the brief
> before it was in the finding. `MOSAIC_GIT_IDENTITY` resolves the **token** (which per-slot credential
> the wrappers act with). The **commit author** comes from `git config user.name` / `user.email`, which
> is a **separate setting** — it fell back to the generic value because nothing set it. Exporting the
> identity could never have fixed authorship. **My worker brief instructed only the export, so the
> seat did exactly what it was told and the commits were still mis-attributed.**
>
> **The requirement is coherence: token and authorship must agree.** A seat acting with
> `gitea-mosaicstack-f10-coder` must also commit as `f10-coder <[email protected]>`.
> Either half alone is identity drift — one produces the right credential with the wrong author, the
> other the reverse. That coherence _is_ P-WRAPPER-001, and it belongs in seat setup, not in prose
> instructions a seat may follow correctly and still end up wrong.
**(b) Capability opacity.** Nothing at dispatch time revealed that `f10-coder` had no credential for
the target provider. Per-slot tokens live at `~/.config/mosaic/secrets/gitea-tokens/`; the seat holds
`gitea-usc-f10-coder` but not `gitea-mosaicstack-f10-coder`. This surfaced only when the seat failed
**mid-task, after ~$9 and 69% of its context.** The orchestrator (me) selected a seat without any way
to check it could act on the target repo — and there was no way to check.
`get_gitea_token` behaved **correctly**: it refused to fall through and borrow another slot's token,
failing loud precisely to protect gate-16 attribution. The tooling was right; the _dispatch-time
information_ did not exist.
**This is P-RECOVERY-001's "honest capability labeling" applied to seats rather than services.** A seat
should declare what it can actually do — which providers, which repos, which credentials — and that
declaration must be **checkable before dispatch**, not discovered by failure after the budget is spent.
**Requirements.**
- **RM-04 (activation coherence)** gains the identity-binding half: seat setup must set **both** the
token identity **and** `git config user.name`/`user.email`, coherently. Verified by an
exit-asserting test that makes a commit and asserts its author — never assumed from an instruction
in a brief.
- **RM-50 (roster ownership)** gains per-seat capability declaration plus a **pre-dispatch capability
check**. Mos (who owns provisioning) confirms the check is mechanically trivial: **capability is
token-file existence.** Before dispatching seat `X` to provider `Y`, test that
`~/.config/mosaic/secrets/gitea-tokens/gitea-<Y>-<X>.token` exists; if absent, provision it or pick a
provisioned seat. **The token-file set is the authoritative capability registry.** A one-second check
would have replaced a mid-task failure that cost ~$9 and 69% of a seat's context.
### D-10 — the queue guard's failure modes are exactly backwards
`ci-queue-wait.sh` — a **required** pre-push/pre-merge gate — was observed this session doing both of
these:
- **Fails OPEN on an unknown result.** `state=unknown ⇒ exit 0`, five times, during real pushes and
real merges. It also evaluates `branch=main` rather than the branch being acted on.
- **Fails CLOSED on credential resolution.** In a worker seat it aborted with
`Gitea token not found`, hard-blocking a legitimate push of completed, tested work. The worker
correctly stopped (Constitution gate 8). The identical command run from that worker's _own worktree_
in another shell succeeded, so the checkout and remote were fine — the difference was the worker's
process environment.
**A gate that waves through work it never checked, and blocks work that is ready, has its failure
modes inverted.** Availability failures (cannot reach the provider, cannot resolve a credential)
should degrade to a loud, auditable _inability to assert_ — never to a hard stop on delivery, and
never to a silent pass. Correctness failures (unknown, malformed, terminal-failure) are what must
block.
This is also the **Pi-brick shape** (P-RECOVERY-001): a gate whose own unavailability prevents the
work needed to recover from it.
**Requirement on RM-03, extending its existing two defects:** the guard must distinguish
`CANNOT_ASSERT` (credential/transport/provider unavailable — loud, audited, does not silently pass and
does not permanently block) from `ASSERTED_NOT_READY` (a real non-green CI state — blocks). Both are
registered R-002 cases with must-fail controls; neither may exit 0 silently.
### D-9 — the comms path shell-interprets message bodies (injection-shaped, found by accident)
Sending a status message with `agent-send.sh -m "...`backticks`..."` caused bash to **execute** the
backticked text as command substitution. The recipient received a mangled body plus a
`No such file or directory` error; the intended sentence never arrived. The message was reported as
delivered.
This is the **same class** as the already-noted `pr-create.sh` backtick-quoting bug (M2 scratchpad):
**two tools in the comms path treat a message body as shell input.** A body that can execute on the
sender is a _correctness_ bug before it is ever a security one — and note the failure mode: the
send reported success while silently transmitting something other than what was written. Silent
corruption with a success receipt is precisely the pattern this mission exists to eliminate.
**Requirement on RM-40 / RM-42 (comms/v1), hardened by Mos.** The envelope must carry its payload
**verbatim** and must not be subject to shell interpretation at **any** hop — sender, transport, or
adapter. Concretely: **file/stdin transport, never argv interpolation.**
**Standing interim rule, effective now (Mos).** Until the envelope lands, use `agent-send.sh -f
<file>` for any message body containing special characters — **never `-m`**. Passing a file sidesteps
argv interpolation entirely. **This rule is mandatory in every worker brief this mission issues**,
alongside the D-8 "if a check is unrunnable, say so" clause. Round-trip fidelity (send a body containing backticks, `$(…)`, quotes, and newlines; assert
byte-identical receipt) is a required registered test case under RM-02, including a must-fail control
proving the assertion can detect corruption.
### D-8 — a PRE-REGISTERED acceptance check that was not runnable as written
On PR #1025 the author (me) pre-registered AC2 with the fixture snippet `mkdir -p apps/*/venv/lib`.
@@ -297,17 +811,18 @@ spread is itself information, and X1 says we calibrate on real merged PRs.
### P5 — Retirements, hygiene, conformance
| id | task | src | depends_on | est (S/O) | tier |
| ----- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------- | -------------------------- | ---------------- | ------ |
| RM-50 | One roster-owned socket/host; quarantine unmanaged; **deterministic reaper for stale sessions AND dead-session disk scratch** (D-7) | O+S+live | RM-04 | 14K / 150K | sonnet |
| RM-51 | Auto-sync **allowlist** (never auto-stage unknown paths) + worktree/lease isolation | O+S | RM-02 | 8K / 110K | sonnet |
| RM-52 | Retire the Python controller + duplicate MACP islands (3 → 1) | O+S | RM-26, RM-27, RM-25, RM-28 | 14K / 110K | codex |
| RM-53 | Flat-file orchestration → DB hard cutover, with rehearsed rollback artifact | O+S | RM-27, RM-30, RM-34, RM-29 | (in S-10) / 200K | opus |
| RM-54 | Fleet-wide inert-gate audit against the RM-02 registry | O | RM-02 | — / 120K | sonnet |
| RM-55 | **Conformance harness:** fault-inject the live failure classes on real artifacts | O+S | RM-35, RM-41, RM-53 | 18K / 260K | opus |
| RM-56 | Retirement proof: CI asserts all three retirements are complete **and stay complete** | O | RM-52, RM-45, RM-53 | — / 90K | codex |
| RM-57 | Operator cutover docs + activation proof; map all 15 decisions to evidence | S | RM-04, RM-36, RM-45, RM-55 | 6K / — | codex |
| RM-58 | **Mechanical pre-dispatch context reset** the orchestrator resets a seat out-of-band and verifies it, rather than asking the agent to reset itself | mos-remediation (D-4) | RM-31, RM-50 | 8K | sonnet |
| id | task | src | depends_on | est (S/O) | tier |
| ----- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------- | -------------------------- | ---------------- | ------ |
| RM-50 | One roster-owned socket/host; quarantine unmanaged; **deterministic reaper for stale sessions AND dead-session disk scratch** (D-7) | O+S+live | RM-04 | 14K / 150K | sonnet |
| RM-51 | Auto-sync **allowlist** (never auto-stage unknown paths) + worktree/lease isolation | O+S | RM-02 | 8K / 110K | sonnet |
| RM-52 | Retire the Python controller + duplicate MACP islands (3 → 1) | O+S | RM-26, RM-27, RM-25, RM-28 | 14K / 110K | codex |
| RM-53 | Flat-file orchestration → DB hard cutover, with rehearsed rollback artifact | O+S | RM-27, RM-30, RM-34, RM-29 | (in S-10) / 200K | opus |
| RM-54 | Fleet-wide inert-gate audit against the RM-02 registry | O | RM-02 | — / 120K | sonnet |
| RM-55 | **Conformance harness:** fault-inject the live failure classes on real artifacts | O+S | RM-35, RM-41, RM-53 | 18K / 260K | opus |
| RM-56 | Retirement proof: CI asserts all three retirements are complete **and stay complete** | O | RM-52, RM-45, RM-53 | — / 90K | codex |
| RM-57 | Operator cutover docs + activation proof; map all 15 decisions to evidence | S | RM-04, RM-36, RM-45, RM-55 | 6K / — | codex |
| RM-59 | **Close the D-19 residual risk** — generated-state verification anchored **outside** the worktree's authority (executor/spine-side attestation), retiring the same-UID self-authentication gap | mos-remediation (D-19) | RM-12, RM-21, RM-25 | 20K | opus |
| RM-58 | **Mechanical pre-dispatch context reset** — the orchestrator resets a seat out-of-band and verifies it, rather than asking the agent to reset itself | mos-remediation (D-4) | RM-31, RM-50 | 8K | sonnet |
**Critical path:** `RM-01 → RM-02 → RM-10 → RM-11 → RM-12 → RM-21 → RM-23 → RM-31 → RM-33 → RM-34 → RM-53 → RM-55`.