merge: absorb main into next — 23-commit divergence (08-05..13 base=main window)
ci/woodpecker/pr/ci Pipeline failed

17 content commits + 3 merge bubbles were genuinely missing from next (~8,000 lines:
goal controller #1152, framework enforcement #1174/#1195, pr-edit wrapper #1173/#1200,
pipefail series #1100/#1105/#1106/#1107, git-tools fixes #1073/#1085/#1086/#1089,
#991, #1007, enrollment tolerance #1094). 3 commits were already in next by content
(#1060 identical, #1066/#1062 evolved twins — conflicts resolved to next's side).

Per-commit classification and evidence: mosaic-brain fleet/lanes/stack-remediation/main-next-divergence.md.
Conflict resolutions (6 files) itemized in the PR body.
This commit is contained in:
2026-08-19 16:27:17 -05:00
105 changed files with 7967 additions and 208 deletions
@@ -60,6 +60,52 @@ If a repo does not expose these scripts, run equivalent local workflow commands
- Do not auto-resolve data conflicts in shared state files.
- Keep commits scoped to a single logical change set.
## Model Tiering
Model choice is a standard, not a preference. Delegating a mechanical grep to a
frontier reasoning model wastes budget; sending a security review to a cheap tier
produces a review that passes and proves nothing. Both are defects.
Tiers are named by **capability class**, so the standard survives a model
generation. An operator binds each class to a concrete model id.
| Class | Use for |
| ------------- | ----------------------------------------------------------------------------------------- |
| `search` | grep/glob, file location, status and health checks, one-line mechanical edits |
| `build` | feature implementation, test writing, bugfixes, routine refactors |
| `judge` | code review, planning, API/compat-sensitive changes |
| `adversarial` | security review, ambiguous architecture, anything where a wrong "looks fine" is expensive |
Rules:
1. **Start at the cheapest class that can do the task; escalate on evidence, not
on nerves.** Omitting a tier is not neutral — it inherits the caller's model,
which is usually the most expensive one.
2. **Compat-sensitive work escalates one class.** A change that must interoperate
with an existing contract is judged, not just built.
3. **A tier assignment is benchmarked, not asserted.** Move a task class to a
cheaper tier only against a blind A/B on real work from this codebase, ranked
by someone other than the author. "It seemed fine" is not evidence.
4. **Reviewer independence beats reviewer size.** An `adversarial` verdict from
the model that wrote the code is not a second opinion (see Constitution gate 16).
### Where the binding lives
The class→model map is operator configuration, never framework source: model
availability, cost, and quotas differ per operator and per host.
Resolution order, first hit wins:
1. the config service (DB-backed, surfaced and editable in the Mosaic webUI)
2. a local operator file (`STANDARDS.local.md`, or `policy/` where the runtime
injects it)
3. the framework default — the class names above, with no binding
Only layer 1 is auditable across a fleet, so it is the target end state; layers 2
and 3 exist so a host with no config service still runs. A local override that
silently disagrees with the config service is drift — the same failure class the
tool-index gate exists to catch, and it belongs in `mosaic doctor`.
## Prompting Contract
All runtime adapters should inject: