Compare commits

..
Author SHA1 Message Date
fargo 2097379e25 fix(macp): fail-closed typed gate states — no placeholder/empty-command passes (#1275)
ci/woodpecker/pr/ci Pipeline was successful
RI-N2 / SDLC-D-035. Every GateResult now carries a typed status
discriminator (passed|failed|simulated|waiting|capability_failure);
passed:true remains true only for really-executed, really-green gates.

- ci-pipeline without a CI provider → capability_failure (MACP_NO_CI_PIPELINE),
  never the placeholder pass
- empty-command gate → capability_failure (MACP_NO_COMMAND / MACP_NO_REVIEWER),
  and runGates no longer silently skips it
- manual gate type waits (MACP_AUTHORITY_REQUIRED) — neither pass nor fail
- explicit simulation only (simulate option / --simulate): typed simulated
  results can never make the aggregate passed (RunGatesResult.state)
- CLI stubs (tasks list, submit, events tail) exit nonzero with typed
  MACP_NOT_IMPLEMENTED; macp gate is now wired to runGates with --simulate
- legacy __tests__/gate-runner.test.ts consolidated into src/gate-runner.spec.ts
  (root eslint project service does not cover packages/macp/__tests__)
2026-08-17 16:40:16 -05:00
13 changed files with 964 additions and 629 deletions
+2 -2
View File
@@ -138,9 +138,9 @@ mosaic brain tasks
mosaic brain conversations
# Agent forge pipeline
mosaic forge run [--simulate] # fails closed (FORGE_NO_EXECUTOR) with no executor wired; --simulate for typed simulated runs
mosaic forge run
mosaic forge status
mosaic forge resume [--simulate] # same fail-closed rule as forge run
mosaic forge resume
mosaic forge personas
# Structured logging
-35
View File
@@ -1368,38 +1368,3 @@ All work is **alpha** (< 0.1.0) until Jason approves 0.1.0 beta release.
10. ASSUMPTION: **Conversations and messages get their own PG tables** (not stored in brain's entity model). They follow a chat-specific schema with proper foreign keys to users and projects. Rationale: Chat has different access patterns (streaming, pagination, search) than brain entities.
11. RESOLVED: **Pi handles all target LLM providers natively.** Anthropic, OpenAI/Codex, Z.ai, Ollama, LM Studio, and llama.cpp are all supported via Pi's built-in providers or `models.json` configuration with `openai-completions` API type. No custom provider adapters needed in @mosaicstack/agent — only configuration management.
---
## Release Integrity Workstream (RI, #1275)
### Problem and objective
At `next` 476db12b (review of 2026-08-17), publication from `next` is not bound to the full verification pipeline for the same commit: the publish pipeline's publish steps depend on `build` only, while ordinary push CI excludes `next`. Public Forge/MACP paths contain false-success placeholders: a stub executor that reports `completed` with exit zero, planning/remediation gates that execute literal `true`, a review gate that echoes an approving verdict, and a gate runner that treats empty commands and unimplemented CI-provider gates as passing. Shipping UI surfaces can render a failed fetch as an empty, healthy collection.
Objective: for alpha 0.0.50, the release cannot publish, report, or display work state that the repository has not actually verified. Decisions SDLC-D-033 through SDLC-D-038 (Jason, 2026-08-17) scope this floor; full decision text and required-behavior lists live in jarvis-brain `docs/plans/2026-08-16_mosaic-stack-sdlc-protocol.md` and `data/decisions/mosaic-stack-sdlc-protocol.json`. This section restates only the normative requirements.
### Normative requirements
1. **RI-N1 Exact-commit publication verification (SDLC-D-034).** One canonical terminal verification command performs self-contained re-verification in the publish pipeline against the job's checked-out commit before any external publication effect. The command contains or invokes the complete mandatory verification set (semantic parity with the PR merge gate, including sanitization, upgrade-guard, typecheck, lint, format check, tests, and build); CI and publication do not maintain separate semantic checklists. Every publish step depends on the verification step in the executable pipeline DAG. Provider commit identity and `git rev-parse HEAD` must identify the same commit. Missing, skipped, cancelled, stale, or inconclusive checks fail closed. Documentation-only runs may skip publication but cannot bypass verification when a publication effect will occur. A negative control must prove that a broken check blocks every publish step.
2. **RI-N2 Fail-closed Forge/MACP with explicit simulation (SDLC-D-035).** Simulation requires explicit caller intent (e.g. `--simulate`) and produces a distinct typed `simulated` state that can never satisfy dependencies, acceptance criteria, gates, merge, or release. Normal execution exits nonzero with a typed capability failure when a required executor, reviewer, command, or CI provider is absent — no stub completion, no literal-`true` gates, no synthetic approvals, no empty-command passes. A manual gate with no automation enters a waiting state; it does not pass. Positive tests prove explicit simulation still works; negative controls prove simulation and every missing-provider case cannot advance lifecycle state.
3. **RI-N3 One transitional PRD authority (SDLC-D-036).** `@mosaicstack/prdy` structured storage under `docs/prdy/`, driven by `mosaic mission --plan`, is the authoritative PRD representation for the alpha. `mosaic prdy` either routes through the same application service or operates only as an explicit, named Markdown import/export adapter; `docs/PRD.md` is not a peer authority. `mission --plan` must persist the mission↔PRD linkage (mission id/version, PRD id/version, selected requirements). Markdown output is a generated view carrying source identity; editing it cannot mutate authority silently. Import is explicit, validated, and conflict-aware (proposed successor, never overwrite). Structural validity is separate from approval.
4. **RI-N4 One quality-rails evaluator (SDLC-D-037).** The TypeScript quality-rails package is the sole authoritative evaluator. A complete probe inventory maps every current TypeScript and shell check to one canonical check with disposition (preserve/strengthen/retire, each named). Effective shell enforcement probes are absorbed before their independent paths retire; expected-file presence alone is not parity. The evaluator returns typed results (`passed`/`failed`/`blocked`/`error`/`not-applicable`) with check version, subject, and reason; missing implementation, missing input, unknown check, process error, timeout, or malformed output can never become `passed` or an unqualified skip. Check definitions and policy are versioned and digested. Shell commands become thin adapters with no separate verdict logic. The canonical terminal verification command (RI-N1) invokes this evaluator rather than duplicating its logic. Contract, parity, and negative-control tests are required, plus independent review of probe equivalence.
5. **RI-N5 Consequence-aware stale UI (SDLC-D-038).** Mission Control distinguishes typed freshness states (`current`, `stale`, `partial`, `unknown`, `unavailable`) rather than inferring from empty arrays or null. A failed fetch never renders as an empty healthy collection. Last-known data may display for situational awareness only with source identity, version, and age visibly labeled; any derived completion/assurance/release verdict whose inputs are stale becomes `unknown`; all state-changing actions are disabled until fresh state loads and is revalidated. With no verified snapshot, surfaces show an explicit unavailable state. Cache corruption, cross-workspace data, schema mismatch, and version regression invalidate the snapshot. Tests cover the failure matrix (network, auth, malformed, partial, corruption, stale age, schema mismatch, recovery, stale-action rejection) with negative controls proving no case yields a current green verdict or enabled mutation.
### Acceptance criteria
- AC-RI-1: A push to `next` that fails any mandatory verification step publishes nothing (no npm package, no image), demonstrated by a checked-in negative control and by pipeline evidence on a real `next` publish run where the verification step is green and every publish step depends on it.
- AC-RI-2: With no executor/reviewer/CI provider wired, Forge and MACP normal runs exit nonzero with typed capability failures; with `--simulate`, runs complete but every result is typed `simulated` and cannot satisfy any gate, dependency, or completion state — proven by unit tests including negative controls.
- AC-RI-3: A PRD created or revised through either `mosaic mission --plan` or `mosaic prdy` resolves to one authority under `docs/prdy/` with stable identities and versions; the mission↔PRD linkage survives restart; a Markdown export is labeled as generated and cannot silently become a second writer; divergent legacy content blocks baseline claims until explicitly resolved — proven by contract tests.
- AC-RI-4: `quality-rails check` through any entry point (TS CLI, framework shell adapter) returns the same typed verdict for the same subject; the probe inventory names every legacy check's disposition; a deliberately broken probe fails closed — proven by contract/parity/negative-control tests and independent review of probe equivalence.
- AC-RI-5: No shipping surface renders a failed fetch as an empty healthy state; stale/partial/unavailable states are typed, labeled, and mutation-disabled — proven by the failure-matrix tests.
- AC-RI-6: All cards merged to `next` via squash PR with terminal-green CI; release evidence for 0.0.50 records commit, verification run, and published artifacts.
### Out of scope
The canonical dispatcher/control-plane vertical slice (work graph, execution attempts, fenced leases, typed check-in, independent verifier dispatch) is decided post-alpha (SDLC-D-033, option B). Multi-pipeline verification certificates (SDLC-D-034 option B) are post-alpha. Full AF-1..AF-4 objective matrices and Mission Control portfolio surfaces are post-alpha.
-42
View File
@@ -1,42 +0,0 @@
# Tasks — Release Integrity Workstream (RI-050, #1275)
> Single-writer: the RI-050 orchestrator (jarvis, dragon-lin) only. Workers read but never modify.
>
> **Mission:** alpha 0.0.50 release-integrity floor (decisions SDLC-D-033..038).
> **PRD:** [docs/PRD.md § Release Integrity Workstream](../PRD.md#release-integrity-workstream-ri-1275)
> **Issue:** #1275 (remains open until RI-V-001 closes)
> **Base branch:** `next` (all cards branch from `origin/next`, squash-merge via PR)
>
> **Execution note:** the `agent` column uses `pi-glm-5.3` — outside the pipeline-cron model
> table on purpose. This workstream is executed by jarvis on dragon-lin with local pi workers
> (`pi --model zai/glm-5.3:high`); pipeline crons must not auto-claim these rows.
>
> **Status values:** `not-started` | `in-progress` | `done` | `blocked` | `failed` | `needs-qa`
> `done` requires: repo quality gates green, independent review recorded, terminal-green CI on
> the PR head, squash merge to `next`, and acceptance evidence in notes.
| id | status | description | issue | agent | repo | branch | depends_on | estimate | notes |
| -------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----- | ---------- | ----------------- | --------------------------------- | ---------------------------------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| RI-0-001 | in-progress | Bootstrap: issue #1275, PRD section, this DAG, scratchpad (docs only) | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-mission-bootstrap | — | 6K | PR #1276 (head 758659dd): docs-only, CI green (2475). Review requested from fargo. Merges first (no publish run). |
| RI-1-001 | in-progress | RI-N1: canonical terminal verification command + publish-pipeline exact-commit gate (every publish step depends on verify; commit identity check; fail closed) | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-publish-gate | RI-0-001 | 25K | PR #1277 (head 46784c8d): CI GREEN at head after serialized retry (pipeline 2476, 2026-08-18) - earlier red was CI-agent contention (web SPA timeouts under concurrent pipelines), not code. Review requested from fargo at pinned head (comms 20260818T021025Z). |
| RI-1-002 | not-started | RI-N1 negative control: checked-in tests proving a broken mandatory check blocks every publish step and that DAG edges cannot be bypassed | #1275 | pi-glm-5.3 | mosaicstack/stack | test/ri-050-publish-gate-negative | RI-1-001 | 12K | |
| RI-2-001 | in-progress | RI-N2 (Forge): remove stub-executor false success; `--simulate` typed `simulated` results that satisfy nothing; literal-`true` gates and echo-review replaced with real gates or typed waiting-for-authority | #1275 | pi-glm-5.3 | mosaicstack/stack | fix/ri-050-forge-fail-closed | RI-0-001 | 20K | Independent review APPROVED 2026-08-17 (Gitea review 172 on PR #1278, head 99b8f6ea; reviewing seat fargo — recorded under shared host principal mos-dt-0, provenance correction posted by fred; wrapper gap filed by fred). Executed at head: forge tests 116/116, lint green, typecheck green after building macp dist (minimal-install artifact, not a defect), workspace typecheck 45/45, no external type consumers of the changed interfaces. CI red = known lane-wide fleet-test failure only, carries no information about this change (fred, log-content analysis, pipelines 2456-2458). Non-blocking finding: README L141-143 + skills/mosaic-forge/SKILL.md document bare forge run/resume, which now fails closed — fast-follow docs touch. Merge queued behind #1270. UPDATE 2026-08-18: #1270 merged; CI GREEN at head 4917df1f via serialized retry (pipeline 2477) - root cause of prior reds was CI-agent contention (web SPA timeouts under concurrent pipelines), superseding the fleet-test-failure theory. |
| RI-2-002 | in-progress | RI-N2 (MACP): gate runner fails closed on empty commands, stub executors, and unimplemented CI-provider gates unless explicit simulate; typed capability failures | #1275 | pi-glm-5.3 | mosaicstack/stack | fix/ri-050-macp-fail-closed | RI-0-001 | 15K | PR #1293 (head 2097379e): CI green (pipeline 2465), independent review APPROVED (Gitea review 173, jarvis seat, 2026-08-17) - macp 109/109 verified at head. Merge queued behind #1276/#1277/#1278. |
| RI-3-001 | not-started | RI-N4: complete probe inventory mapping every TS and shell quality-rail check to one canonical check with disposition (preserve/strengthen/retire, each named) | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-qr-probe-inventory | RI-0-001 | 12K | |
| RI-3-002 | not-started | RI-N4: TS evaluator absorbs effective shell probes; typed results (passed/failed/blocked/error/not-applicable) with versioned digested check definitions; shell commands become thin adapters; contract/parity/negative-control tests | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-qr-evaluator | RI-3-001 | 30K | |
| RI-4-001 | in-progress | RI-N3: one PRD application service — `mission --plan` persists mission↔PRD linkage (ids/versions/selected requirements); `mosaic prdy` routes through the service or becomes a named import/export adapter; Markdown is a labeled generated view; explicit conflict-aware import | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-prd-authority | RI-0-001 | 35K | PR #1294 (head 8d258e1d): CI green (pipeline 2466), independent review APPROVED (Gitea review 174, jarvis seat, 2026-08-17) - prdy 20/20 + command specs 9/9 at head. Merge queued behind #1276/#1277/#1278. |
| RI-5-001 | not-started | RI-N5: typed freshness states (current/stale/partial/unknown/unavailable); no failed-fetch-renders-empty; stale derived verdicts → unknown; mutations disabled when stale; failure-matrix tests | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-web-stale-safety | RI-0-001 | 25K | |
| RI-V-001 | not-started | Final verification + release evidence: all cards verified merged, negative controls demonstrated, real `next` publish run green on exact commit, evidence pack recorded | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-release-evidence | RI-1-002, RI-2-001, RI-2-002, RI-3-002, RI-4-001, RI-5-001 | 10K | |
## Dispatch waves (max 2 parallel workers)
1. RI-1-001 + RI-2-001
2. RI-2-002 + RI-4-001
3. RI-3-001 + RI-5-001
4. RI-1-002 + RI-3-002
5. RI-V-001
## Budget
Derived soft cap: 250K tokens (no explicit cap given). Projected total: 190K.
Conservative mode (1 worker) above 70% projected; freeze above 90%.
-242
View File
@@ -1,242 +0,0 @@
# Scratchpad — RI-050 orchestrator (jarvis, dragon-lin)
Mission: alpha 0.0.50 release-integrity floor. Issue #1275. Base `next` @ 476db12b.
Design SSOT: jarvis-brain `docs/plans/2026-08-16_mosaic-stack-sdlc-protocol.md` (SDLC-D-033..038).
## Mode (Jason's directives)
- Orchestrator: jarvis (this session, dragon-lin). NOT mos-claude; work stays on this host.
- Workers: local pi headless — `pi --model zai/glm-5.3:high -p` in the card's worktree, tools read,bash,edit,write.
- Delegation override of stack AGENTS.md `agent` column: rows carry `pi-glm-5.3` (outside cron table so no auto-claim).
- Target branch: `next`. Cards branch from `origin/next`, squash-merge via PR.
## Operational constraints (measured this session)
- Main checkout at `/home/jwoltje/src/mosaic-stack` is a dirty diverged `main` (ahead 1139/behind 711) — NEVER touched. All work in `/home/jwoltje/src/mosaic-stack-worktrees/<branch>`.
- Disk: /home 187G free. /tmp only 8.7G — keep pnpm stores/node_modules under /home.
- `main` and `next` have DIVERGED; PRs target `next`.
- Identity: pin `GITEA_LOGIN=mosaicstack-jarvis` for all wrapper ops. Issue #1275 verified authored by @jarvis.
- `ci-queue-wait.sh` on this host is fail-open (board: fix #1032 not installed) — substitute SHA-status checks via `/commits/{sha}/status` and diff failing step names.
- CI on PRs runs `pull_request` pipelines (any branch) incl. ci-postgres service. Push CI runs on main only; publish runs on push/tag to next + manual.
- Wrapper gaps on this host per board (7 gaps; e.g. no pr-review-list, issue-assign broken, pr-merge makes no trailers): verify outcomes by reading back provider state, never trust rc alone.
- Publish pipeline currently: install → build → publish-npm/publish-next-npm (+image). No verify. CI steps: install, sanitization, upgrade-guard, typecheck, lint, format, test, ci-postgres.
## Budget
Soft cap 250K. Projected 190K across 10 cards. Track per-card used vs estimate in TASKS.md notes.
## Progress log
- 2026-08-16 23:52 — Issue #1275 created (@jarvis verified).
- 2026-08-16 23:5x — Bootstrap branch `docs/ri-050-mission-bootstrap` from origin/next@476db12b; PRD section + TASKS.md + this scratchpad written. RI-0-001 in-progress.
## Wave 1 dispatched (2026-08-17 00:35)
- RI-1-001 worker: pi glm-5.3:high, pid 2322125, worktree ri-1-001, log /var/tmp/ri-050/ri-1-001-run.log
- RI-2-001 worker: pi glm-5.3:high, pid 2322126, worktree ri-2-001, log /var/tmp/ri-050/ri-2-001-run.log
- Gotcha recorded: pi has no -f flag (that's pi-do.sh); pass brief as positional message. First launch died "Unknown option: -f" — relaunched.
- CI lane: PR #1276 (bootstrap) fails `test` at base like every next PR — fred's green #1270 unblocks (comms sent 2026-08-17T05:21Z, `comms/20260817T052148Z__from-jarvis__650fe8.md`). Merge gate for all RI PRs queues behind #1270.
- Live RI-N1 evidence posted to #1275 (comment 22915): pipeline 2439 publish-next-npm SUCCESS beside build-gateway FAILURE.
---
# HANDOFF — RI-050 continuation (written 2026-08-17 ~08:45 UTC, jarvis/dragon-lin)
You are taking over the alpha 0.0.50 release-integrity workstream in place. Everything you
need is on the remote. Read this whole file, then `docs/release-integrity/TASKS.md` (same
branch), then the PRD section (`docs/PRD.md` § Release Integrity Workstream, same branch).
## Identity / mode
- Orchestrator identity: `jarvis` (dragon-lin). You continue as the RI-050 orchestrator under
whatever identity Jason gives you — if you are NOT jarvis, say so in comms and PR bodies.
- Jason's standing directives for this mission: work happens on THIS repo (mosaicstack/stack),
PRs target `next` (NOT main), workers are local pi headless sessions on
`zai/glm-5.3:high`. Do not hand this to mos-claude. Do not borrow other seats' lanes.
- All wrapper ops: pin `GITEA_LOGIN=mosaicstack-jarvis` (issue #1275 was verified authored by
@jarvis; keep identity consistent or verify yours with issue-view and READ BACK user.login).
- CI substitution rule (this host's ci-queue-wait.sh is fail-open; fix #1032 not installed):
judge CI by SHA-status via `/api/v1/repos/mosaicstack/stack/commits/{sha}/status` or the
woodpecker API (`pipeline-status.sh -r mosaicstack/stack -n N -f json`), and DIFF THE
FAILING STEP NAMES rather than trusting rc.
## Mission state at handoff
Mission: alpha 0.0.50 release-integrity floor. Issue #1275 (open, has live-evidence comment).
Decisions SDLC-D-033..038 live in jarvis-brain
`docs/plans/2026-08-16_mosaic-stack-sdlc-protocol.md` (normative text also mirrored in the
PRD section on this branch, so this repo is self-sufficient).
Base: `origin/next` @ 476db12b. NOTE: `main` and `next` have DIVERGED — never base on main.
Branches (all pushed, all clean trees):
- `docs/ri-050-mission-bootstrap` @ 5114faa2 → PR #1276 (open, mergeable) — bootstrap docs +
this scratchpad + TASKS.md DAG. STATUS: CI red on `test` only, which is the known lane-wide
failure (see blocker below); own prettier issue already fixed.
- `feat/ri-050-publish-gate` @ 0aa5ed35 → PR #1277 (open, mergeable) — RI-1-001 COMPLETE
(worker reported success, orchestrator review PASSED: verify step asserts CI_COMMIT_SHA ==
git rev-parse HEAD then runs canonical `pnpm verify:release`; every publish/image step
depends_on verify directly, confirmed by parsing the DAG: publish-npm, publish-next-npm,
build-gateway/appservice/web all -> [build, verify]; invariant test
scripts/verify-release.test.mjs passes 7/7 locally with negative fixtures). CI: same known
lane-red `test` step only.
- `fix/ri-050-forge-fail-closed` @ 99b8f6ea → PR #1278 (open, mergeable) — RI-2-001 worker
reported success (typed `FORGE_*` capability errors, --simulate typed simulated everywhere,
vacuous true/echo gates replaced, closed ForgeOutcome set, 116 tests green incl. 16 new).
ORCHESTRATOR REVIEW NOT YET DONE — your first job. Review the diff
(1391 insertions across forge src), check the fail-closed paths and that simulated
results cannot satisfy any consumer, run `pnpm --filter @mosaicstack/forge test`.
## The one blocker
Every `next` PR pipeline is red on ONE assertion:
`packages/mosaic/framework/tools/fleet/test-start-agent-session.sh:103` ("host provides 'pi'
in the system path"). Pre-existing at base; affects PRs #1276/#1277/#1278 identically.
fred's PR #1270 ("unblocks every PR on next") is green and open — it is HIS to merge; do not
merge it yourself. jarvis sent comms (`comms/20260817T052148Z__from-jarvis__650fe8.md` in
jarvis-brain) asking merge timing; no reply yet as of handoff. Merge gates for ALL RI PRs
queue behind #1270 landing. Until then: review/develop freely, merge nothing that needs the
green gate (docs-only #1276 arguably could merge red-lane with Jason's explicit call — ask,
don't assume).
## Remaining DAG (docs/release-integrity/TASKS.md is canonical)
Wave 2 (next): RI-2-002 MACP fail-closed (brief pattern: mirror RI-2-001 for
packages/macp/src/gate-runner.ts — empty commands, stub executors, unimplemented CI-provider
gates fail closed; explicit simulate) and RI-4-001 PRD authority (one PRD service;
@mosaicstack/prdy docs/prdy authoritative via `mosaic mission --plan`; `mosaic prdy` routes
or becomes named Markdown adapter; mission<->PRD linkage persists — see PRD RI-N3).
Wave 3: RI-3-001 probe inventory (docs), RI-5-001 web stale-safety.
Wave 4: RI-1-002 negative-control tests, RI-3-002 TS evaluator absorbs shell probes.
Final: RI-V-001 evidence pack (real green next publish run post-gate + all cards verified).
## Worker mechanics (measured, reuse)
- Dispatch: create worktree `git -C /home/jwoltje/src/mosaic-stack worktree add
/home/jwoltje/src/mosaic-stack-worktrees/<id> -b <branch> origin/next`, write a brief to
/var/tmp/ri-050/, then run from INSIDE the worktree:
`pi -p --no-session --model zai/glm-5.3:high --tools read,bash,edit,write "$(cat brief.md)"`
(pi has NO -f flag — pass the brief as a positional message; first dispatch died on that).
- Briefs for 1-001/2-001 are at /var/tmp/ri-050/ on dragon-lin (may not survive; the
pattern is fully described above and in TASKS.md).
- Briefs must carry: worktree path, branch, base, requirements, known base-red list (so the
worker doesn't chase it), gates to run, PR creation command with GITEA_LOGIN pin, "do NOT
merge, do NOT touch docs/TASKS.md", and the JSON report format.
- Verify worker claims: read the PR, run their tests yourself, parse pipeline step names.
## Do-not-touch
- Main checkout at /home/jwoltje/src/mosaic-stack (dirty diverged main) — never touch.
- fred's open PRs (#1270 and others) — review evidence welcome, merging his is not yours.
- Other RI PRs' authors' lanes: #1277/#1278 are yours to gate and merge ONCE lane is green
and review is recorded.
- Never `--no-verify`; never bypass the wrapper-fails-closed rule (wrapper failure ⇒
`blocked + report exact command + stop`).
## Session-restore command sequence
1. `git -C /home/jwoltje/src/mosaic-stack-worktrees/ri-050 fetch origin --prune`
2. Read this file + `docs/release-integrity/TASKS.md` + PRD section.
3. Check PR states (#1270, #1276, #1277, #1278) and lane CI (SHA-status per above).
4. Review RI-2-001 (PR #1278) if not yet done; then dispatch wave 2.
— jarvis, 2026-08-17
---
# CONTINUATION — fargo (sb-it-1-dt)
Orchestrator seat is now **fargo** on sb-it-1-dt (Jason, 2026-08-17): Claude seat, worktree discipline
per fred's ruling (`~/agent-work/<slug>`, create → work → commit → push → remove as one act; the
helper's `/src` refusal is a web1 convention, does not bind here). fred supports; lane rulings are
his. Workers remain local pi `zai/glm-5.3:high` + limited Claude per Jason.
## 2026-08-17 — RI-2-001 independent review DONE
- **PR #1278 APPROVED** (Gitea review 172, pinned to head 99b8f6ea). Executed evidence, not read-only:
forge suite 116/116 at head (matches PR claim), forge lint green, forge typecheck green after
building `@mosaicstack/macp` dist (TS2307 on bare `pnpm install --frozen-lockfile` is a
minimal-install build-order artifact — the macp import is type-only, vitest passes unbuilt; CI
installs build workspace deps, hence green there), **workspace typecheck 45/45 at head**,
consumer sweep: no external type consumers of RunManifest/StageStatus/ForgeTaskResult/
TaskExecutor; only importer of the package is packages/mosaic via registerForgeCommand
(smoke test asserts registration/help only — cannot break). Digest gate (shaggy's) before==after
with both-arm reactivity controls.
- CI red on #1276/#1277/#1278: lane-wide `test` failure only
(test-start-agent-session.sh:103, fred's guard mis-wired; #1270 unwires it). Fred measured log
content: one real byte-identical failure per pipeline (2456/2457/2458); 13 of ~14 `FAIL` grep
hits are passing fail-loud test NAMES. **The red carries no information about the RI changes.**
- Non-blocking finding: README L141-143 + skills/mosaic-forge/SKILL.md document bare
`mosaic forge run`/`resume`, which now exits 1 FORGE_NO_EXECUTOR — fast-follow docs touch.
- **Identity incident, ruled on by fred:** review 172 recorded under shared host principal
mos-dt-0, not fargo. Mechanism (measured, wrapper source): pr-review.sh resolves its acting login
from the tea login list only; no fargo tea login on this host → silent host-default fallback;
MOSAIC_GIT_IDENTITY is only read in detect-platform.sh get_gitea_token's fallback arm, never
reached. Exact-id read-back verifies against the writing token, so it passed while attribution
was wrong — durable-provenance machinery proves the write, not the seat. Fred's ruling: review
172 stands (substance/verdict/pin correct; label wrong); NO re-approval (one approval,
annotated, is the stronger record); fred posts the provenance correction under @fred with
--login fred-ms (hard-fail path); no fargo tea login ever (freeze + Jason's to authorize);
tooling gap filed by fred. Also explains (does not reopen) #1228's mos-dt-0 attribution.
- Merge gate: all RI PRs queue behind fred's green #1270 (Jason's call).
## Next
1. Wave 2 dispatch: RI-2-002 (MACP fail-closed, mirror RI-2-001 pattern for
packages/macp/src/gate-runner.ts) + RI-4-001 (PRD authority). Two parallel workers max.
2. Docs fast-follow (README + mosaic-forge skill) — fold into #1276 or a tiny docs card.
3. RI-V-001 evidence at the end.
— fargo, 2026-08-17
---
# RESUMPTION + DAILY-HANDOFF PROTOCOL (Jason, 2026-08-17)
Orchestrator seat is back with **jarvis** (dragon-lin). Expect daily handoff between jarvis
and fargo. Protocol (both seats, every handoff):
1. **This file is the shared mission log.** Append a dated section per session: state
measured, actions taken, PR/review states, next actions. Never rewrite prior sections.
2. **TASKS.md stays current within one session** — status, PR number in notes, review
evidence. Stale rows are handoff debt.
3. **Cross-review rule (SDLC-D-011 in practice):** the reviewing seat must differ from the
producing seat. jarvis reviews fargo-dispatched PRs, fargo reviews jarvis-dispatched
PRs. Producers are always pi workers; dispatching seats verify before push; the other
seat records the Gitea review.
4. Handoff = append here + push + (optional) issue #1275 comment if a decision changed.
## RESUMED — jarvis/dragon-lin, 2026-08-17 (afternoon)
- Measured: next = 8199261c (#1270 merged — lane unblocked for new PRs). #1293/#1294
(fargo, wave 2) CI-green, mergeable, no recorded reviews. #1276/#1277/#1278 still based
on 476db12b with stale red CI → need rebase onto 8199261c. #1278 review pinned to old
head 99b8f6ea by @mos-dt-0 (fargo's, mis-attributed per his note) — rebase will dismiss
it; re-approval must come from fargo/fred (author is @jarvis, cannot self-approve).
- Live evidence #2: push pipeline 2462 (the #1270 merge itself) ran publish-next-npm
SUCCESS beside build-gateway FAILURE again.
- Plan: rebase the three original branches; independently review #1293/#1294; merge order
once green+reviewed: #1276 (docs) → #1277 (publish gate) → #1278/#1293/#1294 (code).
After #1277 merges, watch the next push pipeline prove the verify gate live.
- fargo's non-RI PRs (#1291/#1296/#1297/#1281) stay strictly his lane.
## jarvis session 2026-08-17 (evening) — reviews, rebases, merge plan
- Rebased #1276/#1277/#1278 onto 8199261c (heads 59e2c460 / 46784c8d / 4917df1f);
invariant tests 7/7 and forge 116/116 re-run green at new heads. #1270 touched
test-enumeration-exclusions.txt + package.json, NOT ci.yml — no semantic overlap with
#1277's ci.yml changes (checked, was a real concern).
- Independent reviews recorded: #1293 APPROVED (review 173; macp 109/109; fail-closed paths
+ aggregate state machine verified), #1294 APPROVED (review 174; prdy 20/20 + command
specs 9/9; single-writer + linkage persistence + labeled export + conflict-aware import
verified). Note: 19 unrelated mosaic suites fail on bare minimal install (known workspace
build-order artifact, documented by fargo) — not this change.
- Measured: `next` has NO branch protection (API: only main listed). Cross-seat review
discipline is protocol-enforced, not Gitea-enforced. Flagged to fargo for Jason: direct
pushes to next trigger ungated publishes; protection is Jason's call (#1231 adjacent).
- Merge order planned: #1276 (docs-only — no publish run) -> #1277 (first gated publish)
-> #1278 -> #1293 -> #1294. Sent fargo review requests with pinned head SHAs
(comms/20260818T011932Z__from-jarvis__a9c02b.md). Not merging #1293/#1294 before my three
clear fargo's review — order optimality beats speed; every pre-#1277 merge publishes ungated.
- CI on the three rebased heads: pending at time of this entry.
-253
View File
@@ -1,253 +0,0 @@
import { mkdirSync, readFileSync, rmSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { randomUUID } from 'node:crypto';
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { normalizeGate, countAIFindings, runGate, runGates } from '../src/gate-runner.js';
function makeTmpDir(): string {
const dir = join(tmpdir(), `macp-gate-${randomUUID()}`);
mkdirSync(dir, { recursive: true });
return dir;
}
describe('normalizeGate', () => {
it('normalizes a string to mechanical gate', () => {
expect(normalizeGate('echo test')).toEqual({
command: 'echo test',
type: 'mechanical',
fail_on: 'blocker',
});
});
it('normalizes an object gate with defaults', () => {
expect(normalizeGate({ command: 'lint' })).toEqual({
command: 'lint',
type: 'mechanical',
fail_on: 'blocker',
});
});
it('preserves explicit type and fail_on', () => {
expect(normalizeGate({ command: 'review', type: 'ai-review', fail_on: 'any' })).toEqual({
command: 'review',
type: 'ai-review',
fail_on: 'any',
});
});
it('handles non-string/non-object input', () => {
expect(normalizeGate(42)).toEqual({ command: '', type: 'mechanical', fail_on: 'blocker' });
expect(normalizeGate(null)).toEqual({ command: '', type: 'mechanical', fail_on: 'blocker' });
});
});
describe('countAIFindings', () => {
it('returns zeros for non-object', () => {
expect(countAIFindings(null)).toEqual({ blockers: 0, total: 0 });
expect(countAIFindings('string')).toEqual({ blockers: 0, total: 0 });
expect(countAIFindings([])).toEqual({ blockers: 0, total: 0 });
});
it('counts from stats block', () => {
const output = { stats: { blockers: 2, should_fix: 3, suggestions: 1 } };
expect(countAIFindings(output)).toEqual({ blockers: 2, total: 6 });
});
it('counts from findings array when stats has no blockers', () => {
const output = {
stats: { blockers: 0 },
findings: [{ severity: 'blocker' }, { severity: 'warning' }, { severity: 'blocker' }],
};
expect(countAIFindings(output)).toEqual({ blockers: 2, total: 3 });
});
it('uses stats blockers over findings array when stats has blockers', () => {
const output = {
stats: { blockers: 5 },
findings: [{ severity: 'blocker' }, { severity: 'warning' }],
};
// stats.blockers = 5, total from stats = 5+0+0 = 5, findings not used for total since stats total is non-zero
expect(countAIFindings(output)).toEqual({ blockers: 5, total: 5 });
});
it('counts findings length as total when stats has zero total', () => {
const output = {
findings: [{ severity: 'warning' }, { severity: 'info' }],
};
expect(countAIFindings(output)).toEqual({ blockers: 0, total: 2 });
});
});
describe('runGate', () => {
let tmp: string;
let logPath: string;
beforeEach(() => {
tmp = makeTmpDir();
logPath = join(tmp, 'gate.log');
});
afterEach(() => {
rmSync(tmp, { recursive: true, force: true });
});
it('passes mechanical gate on exit 0', () => {
const result = runGate('echo hello', tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.exit_code).toBe(0);
expect(result.type).toBe('mechanical');
expect(result.output).toContain('hello');
});
it('fails mechanical gate on non-zero exit', () => {
const result = runGate('exit 1', tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.exit_code).toBe(1);
});
it('ci-pipeline always passes', () => {
const result = runGate({ command: 'anything', type: 'ci-pipeline' }, tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.type).toBe('ci-pipeline');
expect(result.output).toBe('CI pipeline gate placeholder');
});
it('empty command passes', () => {
const result = runGate({ command: '' }, tmp, logPath, 30);
expect(result.passed).toBe(true);
});
it('ai-review gate parses JSON output', () => {
const json = JSON.stringify({ stats: { blockers: 0, should_fix: 1 } });
const result = runGate({ command: `echo '${json}'`, type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.blockers).toBe(0);
expect(result.findings).toBe(1);
});
it('ai-review gate fails on blockers', () => {
const json = JSON.stringify({ stats: { blockers: 2 } });
const result = runGate({ command: `echo '${json}'`, type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.blockers).toBe(2);
});
it('ai-review gate with fail_on=any fails on any findings', () => {
const json = JSON.stringify({ stats: { blockers: 0, should_fix: 1 } });
const result = runGate(
{ command: `echo '${json}'`, type: 'ai-review', fail_on: 'any' },
tmp,
logPath,
30,
);
expect(result.passed).toBe(false);
expect(result.fail_on).toBe('any');
});
it('ai-review gate fails on invalid JSON output', () => {
const result = runGate({ command: 'echo "not json"', type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.parse_error).toBeDefined();
});
it('writes to log file', () => {
runGate('echo logged', tmp, logPath, 30);
const log = readFileSync(logPath, 'utf-8');
expect(log).toContain('COMMAND: echo logged');
expect(log).toContain('logged');
expect(log).toContain('EXIT:');
});
});
describe('runGates', () => {
let tmp: string;
let logPath: string;
let eventsPath: string;
beforeEach(() => {
tmp = makeTmpDir();
logPath = join(tmp, 'gates.log');
eventsPath = join(tmp, 'events.ndjson');
});
afterEach(() => {
rmSync(tmp, { recursive: true, force: true });
});
it('runs multiple gates and returns results', () => {
const { allPassed, gateResults } = runGates(
['echo one', 'echo two'],
tmp,
logPath,
30,
eventsPath,
'task-1',
);
expect(allPassed).toBe(true);
expect(gateResults).toHaveLength(2);
});
it('reports failure when any gate fails', () => {
const { allPassed, gateResults } = runGates(
['echo ok', 'exit 1'],
tmp,
logPath,
30,
eventsPath,
'task-2',
);
expect(allPassed).toBe(false);
expect(gateResults[0]!.passed).toBe(true);
expect(gateResults[1]!.passed).toBe(false);
});
it('emits events for each gate', () => {
runGates(['echo test'], tmp, logPath, 30, eventsPath, 'task-3');
const events = readFileSync(eventsPath, 'utf-8')
.trim()
.split('\n')
.map((l) => JSON.parse(l));
expect(events).toHaveLength(2); // started + passed
expect(events[0].event_type).toBe('rail.check.started');
expect(events[1].event_type).toBe('rail.check.passed');
});
it('skips gates with empty command (non ci-pipeline)', () => {
const { gateResults } = runGates(
[{ command: '', type: 'mechanical' }, 'echo real'],
tmp,
logPath,
30,
eventsPath,
'task-4',
);
expect(gateResults).toHaveLength(1);
});
it('does not skip ci-pipeline even with empty command', () => {
const { gateResults } = runGates(
[{ command: '', type: 'ci-pipeline' }],
tmp,
logPath,
30,
eventsPath,
'task-5',
);
expect(gateResults).toHaveLength(1);
expect(gateResults[0]!.passed).toBe(true);
});
it('emits failed event with correct message', () => {
runGates(['exit 42'], tmp, logPath, 30, eventsPath, 'task-6');
const events = readFileSync(eventsPath, 'utf-8')
.trim()
.split('\n')
.map((l) => JSON.parse(l));
const failEvent = events.find(
(e: Record<string, unknown>) => e.event_type === 'rail.check.failed',
);
expect(failEvent).toBeDefined();
expect(failEvent.message).toContain('Gate failed (');
});
});
+163 -1
View File
@@ -1,5 +1,8 @@
import { describe, it, expect } from 'vitest';
import { describe, it, expect, afterEach, beforeEach, vi } from 'vitest';
import { Command } from 'commander';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { registerMacpCommand } from './cli.js';
describe('registerMacpCommand', () => {
@@ -75,3 +78,162 @@ describe('registerMacpCommand', () => {
expect(topLevel).toContain('events');
});
});
/**
* RI-N2 fail-closed CLI behavior: an unimplemented capability is a failure,
* never a success. Every stub exits nonzero with a typed message, and the
* implemented `macp gate` mirrors the typed gate-runner states.
*/
describe('registerMacpCommand fail-closed (RI-N2)', () => {
let tmpDir: string;
function buildProgram(): Command {
const program = new Command();
program.exitOverride();
program.configureOutput({ writeErr: () => {} });
registerMacpCommand(program);
return program;
}
beforeEach(() => {
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'macp-cli-failclosed-'));
process.exitCode = 0;
});
afterEach(() => {
process.exitCode = 0;
fs.rmSync(tmpDir, { recursive: true, force: true });
});
it('macp tasks list exits nonzero (unimplemented capability)', async () => {
const program = buildProgram();
await program.parseAsync(['macp', 'tasks', 'list'], { from: 'user' });
expect(process.exitCode).not.toBe(0);
});
it('macp submit exits nonzero with a typed MACP_NOT_IMPLEMENTED message', async () => {
const program = buildProgram();
const errSpy = vi.spyOn(console, 'error').mockImplementation(() => {});
try {
await program.parseAsync(['macp', 'submit', 'spec.json'], { from: 'user' });
expect(process.exitCode).not.toBe(0);
const errText = errSpy.mock.calls.map((c) => String(c[0])).join('\n');
expect(errText).toContain('MACP_NOT_IMPLEMENTED');
} finally {
errSpy.mockRestore();
}
});
it('macp events tail exits nonzero (unimplemented capability)', async () => {
const program = buildProgram();
await program.parseAsync(['macp', 'events', 'tail'], { from: 'user' });
expect(process.exitCode).not.toBe(0);
});
it('macp gate runs a green inline command and exits 0', async () => {
const program = buildProgram();
await program.parseAsync(
[
'macp',
'gate',
'exit 0',
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
expect(process.exitCode).toBe(0);
});
it('macp gate exits nonzero on a failing command', async () => {
const program = buildProgram();
await program.parseAsync(
[
'macp',
'gate',
'exit 9',
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
expect(process.exitCode).not.toBe(0);
});
it('macp gate with an unimplemented ci-pipeline capability exits nonzero', async () => {
const program = buildProgram();
const specPath = path.join(tmpDir, 'gates.json');
fs.writeFileSync(specPath, JSON.stringify([{ type: 'ci-pipeline' }]));
await program.parseAsync(
[
'macp',
'gate',
specPath,
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
expect(process.exitCode).not.toBe(0);
});
it('macp gate --simulate completes (exit 0) but reports simulated results', async () => {
const program = buildProgram();
const logSpy = vi.spyOn(console, 'log').mockImplementation(() => {});
try {
await program.parseAsync(
[
'macp',
'gate',
'exit 0',
'--simulate',
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
// completes only because the caller explicitly asked to simulate
expect(process.exitCode).toBe(0);
const outText = logSpy.mock.calls.map((c) => String(c[0])).join('\n');
expect(outText).toContain('simulated');
expect(outText).toContain('SIMULATED');
} finally {
logSpy.mockRestore();
}
});
it('macp gate with an empty spec exits nonzero with a typed error', async () => {
const program = buildProgram();
await program.parseAsync(
[
'macp',
'gate',
' ',
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
expect(process.exitCode).not.toBe(0);
});
});
+129 -19
View File
@@ -1,5 +1,73 @@
import { existsSync, readFileSync } from 'node:fs';
import type { Command } from 'commander';
import { runGates } from './gate-runner.js';
import { MACPCapabilityError, type MacpErrorCode } from './errors.js';
/**
* Load gates from a spec: an existing file (JSON gates array, a JSON object
* with `quality_gates`, a JSON gate object, or one command per line) or an
* inline command string. Fails closed with a typed capability error when the
* spec contains no executable gate definition.
*/
function loadGateSpec(spec: string): unknown[] {
if (existsSync(spec)) {
const raw = readFileSync(spec, 'utf-8');
try {
const parsed = JSON.parse(raw) as unknown;
if (Array.isArray(parsed)) {
if (parsed.length === 0) {
throw new MACPCapabilityError(
'MACP_NO_COMMAND',
'gate-spec',
`gate spec file '${spec}' contains an empty gates array`,
);
}
return parsed;
}
if (typeof parsed === 'object' && parsed !== null) {
const obj = parsed as Record<string, unknown>;
if (Array.isArray(obj['quality_gates'])) {
return obj['quality_gates'];
}
return [parsed];
}
throw new MACPCapabilityError(
'MACP_NO_COMMAND',
'gate-spec',
`gate spec file '${spec}' parsed to ${typeof parsed} — expected a gates array, a task with quality_gates, or a gate object`,
);
} catch (exc) {
if (exc instanceof MACPCapabilityError) throw exc;
// Not JSON — treat each non-empty line as a command gate.
const lines = raw
.split('\n')
.map((l) => l.trim())
.filter((l) => l.length > 0);
if (lines.length > 0) return lines;
throw new MACPCapabilityError(
'MACP_NO_COMMAND',
'gate-spec',
`gate spec file '${spec}' contains no gates`,
);
}
}
if (spec.trim().length > 0) return [spec];
throw new MACPCapabilityError('MACP_NO_COMMAND', 'gate-spec', 'gate spec is empty');
}
/** Print a typed not-implemented failure and exit nonzero (RI-N2 fail-closed). */
function notImplemented(subcommand: string, capability: string, hint: string): void {
const err = new MACPCapabilityError(
'MACP_NOT_IMPLEMENTED',
capability,
`${subcommand} is not implemented in @mosaicstack/macp yet (${capability} capability absent) — ${hint}`,
);
console.error(`[macp] ${subcommand}: ${err.message} [${err.code}]`);
process.exitCode = 1;
}
/**
* Register macp subcommands on an existing Commander program.
* This avoids cross-package Commander version mismatches by using the
@@ -24,15 +92,14 @@ export function registerMacpCommand(parent: Command): void {
'Filter by task type (coding|deploy|research|review|documentation|infrastructure)',
)
.action((opts: { status?: string; type?: string }) => {
// not yet wired — task persistence layer is not present in @mosaicstack/macp
console.log('[macp] tasks list: not yet wired — use macp package programmatically');
// unimplemented capability — a failure, never a success (RI-N2)
if (opts.status) {
console.log(` status filter: ${opts.status}`);
}
if (opts.type) {
console.log(` type filter: ${opts.type}`);
}
process.exitCode = 0;
notImplemented('tasks list', 'task-persistence', 'use the macp package programmatically');
});
// ─── submit ──────────────────────────────────────────────────────────────
@@ -41,12 +108,11 @@ export function registerMacpCommand(parent: Command): void {
.command('submit <path>')
.description('Submit a task from a JSON/YAML spec file')
.action((specPath: string) => {
// not yet wired — task submission requires a running MACP server
console.log('[macp] submit: not yet wired — use macp package programmatically');
// unimplemented capability — a failure, never a success (RI-N2)
console.log(` spec path: ${specPath}`);
console.log(' task id: (unavailable — no MACP server connected)');
console.log(' status: (unavailable — no MACP server connected)');
process.exitCode = 0;
notImplemented('submit', 'macp-server', 'use the macp package programmatically');
});
// ─── gate ────────────────────────────────────────────────────────────────
@@ -58,16 +124,58 @@ export function registerMacpCommand(parent: Command): void {
.option('--cwd <path>', 'Working directory for gate execution', process.cwd())
.option('--log <path>', 'Path to write gate log output', '/tmp/macp-gate.log')
.option('--timeout <seconds>', 'Gate timeout in seconds', '60')
.action((spec: string, opts: { failOn: string; cwd: string; log: string; timeout: string }) => {
// not yet wired — gate execution requires a task context and event sink
console.log('[macp] gate: not yet wired — use macp package programmatically');
console.log(` spec: ${spec}`);
console.log(` fail-on: ${opts.failOn}`);
console.log(` cwd: ${opts.cwd}`);
console.log(` log: ${opts.log}`);
console.log(` timeout: ${opts.timeout}s`);
process.exitCode = 0;
});
.option(
'--simulate',
'Simulate gates instead of executing them; results are typed simulated and never satisfy a check',
)
.action(
(
spec: string,
opts: { failOn: string; cwd: string; log: string; timeout: string; simulate?: boolean },
) => {
let gates: unknown[];
try {
gates = loadGateSpec(spec);
} catch (exc) {
if (exc instanceof MACPCapabilityError) {
console.error(`[macp] gate: ${exc.message} [${exc.code}]`);
} else {
console.error(`[macp] gate: ${String(exc)}`);
}
process.exitCode = 1;
return;
}
const timeoutSec = Number.parseInt(opts.timeout, 10) || 60;
const eventsPath = `${opts.log}.events.ndjson`;
const { state, gateResults } = runGates(
gates,
opts.cwd,
opts.log,
timeoutSec,
eventsPath,
'macp-cli-gate',
{
simulate: opts.simulate,
},
);
for (const r of gateResults) {
const label = r.command || r.type;
const reason = r.reason ? `${r.reason}` : '';
console.log(`[macp] gate ${r.status}: ${label}${reason}`);
}
if (opts.simulate) {
console.log(
'[macp] SIMULATED run — every result is typed simulated and can never satisfy a gate, dependency, or release check',
);
}
// Simulated runs may complete (exit 0) only because the caller
// explicitly passed --simulate; the typed state stays 'simulated'.
process.exitCode = state === 'passed' || state === 'simulated' ? 0 : 1;
},
);
// ─── events ──────────────────────────────────────────────────────────────
@@ -79,14 +187,16 @@ export function registerMacpCommand(parent: Command): void {
.option('--file <path>', 'Path to the MACP events NDJSON file')
.option('--follow', 'Follow the file for new events (like tail -f)')
.action((opts: { file?: string; follow?: boolean }) => {
// not yet wired — event streaming requires a live event source
console.log('[macp] events tail: not yet wired — use macp package programmatically');
// unimplemented capability — a failure, never a success (RI-N2)
if (opts.file) {
console.log(` file: ${opts.file}`);
}
if (opts.follow) {
console.log(' mode: follow');
}
process.exitCode = 0;
notImplemented('events tail', 'event-source', 'use the macp package programmatically');
});
}
// Re-export so CLI consumers can surface typed capability codes.
export type { MacpErrorCode };
+35
View File
@@ -0,0 +1,35 @@
/** Typed error code from the closed MACP_ERROR_CODES set. */
export type MacpErrorCode = (typeof MACP_ERROR_CODES)[number];
/**
* Typed fail-closed capability errors (RI-N2, SDLC-D-035).
*
* MACP must fail closed when a required capability (executor, reviewer,
* command, CI provider, human authority) is absent. These typed codes mirror
* the Forge failure vocabulary (FORGE_NO_*) so both packages speak the same
* language: an unimplemented capability is a failure, never a stub success.
*/
/** Closed set of typed MACP capability error codes. */
export const MACP_ERROR_CODES = [
'MACP_NOT_IMPLEMENTED',
'MACP_NO_COMMAND',
'MACP_NO_REVIEWER',
'MACP_NO_CI_PIPELINE',
'MACP_NO_PROVIDER',
'MACP_AUTHORITY_REQUIRED',
] as const;
/** Raised when a required capability is missing and execution must fail closed. */
export class MACPCapabilityError extends Error {
/** Typed error code from the closed MACP_ERROR_CODES set. */
readonly code: MacpErrorCode;
/** The missing capability, e.g. `ci-provider`, `task-persistence`, `command`. */
readonly capability: string;
constructor(code: MacpErrorCode, capability: string, message: string) {
super(message);
this.name = 'MACPCapabilityError';
this.code = code;
this.capability = capability;
}
}
+429
View File
@@ -0,0 +1,429 @@
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { afterEach, beforeEach, describe, expect, it } from 'vitest';
import { countAIFindings, normalizeGate, runGate, runGates } from './gate-runner.js';
function makeTmpDir(): string {
return fs.mkdtempSync(path.join(os.tmpdir(), 'macp-gate-'));
}
describe('normalizeGate', () => {
it('normalizes a string to mechanical gate', () => {
expect(normalizeGate('echo test')).toEqual({
command: 'echo test',
type: 'mechanical',
fail_on: 'blocker',
});
});
it('normalizes an object gate with defaults', () => {
expect(normalizeGate({ command: 'lint' })).toEqual({
command: 'lint',
type: 'mechanical',
fail_on: 'blocker',
});
});
it('preserves explicit type and fail_on', () => {
expect(normalizeGate({ command: 'review', type: 'ai-review', fail_on: 'any' })).toEqual({
command: 'review',
type: 'ai-review',
fail_on: 'any',
});
});
it('handles non-string/non-object input', () => {
expect(normalizeGate(42)).toEqual({ command: '', type: 'mechanical', fail_on: 'blocker' });
expect(normalizeGate(null)).toEqual({ command: '', type: 'mechanical', fail_on: 'blocker' });
});
});
describe('countAIFindings', () => {
it('returns zeros for non-object', () => {
expect(countAIFindings(null)).toEqual({ blockers: 0, total: 0 });
expect(countAIFindings('string')).toEqual({ blockers: 0, total: 0 });
expect(countAIFindings([])).toEqual({ blockers: 0, total: 0 });
});
it('counts from stats block', () => {
const output = { stats: { blockers: 2, should_fix: 3, suggestions: 1 } };
expect(countAIFindings(output)).toEqual({ blockers: 2, total: 6 });
});
it('counts from findings array when stats has no blockers', () => {
const output = {
stats: { blockers: 0 },
findings: [{ severity: 'blocker' }, { severity: 'warning' }, { severity: 'blocker' }],
};
expect(countAIFindings(output)).toEqual({ blockers: 2, total: 3 });
});
it('uses stats blockers over findings array when stats has blockers', () => {
const output = {
stats: { blockers: 5 },
findings: [{ severity: 'blocker' }, { severity: 'warning' }],
};
// stats.blockers = 5, total from stats = 5+0+0 = 5, findings not used for total since stats total is non-zero
expect(countAIFindings(output)).toEqual({ blockers: 5, total: 5 });
});
it('counts findings length as total when stats has zero total', () => {
const output = {
findings: [{ severity: 'warning' }, { severity: 'info' }],
};
expect(countAIFindings(output)).toEqual({ blockers: 0, total: 2 });
});
});
describe('runGate', () => {
let tmp: string;
let logPath: string;
beforeEach(() => {
tmp = makeTmpDir();
logPath = path.join(tmp, 'gate.log');
});
afterEach(() => {
fs.rmSync(tmp, { recursive: true, force: true });
});
it('passes mechanical gate on exit 0', () => {
const result = runGate('echo hello', tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.exit_code).toBe(0);
expect(result.type).toBe('mechanical');
expect(result.output).toContain('hello');
});
it('fails mechanical gate on non-zero exit', () => {
const result = runGate('exit 1', tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.exit_code).toBe(1);
});
it('ci-pipeline fails closed without a CI provider (no placeholder pass)', () => {
const result = runGate({ command: 'anything', type: 'ci-pipeline' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.status).toBe('capability_failure');
expect(result.capability_code).toBe('MACP_NO_CI_PIPELINE');
expect(result.type).toBe('ci-pipeline');
expect(result.output).not.toBe('CI pipeline gate placeholder');
});
it('empty command is a typed capability failure, never a pass', () => {
const result = runGate({ command: '' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.status).toBe('capability_failure');
expect(result.capability_code).toBe('MACP_NO_COMMAND');
});
it('ai-review gate parses JSON output', () => {
const json = JSON.stringify({ stats: { blockers: 0, should_fix: 1 } });
const result = runGate({ command: `echo '${json}'`, type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.blockers).toBe(0);
expect(result.findings).toBe(1);
});
it('ai-review gate fails on blockers', () => {
const json = JSON.stringify({ stats: { blockers: 2 } });
const result = runGate({ command: `echo '${json}'`, type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.blockers).toBe(2);
});
it('ai-review gate with fail_on=any fails on any findings', () => {
const json = JSON.stringify({ stats: { blockers: 0, should_fix: 1 } });
const result = runGate(
{ command: `echo '${json}'`, type: 'ai-review', fail_on: 'any' },
tmp,
logPath,
30,
);
expect(result.passed).toBe(false);
expect(result.fail_on).toBe('any');
});
it('ai-review gate fails on invalid JSON output', () => {
const result = runGate({ command: 'echo "not json"', type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.parse_error).toBeDefined();
});
it('writes to log file', () => {
runGate('echo logged', tmp, logPath, 30);
const log = fs.readFileSync(logPath, 'utf-8');
expect(log).toContain('COMMAND: echo logged');
expect(log).toContain('logged');
expect(log).toContain('EXIT:');
});
});
describe('runGates', () => {
let tmp: string;
let logPath: string;
let eventsPath: string;
beforeEach(() => {
tmp = makeTmpDir();
logPath = path.join(tmp, 'gates.log');
eventsPath = path.join(tmp, 'events.ndjson');
});
afterEach(() => {
fs.rmSync(tmp, { recursive: true, force: true });
});
it('runs multiple gates and returns results', () => {
const { allPassed, gateResults } = runGates(
['echo one', 'echo two'],
tmp,
logPath,
30,
eventsPath,
'task-1',
);
expect(allPassed).toBe(true);
expect(gateResults).toHaveLength(2);
});
it('reports failure when any gate fails', () => {
const { allPassed, gateResults } = runGates(
['echo ok', 'exit 1'],
tmp,
logPath,
30,
eventsPath,
'task-2',
);
expect(allPassed).toBe(false);
expect(gateResults[0]!.passed).toBe(true);
expect(gateResults[1]!.passed).toBe(false);
});
it('emits events for each gate', () => {
runGates(['echo test'], tmp, logPath, 30, eventsPath, 'task-3');
const events = fs
.readFileSync(eventsPath, 'utf-8')
.trim()
.split('\n')
.map((l) => JSON.parse(l));
expect(events).toHaveLength(2); // started + passed
expect(events[0].event_type).toBe('rail.check.started');
expect(events[1].event_type).toBe('rail.check.passed');
});
it('does not silently skip gates with empty command — they become capability failures', () => {
const { gateResults, allPassed, state } = runGates(
[{ command: '', type: 'mechanical' }, 'echo real'],
tmp,
logPath,
30,
eventsPath,
'task-4',
);
expect(gateResults).toHaveLength(2);
expect(gateResults[0]!.status).toBe('capability_failure');
expect(gateResults[1]!.status).toBe('passed');
expect(allPassed).toBe(false);
expect(state).toBe('capability_failure');
});
it('does not skip ci-pipeline even with empty command — typed capability failure', () => {
const { gateResults, allPassed, state } = runGates(
[{ command: '', type: 'ci-pipeline' }],
tmp,
logPath,
30,
eventsPath,
'task-5',
);
expect(gateResults).toHaveLength(1);
expect(gateResults[0]!.passed).toBe(false);
expect(gateResults[0]!.status).toBe('capability_failure');
expect(allPassed).toBe(false);
expect(state).toBe('capability_failure');
});
it('emits failed event with correct message', () => {
runGates(['exit 42'], tmp, logPath, 30, eventsPath, 'task-6');
const events = fs
.readFileSync(eventsPath, 'utf-8')
.trim()
.split('\n')
.map((l) => JSON.parse(l));
const failEvent = events.find(
(e: Record<string, unknown>) => e.event_type === 'rail.check.failed',
);
expect(failEvent).toBeDefined();
expect(failEvent.message).toContain('Gate failed (');
});
});
/**
* RI-N2 / SDLC-D-035 fail-closed controls for the MACP gate runner.
*
* Invariant under test: `passed: true` occurs ONLY when a gate really executed
* and really exited green (`status === 'passed'`). Absent capabilities,
* manual sign-offs, and simulated runs are typed distinctly and can never
* make the aggregate `passed`.
*/
describe('gate-runner fail-closed (RI-N2)', () => {
let tmpDir: string;
let logPath: string;
let eventsPath: string;
beforeEach(() => {
tmpDir = makeTmpDir();
logPath = path.join(tmpDir, 'gate.log');
eventsPath = path.join(tmpDir, 'events.ndjson');
});
afterEach(() => {
fs.rmSync(tmpDir, { recursive: true, force: true });
});
function run(gates: unknown[], options?: { simulate?: boolean }) {
return runGates(gates, tmpDir, logPath, 10, eventsPath, 'spec-task', options);
}
// ─── positive controls ───────────────────────────────────────────────────
it('a really-executed green command gate still passes', () => {
const result = run([{ command: 'exit 0', type: 'mechanical' }]);
expect(result.gateResults[0]!.status).toBe('passed');
expect(result.gateResults[0]!.passed).toBe(true);
expect(result.allPassed).toBe(true);
expect(result.state).toBe('passed');
});
it('explicit simulate completes and types every result simulated', () => {
const result = run([{ command: 'exit 0', type: 'mechanical' }, 'echo hello'], {
simulate: true,
});
expect(result.gateResults).toHaveLength(2);
for (const gate of result.gateResults) {
expect(gate.status).toBe('simulated');
expect(gate.passed).toBe(false);
}
expect(result.state).toBe('simulated');
});
it('a really-executed red command gate fails with typed status failed', () => {
const result = run([{ command: 'exit 3', type: 'mechanical' }]);
expect(result.gateResults[0]!.status).toBe('failed');
expect(result.gateResults[0]!.passed).toBe(false);
expect(result.allPassed).toBe(false);
expect(result.state).toBe('failed');
});
// ─── negative controls — each asserts typed status AND aggregate not passed ──
it('an empty-command gate is a capability_failure, not skipped and not passed', () => {
const result = run([{ command: '', type: 'mechanical' }]);
// runGates must not silently skip it — it produces a typed result
expect(result.gateResults).toHaveLength(1);
const gate = result.gateResults[0]!;
expect(gate.status).toBe('capability_failure');
expect(gate.capability_code).toBe('MACP_NO_COMMAND');
expect(gate.passed).toBe(false);
// aggregate is not passed
expect(result.allPassed).toBe(false);
expect(result.state).toBe('capability_failure');
expect(result.state).not.toBe('passed');
});
it('a commandless ai-review gate is a typed MACP_NO_REVIEWER capability_failure', () => {
const result = run([{ command: '', type: 'ai-review' }]);
expect(result.gateResults[0]!.status).toBe('capability_failure');
expect(result.gateResults[0]!.capability_code).toBe('MACP_NO_REVIEWER');
expect(result.allPassed).toBe(false);
expect(result.state).not.toBe('passed');
});
it('a ci-pipeline gate without a provider implementation is a capability_failure, never a placeholder pass', () => {
const result = run([{ command: '', type: 'ci-pipeline' }]);
const gate = result.gateResults[0]!;
expect(gate.status).toBe('capability_failure');
expect(gate.capability_code).toBe('MACP_NO_CI_PIPELINE');
expect(gate.passed).toBe(false);
// the old false-success placeholder must be gone
expect(gate.output).not.toBe('CI pipeline gate placeholder');
expect(result.allPassed).toBe(false);
expect(result.state).not.toBe('passed');
});
it('a ci-pipeline gate fails closed even alongside an otherwise green run', () => {
const result = run(['exit 0', { type: 'ci-pipeline', command: 'fake-ci' }]);
expect(result.gateResults[1]!.status).toBe('capability_failure');
expect(result.gateResults[0]!.status).toBe('passed');
expect(result.allPassed).toBe(false);
expect(result.state).toBe('capability_failure');
});
it('a manual gate with no automation enters typed waiting — neither pass nor fail', () => {
const result = run([{ type: 'manual' }]);
const gate = result.gateResults[0]!;
expect(gate.status).toBe('waiting');
expect(gate.passed).toBe(false);
expect(gate.exit_code).toBe(0);
// aggregate is not passed while any gate is waiting
expect(result.allPassed).toBe(false);
expect(result.state).toBe('waiting');
expect(result.state).not.toBe('passed');
});
it('a simulated result can never make the aggregate passed', () => {
const result = run(['exit 0', 'exit 0'], { simulate: true });
expect(result.gateResults.every((g) => g.status === 'simulated')).toBe(true);
expect(result.allPassed).toBe(false);
expect(result.state).toBe('simulated');
expect(result.state).not.toBe('passed');
});
it('waiting dominates an otherwise green aggregate', () => {
const result = run(['exit 0', { type: 'manual' }]);
expect(result.allPassed).toBe(false);
expect(result.state).toBe('waiting');
});
});
describe('runGate fail-closed (RI-N2)', () => {
let tmpDir: string;
let logPath: string;
beforeEach(() => {
tmpDir = makeTmpDir();
logPath = path.join(tmpDir, 'gate.log');
});
afterEach(() => {
fs.rmSync(tmpDir, { recursive: true, force: true });
});
it('simulate: true returns a typed simulated result without executing', () => {
const result = runGate('this-command-does-not-exist-xyz', tmpDir, logPath, 10, {
simulate: true,
});
expect(result.status).toBe('simulated');
expect(result.passed).toBe(false);
expect(result.exit_code).toBe(0);
});
it('normal mode executes for real and types a green gate passed', () => {
const result = runGate('echo ok', tmpDir, logPath, 10);
expect(result.status).toBe('passed');
expect(result.passed).toBe(true);
expect(result.output).toContain('ok');
});
it('a bare string gate normalizes to mechanical and executes', () => {
const result = runGate('exit 7', tmpDir, logPath, 10);
expect(result.type).toBe('mechanical');
expect(result.status).toBe('failed');
expect(result.passed).toBe(false);
});
});
+148 -25
View File
@@ -4,7 +4,20 @@ import { dirname } from 'node:path';
import { emitEvent } from './event-emitter.js';
import { nowISO } from './event-emitter.js';
import type { GateResult } from './types.js';
import type { GateResult, GateStatus, RunGatesResult } from './types.js';
/** Typed reason stamped on every simulated gate result. */
export const SIMULATED_GATE_REASON =
'simulated execution (explicit simulate opt-in): gate was not evaluated by a real implementation';
/** Options for gate execution (RI-N2 fail-closed / explicit simulation). */
export interface RunGateOptions {
/**
* Explicit caller opt-in to simulation. Simulated gates are NOT executed;
* every result is typed `simulated` and never satisfies anything.
*/
simulate?: boolean;
}
export interface NormalizedGate {
command: string;
@@ -103,36 +116,91 @@ export function countAIFindings(parsedOutput: unknown): { blockers: number; tota
return { blockers, total };
}
function simulatedResult(gateEntry: NormalizedGate): GateResult {
return {
command: gateEntry.command,
exit_code: 0,
type: gateEntry.type,
output: SIMULATED_GATE_REASON,
timed_out: false,
passed: false,
status: 'simulated',
reason: SIMULATED_GATE_REASON,
};
}
function capabilityFailureResult(
gateEntry: NormalizedGate,
code: GateResult['capability_code'],
reason: string,
): GateResult {
return {
command: gateEntry.command,
exit_code: 1,
type: gateEntry.type,
output: '',
timed_out: false,
passed: false,
status: 'capability_failure',
capability_code: code,
reason,
};
}
function waitingResult(gateEntry: NormalizedGate, reason: string): GateResult {
return {
command: gateEntry.command,
exit_code: 0,
type: gateEntry.type,
output: '',
timed_out: false,
passed: false,
status: 'waiting',
capability_code: 'MACP_AUTHORITY_REQUIRED',
reason,
};
}
export function runGate(
gate: unknown,
cwd: string,
logPath: string,
timeoutSec: number,
options: RunGateOptions = {},
): GateResult {
const gateEntry = normalizeGate(gate);
const gateType = gateEntry.type;
const command = gateEntry.command;
// Explicit simulation only: never executes, typed simulated, never satisfying.
if (options.simulate) {
return simulatedResult(gateEntry);
}
// Fail closed: no CI provider implementation exists in @mosaicstack/macp,
// so a ci-pipeline gate is an absent capability — never a placeholder pass.
if (gateType === 'ci-pipeline') {
return {
command,
exit_code: 0,
type: gateType,
output: 'CI pipeline gate placeholder',
timed_out: false,
passed: true,
};
return capabilityFailureResult(
gateEntry,
'MACP_NO_CI_PIPELINE',
`ci-pipeline gate '${gateEntry.command || gateType}' has no CI provider implementation wired — refusing placeholder pass`,
);
}
if (!command) {
return {
command: '',
exit_code: 0,
type: gateType,
output: '',
timed_out: false,
passed: true,
};
// A manual gate with no automation waits for human sign-off: not pass, not fail.
if (gateType === 'manual') {
return waitingResult(
gateEntry,
`manual gate has no automation — waiting for human sign-off (type: ${gateType})`,
);
}
// Any other commandless gate is an absent capability — never a vacuous pass.
return capabilityFailureResult(
gateEntry,
gateType === 'ai-review' ? 'MACP_NO_REVIEWER' : 'MACP_NO_COMMAND',
`gate of type '${gateType}' has no command to execute — refusing empty-command pass`,
);
}
const { exitCode, output, timedOut } = runShell(command, cwd, logPath, timeoutSec);
@@ -143,10 +211,12 @@ export function runGate(
output,
timed_out: timedOut,
passed: false,
status: 'failed',
};
if (gateType !== 'ai-review') {
result.passed = exitCode === 0;
result.status = result.passed ? 'passed' : 'failed';
return result;
}
@@ -170,6 +240,7 @@ export function runGate(
} else {
result.passed = exitCode === 0 && blockers === 0 && !timedOut && parseError === undefined;
}
result.status = result.passed ? 'passed' : 'failed';
result.fail_on = failOn;
result.blockers = blockers;
@@ -191,16 +262,19 @@ export function runGates(
timeoutSec: number,
eventsPath: string,
taskId: string,
): { allPassed: boolean; gateResults: GateResult[] } {
let allPassed = true;
options: RunGateOptions = {},
): RunGatesResult {
const gateResults: GateResult[] = [];
let hasCapabilityFailure = false;
let hasSimulated = false;
let hasFailed = false;
let hasWaiting = false;
for (const gate of gates) {
const gateEntry = normalizeGate(gate);
const gateCmd = gateEntry.command;
if (!gateCmd && gateEntry.type !== 'ci-pipeline') continue;
const label = gateCmd || gateEntry.type;
// NOTE: no silent skip — every gate produces a typed result (RI-N2).
emitEvent(
eventsPath,
'rail.check.started',
@@ -209,10 +283,10 @@ export function runGates(
'quality-gate',
`Running gate: ${label}`,
);
const result = runGate(gate, cwd, logPath, timeoutSec);
const result = runGate(gate, cwd, logPath, timeoutSec, options);
gateResults.push(result);
if (result.passed) {
if (result.status === 'passed') {
emitEvent(
eventsPath,
'rail.check.passed',
@@ -224,7 +298,46 @@ export function runGates(
continue;
}
allPassed = false;
if (result.status === 'waiting') {
hasWaiting = true;
emitEvent(
eventsPath,
'rail.check.waiting',
taskId,
'gated',
'quality-gate',
`Gate waiting: ${label}${result.reason ?? 'manual gate awaits sign-off'}`,
);
continue;
}
if (result.status === 'simulated') {
hasSimulated = true;
emitEvent(
eventsPath,
'rail.check.simulated',
taskId,
'gated',
'quality-gate',
`Gate simulated (non-satisfying): ${label}`,
);
continue;
}
if (result.status === 'capability_failure') {
hasCapabilityFailure = true;
emitEvent(
eventsPath,
'rail.check.failed',
taskId,
'gated',
'quality-gate',
`Gate capability failure (${result.capability_code ?? 'MACP_NO_PROVIDER'}): ${label}${result.reason ?? 'required capability is absent'}`,
);
continue;
}
hasFailed = true;
let message: string;
if (result.timed_out) {
message = `Gate timed out after ${timeoutSec}s: ${label}`;
@@ -236,5 +349,15 @@ export function runGates(
emitEvent(eventsPath, 'rail.check.failed', taskId, 'gated', 'quality-gate', message);
}
return { allPassed, gateResults };
const state: GateStatus = hasCapabilityFailure
? 'capability_failure'
: hasSimulated
? 'simulated'
: hasFailed
? 'failed'
: hasWaiting
? 'waiting'
: 'passed';
return { allPassed: state === 'passed', gateResults, state };
}
+16 -2
View File
@@ -6,11 +6,13 @@ export type {
DependsOnPolicy,
GateType,
GateFailOn,
GateStatus,
GateEntry,
Task,
EventType,
MACPEvent,
GateResult,
RunGatesResult,
TaskResult,
ProviderMeta,
ProviderRegistry,
@@ -18,6 +20,11 @@ export type {
export { CredentialError } from './types.js';
// Typed fail-closed capability errors (RI-N2, SDLC-D-035)
export { MACP_ERROR_CODES, MACPCapabilityError } from './errors.js';
export type { MacpErrorCode } from './errors.js';
// Credential resolver
export {
DEFAULT_CREDENTIALS_DIR,
@@ -35,9 +42,16 @@ export {
export type { ResolveCredentialsOptions } from './credential-resolver.js';
// Gate runner
export { normalizeGate, runShell, countAIFindings, runGate, runGates } from './gate-runner.js';
export {
normalizeGate,
runShell,
countAIFindings,
runGate,
runGates,
SIMULATED_GATE_REASON,
} from './gate-runner.js';
export type { NormalizedGate } from './gate-runner.js';
export type { NormalizedGate, RunGateOptions } from './gate-runner.js';
// Risk-floor (agent reflection loop — diff review classifier)
export { evaluateRiskFloor, DEFAULT_RISK_THRESHOLD } from './risk-floor.js';
+39 -2
View File
@@ -1,3 +1,5 @@
import type { MacpErrorCode } from './errors.js';
/** Task status values. */
export type TaskStatus = 'pending' | 'running' | 'gated' | 'completed' | 'failed' | 'escalated';
@@ -17,7 +19,17 @@ export type DispatchMode = 'yolo' | 'acp' | 'exec';
export type DependsOnPolicy = 'all' | 'any' | 'all_terminal';
/** Quality gate type. */
export type GateType = 'mechanical' | 'ai-review' | 'ci-pipeline';
export type GateType = 'mechanical' | 'ai-review' | 'ci-pipeline' | 'manual';
/**
* Typed execution state of a gate closed set (RI-N2, SDLC-D-035).
*
* Only `passed` means "really executed and green". `simulated` is produced
* exclusively under an explicit simulate opt-in and never satisfies anything.
* `capability_failure` means a required executor/provider/command was absent.
* `waiting` means a manual gate awaits human sign-off (neither pass nor fail).
*/
export type GateStatus = 'passed' | 'failed' | 'simulated' | 'waiting' | 'capability_failure';
/** Gate fail_on mode. */
export type GateFailOn = 'blocker' | 'any';
@@ -67,7 +79,9 @@ export type EventType =
| 'task.retry.scheduled'
| 'rail.check.started'
| 'rail.check.passed'
| 'rail.check.failed';
| 'rail.check.failed'
| 'rail.check.waiting'
| 'rail.check.simulated';
/** Structured event record. */
export interface MACPEvent {
@@ -88,7 +102,14 @@ export interface GateResult {
type: string;
output: string;
timed_out: boolean;
/** Back-compat boolean view — true ONLY when `status === 'passed'`. */
passed: boolean;
/** Typed discriminator — the authoritative gate outcome (RI-N2). */
status: GateStatus;
/** Typed capability error code, set when `status === 'capability_failure'`. */
capability_code?: MacpErrorCode;
/** Why a non-executed state (simulated/waiting/capability_failure) was reached. */
reason?: string;
fail_on?: string;
blockers?: number;
findings?: number;
@@ -96,6 +117,22 @@ export interface GateResult {
parse_error?: string;
}
/**
* Aggregate outcome of `runGates` (RI-N2).
*
* `state` is the typed aggregate: it is `passed` only when every gate really
* executed green. A `simulated` result makes the aggregate `simulated` (never
* `passed`); a `waiting` manual gate keeps the aggregate `waiting`; a missing
* capability makes it `capability_failure`. `allPassed` is exactly
* `state === 'passed'`, so a simulated or waiting result can never satisfy a
* dependency, acceptance criterion, gate, merge, or release check.
*/
export interface RunGatesResult {
allPassed: boolean;
gateResults: GateResult[];
state: GateStatus;
}
/** Result from a completed task. */
export interface TaskResult {
task_id: string;
+3 -6
View File
@@ -133,13 +133,10 @@ When the full `@mosaicstack/forge` package is available, Forge uses MACP task ex
```bash
# Run from CLI
# Fails closed with a typed FORGE_NO_EXECUTOR capability error when no real
# executor is wired — pass --simulate to opt into explicit typed simulation
# (every result carries status `simulated`, which satisfies nothing).
mosaic forge run path/to/brief.md [--simulate]
mosaic forge run path/to/brief.md
# Resume interrupted run (same fail-closed rule as forge run)
mosaic forge resume .forge/runs/20260401-143022/ [--simulate]
# Resume interrupted run
mosaic forge resume .forge/runs/20260401-143022/
# Check status
mosaic forge status .forge/runs/20260401-143022/