Compare commits

..
Author SHA1 Message Date
ops-01 018d96a412 fix(merge): drop fleet test-start-agent-session.sh from framework-shell chain — review B1
ci/woodpecker/pr/ci Pipeline failed
Review stack#1324 id 201 (rev-code-01): the union resolution chained main's #1073
fleet test while keeping next's signed exclusion of it (#1271 burn-down: the test
asserts missing-binary behavior but CI provides pi on PATH — fails by design).
check-test-enumeration CONTRADICTORY EXCLUSION, CI pipeline 2532 terminal failure.

Fix per verified review recommendation: remove the chain entry (55 -> 54 parts),
keeping next's exclusion and #1271 state. Guard now passes:
population 60, enumerated 45, excluded (signed) 16.
2026-08-19 16:44:31 -05:00
ops-01 b01950e92f merge: absorb main into next — 23-commit divergence (08-05..13 base=main window)
ci/woodpecker/pr/ci Pipeline failed
17 content commits + 3 merge bubbles were genuinely missing from next (~8,000 lines:
goal controller #1152, framework enforcement #1174/#1195, pr-edit wrapper #1173/#1200,
pipefail series #1100/#1105/#1106/#1107, git-tools fixes #1073/#1085/#1086/#1089,
#991, #1007, enrollment tolerance #1094). 3 commits were already in next by content
(#1060 identical, #1066/#1062 evolved twins — conflicts resolved to next's side).

Per-commit classification and evidence: mosaic-brain fleet/lanes/stack-remediation/main-next-divergence.md.
Conflict resolutions (6 files) itemized in the PR body.
2026-08-19 16:27:17 -05:00
fargo 1bdeed62eb fix(#1320): placeholder-ize private-network topology, drop raw-curl force-merge recipe (#1322)
ci/woodpecker/push/publish Pipeline was successful
2026-08-19 21:15:54 +00:00
fargo 840c2b0d96 Merge pull request 'skills: fold agent-skills into the monorepo, single install path, promote ms-unslop (plan phase D)' (#1319) from fold-agent-skills into next
ci/woodpecker/push/publish Pipeline was successful
2026-08-19 20:08:09 +00:00
fargo 95d5cb32d4 format: cover folded skills' js/ts scripts with repo prettier
ci/woodpecker/pr/ci Pipeline was successful
The repo format:check glob covers ts/js alongside md; four non-markdown
scripts inside the folded tree were flagged after the md pass. Same pinned
prettier 3.8.1, same markup-only class (verified: node --check still passes
on the js files).
2026-08-19 14:41:18 -05:00
fargo d5f3fae896 skills: promote ms-unslop from skills-local into the package
Phase D3 of plan 2026-08-19 (decision S21): skills-local is the local test
bed; promote individually as each proves out. ms-unslop is the only one of
the seven local skills with working enforcement evidence — a checker
(tools/unslop-hook/unslop-check.js), a machine-source list (lists.json), a
19-test suite, a measured corpus, and a regression fixture.

The checker itself stays fleet-local for now (it binds to a specific
harness extension surface); this promotes the skill document only.

Gates verified on the moved file: sanitization denylist clean, prettier
3.8.1 clean (6730 chars was fred's measure at assignment; 7249 as shipped
today — both pass).
2026-08-19 14:39:46 -05:00
fargo 1556982dbc skills: single install path — canonical skills ship with the framework
Phase D2 of plan 2026-08-19: single package, single install, single command.

The framework installer already treats skills/** as a shipped, manifest-owned
framework subtree, so the folded skills now install into
$MOSAIC_HOME/skills with the rest of the framework — no second repository,
no separate sync step:

- mosaic-sync-skills (bash + powershell): the fetch machinery is gone (clone,
  pull, dirty-state migration, rsync from sources/agent-skills). The script
  now only links installed skills into runtime homes. --link-only is a compat
  no-op; --no-link exits having nothing to do.
- catalog.ts: the sources/agent-skills fallback is dead and removed.
- install.sh, launch.ts, defaults/README.md, README.md, skills/README.md:
  references to the second repo rewritten to describe the shipped path.

Verified: clean install into a fresh MOSAIC_HOME produces 102 skills with no
sources/ directory; the linker then links the selected skills into the four
runtime homes with no git involvement.
2026-08-19 14:39:28 -05:00
fargo 1a822493ba format: apply repo prettier (3.8.1) to the folded skills tree
963 markdown files reformatted with the repository's pinned prettier so
pnpm format:check covers the folded tree like every other repo file.

The formatter's embedded-language pass also normalized code fences
(TS semicolons, closed HTML tags in examples, lowercased CSS hex colors,
one renumbered list that skipped an index). Alphanumeric token deltas vs
the fold commit were audited file-by-file; all are formatter-equivalent
markup normalizations plus the four sanitized skills.
2026-08-19 14:37:17 -05:00
fargo d2eeb64433 skills: sanitize operator-identity tokens from folded ops skills
Four folded skills carried operator identity tokens that the sanitization
gate (verify-sanitized.sh) forbids in the public framework package:

- kickstart: template path pointed at a private brain checkout; now uses the
  framework-shipped $MOSAIC_HOME/templates/docs/TASKS.md.template
- mosaic-deploy: dropped one estate-specific stack-name row from the example
  table
- mosaic-portainer, mosaic-woodpecker: credentials now name the framework
  credentials store (load_credentials <service>) instead of a private
  checkout path

Estate-specific values can live in a skills-local override, which the linker
applies with precedence over canonical skills.
2026-08-19 14:34:00 -05:00
fargo 5e58597dbe fold: absorb mosaicstack/agent-skills into the framework skills tree
Fold the agent-skills repository into the monorepo as the shipped canonical
skills package (plan 2026-08-19 phase D1, decision S19: single package, single
install, single command).

History is preserved by rewriting each commit's paths from skills/ to
packages/mosaic/framework/skills/ (git fast-export/import) and merging the
rewritten history with --allow-unrelated-histories, so the original commits
with their authors, dates, and messages remain reachable. Blob content is
untouched by the rewrite; tree fidelity was verified blob-sha-for-blob-sha.

This change must be merged with a real merge commit (not squash) or the
history link is destroyed.
2026-08-19 14:33:03 -05:00
jason.woltje fe4fa20309 Merge pull request 'guides: add SEAT-IDENTITY and FLEET-COMMS; harden CODE-REVIEW evidence rules' (#1313) from fred/guides-seat-identity-fleet-comms into next
ci/woodpecker/push/publish Pipeline was successful
Reviewed-on: #1313
Reviewed-by: rev-code-01 <[email protected]>
2026-08-19 15:44:57 +00:00
fred 5e93ef70bd guides: fix the cross-reference direction in SEAT-IDENTITY
ci/woodpecker/pr/ci Pipeline was successful
rev-code-01's non-blocking nit on #1313 round 2. The no-linking-step paragraph
pointed at the bridge explanation as 'described below'; it is above. Now names the
section, which survives further reordering better than a direction word does.

Text-only. Verified with the repo's PINNED prettier (3.8.1 via pnpm-lock.yaml) and
the sanitization gate, both clean.
2026-08-18 19:10:09 -05:00
fred 3884f2de4d guides: address rev-code-01's review of #1313 (B1, B2, S1, S2)
ci/woodpecker/pr/ci Pipeline was successful
All four findings reproduced before fixing. rev-code-01 was right on each.

B2 (blocker, mine). SEAT-IDENTITY provisioning step 4 said to symlink the
framework store entry to the seat slot, while the same file says those bridges
must not be recreated. The same bridge, told both ways, in one document. I
rewrote the resolution and token-location sections when the deploy made them
stale and did not carry the change into the numbered steps. Step 4 is gone and
the file now says explicitly that no provisioning step links the store to the
slot, so the omission cannot read as an oversight.

S1 (mine). The guide claimed the helper "attempts a fleet notification" on
refusal. The shipped helper does no such thing — its only reference to
notification is a comment saying an alert built on the record is best-effort, and
there is no send or wake call anywhere in the file. Now: it writes a durable
record, the record is what exists, and nobody should wait for a notification that
nothing sends. A guide that promises an alert is worse than one that promises
nothing.

S2. Estate-local content removed from files that ship to every estate: the
~/.mosaic/fleet/bin script paths (dead paths elsewhere) and the 2026-08-18 dates,
which dated a specific host's migration rather than describing behavior. The
bridge-removal passage now states the ORDERING that matters — remove bridges only
after a seat-aware helper can reach the slot, never before — which is the part
that transfers.

B1. prettier reformatted all three files. Reproduced the pipeline 2515 failure
locally before and confirmed clean after; the other three guides prettier flags
are untouched by this branch (0 changes vs origin/next) and are pre-existing.

Sanitization gate re-run and passing.

Verified for the record, since I could not verify my own work: rev-code-01
confirmed the no-fallback claim TRUE against helper content on origin/next, and
judged the evidence rules actionable on the grounds that each names an executable
replacement.
2026-08-18 18:51:15 -05:00
fred a3c50d91ca guides: genericize the operator name in SEAT-IDENTITY provisioning
ci/woodpecker/pr/ci Pipeline failed
Pipeline 2514 failed the sanitization gate on 'Jason mints the token into the
seat slot'. The denylist is jarvis|jason|woltje|... and a shipped framework file
must not carry operator identity. My mistake: I generalized the estate paths and
seat names when promoting this guide and did not check the operator name.

Now reads 'the estate operator', with the accompanying rule that an agent does
not ask another agent to mint one either.

Verified by running tools/quality/scripts/verify-sanitized.sh locally rather than
guessing at the pattern: gate passes.
2026-08-18 18:28:37 -05:00
fred 2fd102e6af guides: state the decree, drop the mechanism
ci/woodpecker/pr/ci Pipeline failed
The #1280 prohibition carried an explanation of how the tools misattribute and
why the failure is invisible from inside them. A reader who is not going to use
the tool cannot act on any of it. Same for rule 2's closing clause about what
reviews commonly miss. Both cut to the decree and the corrective action.

Rules 1 and 3-12 keep their trailing sentences: those are corrective actions or
the detail that makes the case recognizable, not justification.
2026-08-18 18:25:47 -05:00
fred efb3c3a10c guides: add SEAT-IDENTITY and FLEET-COMMS; harden CODE-REVIEW evidence rules
ci/woodpecker/pr/ci Pipeline failed
Three guides that existed only as one host's working copy, promoted to framework
templates so every estate gets them. A working copy under ~/.mosaic binds one
host; only a template here binds all of them.

SEAT-IDENTITY.md (new) documents how a seat's git credential is actually
resolved after #1311: identity from MOSAIC_GIT_IDENTITY, then
mosaic.gitIdentity, then the stdin username; host mapped to a store prefix; then
ONE of two stores chosen by whether the seat directory exists, with no
precedence and no fallback between them. A seat with a directory and an empty
slot fails closed rather than reaching the service store, and that is the point.

It also corrects how to find the helper. credential.helper commonly names an
absolute path, so `command -v git-credential-mosaic` answers a different question
than the one git asks, and the two stop agreeing the moment the PATH copy is
removed. Git also tries EVERY configured helper in order, so a fail-closed helper
in front silently hands the request to whatever is configured behind it. The
guide says to read the whole list.

FLEET-COMMS.md (new) documents agent-send.sh: the class table, the addressing
preamble, and the exit codes — including that rc=2 means the text reached the
pane as an unsubmitted draft, so retrying double-sends it. Confirm with
capture-pane instead. It also says to measure the fleet rather than trust
roster.yaml, which on a live host was simultaneously naming a socket that did not
exist, listing seats that were not running, and omitting seats that were.

CODE-REVIEW.md gains an Evidence Discipline section: a green is not a result
until you have shown it could go red, measurement and explanation are separate
sentences, verify by content on the ref that ships rather than by ancestry of a
local sha, and confidence is part of a finding. Plus four shell-measurement rules
earned on #1311, each of which produced a wrong conclusion first — `cmd | tail;
echo rc=$?` reports tail's status, a missed glob under pipefail exits 2 and kills
the run under set -e, nonzero-with-no-output is an environment question before it
is a code question, and `git -C` in a non-repo directory answers from the
enclosing repo.

The estate-specific repository exception that lived in the working copy is not
carried here. The template says an estate may document one, scoped to a named
repository and never precedent for a second.

Both new guides are added to the two routing tables that agents read.
2026-08-18 18:15:50 -05:00
jason.woltje d4d32a80b2 Merge pull request 'git credentials: fail closed, and read a seat's token from its own slot' (#1311) from fred/credential-fail-closed-seat-slots into next
ci/woodpecker/push/publish Pipeline was successful
Reviewed-on: #1311
2026-08-18 22:40:46 +00:00
fred c703cc50eb git-credential-mosaic: escape the escalation record, and stop naming a record that was never written
ci/woodpecker/pr/ci Pipeline was successful
Both defects found in review by rev-code-01 on #1311.

F3 — the JSONL record interpolated every field with a bare %s. An identity comes
from git config or the environment and a cwd is whatever directory git ran in, so
either can contain a quote or a backslash. One such refusal turned the day's spool
into unparseable JSONL, and the operator would only discover it while reading the
record that explains an outage. Fields are now JSON-escaped.

F2 — the diagnostic printed "record: <spool>/<date>.jsonl" unconditionally, but
the record is only written inside the branch where mkdir -p succeeded. When the
spool cannot be created the helper named a file that does not exist, on exactly
the hosts where the escalation was lost. It now reports the real path or says
NOT WRITTEN.

Also: prettier on README.md, which was the format-step failure on pipeline 2508.
It reflowed only the two tables this branch added.

Tests: cases 14 and 15 cover both. Verified discriminating — against the previous
helper with these same tests, case 14 fails with the unparseable record printed
and case 15 fails on both assertions; against this one both pass.

The first draft of case 14 used `ls "$spool"/*.jsonl | head -1`, which under
`set -o pipefail` exits 2 on a missed glob and killed the suite with zero output
— the same silent-nonzero failure rev-code-01 hit from a partial tools/ extraction
and the reason this file exists. Replaced with a glob loop and a comment.
2026-08-18 17:04:55 -05:00
fred 3d2b712355 git credentials: fail closed, and read a seat's token from its own slot
ci/woodpecker/pr/ci Pipeline failed
Two changes to one rule: a credential is resolved from exactly one place,
and an identity that cannot be resolved is refused rather than substituted.

FAIL CLOSED. Both readers ended in an unconditional fall-through to the
shared Gitea account whenever an identity did not resolve. Every seat in a
fleet therefore pushed, opened PRs and filed reviews under one account, and
a record made that way cannot be traced to the agent that made it
afterwards. The fallback now applies only where there is no attribution to
lose: a host with no fleet. Where seats exist, an unresolvable request emits
nothing, exits nonzero, explains itself on stderr, and — in the git helper —
appends a record naming the identity, host, reason and cwd, and no token
value, to ${MOSAIC_CREDENTIAL_SPOOL:-~/.local/state/mosaic-credential-escalations}.

A host runs a fleet when <brain>/fleet/agents exists, which is the signal
packages/mosaic/src/fleet/brain-home.ts already uses to decide a brain is
active, resolved the same way (MOSAIC_BRAIN_HOME, else ~/.mosaic). This is
what keeps the change a no-op for an operator who has not provisioned
per-slot tokens: no fleet directory, shared account, unchanged. It is also
why there is no environment variable to restore the old behavior — one would
reintroduce the substitution being removed.

STORE SELECTION. Both readers hardcoded ~/.config/mosaic/secrets/gitea-tokens,
so a seat's own secrets/ slot was invisible to the framework: a seat could
hold a valid credential and still be served the shared account. The store is
now chosen by what the identity is. An identity with a directory under
<brain>/fleet/agents/ is a seat and is read only from
<brain>/fleet/agents/<id>/secrets/; any other identity is a service identity
and is read from the framework store. There is no precedence between them
and no fallback from one to the other, so a seat with an empty slot is
refused even when a same-named token sits in the framework store. Two copies
of one credential are drift rather than redundancy, and drift surfaces as
the stale copy returning 401, which reads as a revoked token and sends
whoever debugs it somewhere else.

detect-platform.sh is in scope alongside git-credential-mosaic because they
are the two readers of these tokens. Patching only the git helper would make
"one credential, one location" true for push and fetch and false for
pr-create.sh, issue-create.sh and pr-review.sh, which is the harder failure
to notice.

TESTS. The three assertions that pinned the shared-account fall-through are
now fail-closed assertions, and a refusal is checked four independent ways:
nonzero exit, empty stdout, a stderr diagnostic naming identity and host,
and no shared token value anywhere in the output. The exit code alone would
pass against a helper that emitted the credential and then failed. Added:
seat-slot resolution, the no-cross-store-fallback case with a control
proving the framework-store file it declines to read is readable, no-identity
on a fleet host, the fleet gate firing on the default ~/.mosaic and not only
on an injected MOSAIC_BRAIN_HOME, and a cross-host leak check. Both suites
were run against the pre-change code as a control and fail there on exactly
the shared-token emission.

shellcheck is not installed on the authoring host, so the rewritten helper
is unlinted locally and CI is the first lint of it.
2026-08-18 16:19:43 -05:00
fargo 245e0c427d feat(quality-rails): typed evaluator absorbs QC-19/QC-20; verify-release wiring (RI-3-002, #1275) (#1308)
ci/woodpecker/push/publish Pipeline failed
2026-08-18 17:54:47 +00:00
jarvisandfargo ff45f7b5d0 docs(ri-050): forge fail-closed docs + TASKS status catch-up (#1275) (#1299)
ci/woodpecker/push/publish Pipeline failed
Co-authored-by: Jarvis <[email protected]>
2026-08-18 16:00:36 +00:00
jarvisandfargo 64350892e7 docs(ri-050): RI-3-001 complete quality-rails probe inventory (#1275) (#1302)
ci/woodpecker/push/publish Pipeline was canceled
Co-authored-by: jarvis <[email protected]>
2026-08-18 15:59:34 +00:00
jason.woltjeandjarvis 6e9df3c640 fix(doctor): greenfield brain lock-in note (fred's #1301 follow-up, #1288) (#1307)
ci/woodpecker/push/publish Pipeline failed
Co-authored-by: Jason Woltje <[email protected]>
2026-08-18 07:57:37 +00:00
jarvis f5ba042dfa test(ri-050): RI-1-002 publish-gate negative controls (#1275) (#1305)
ci/woodpecker/push/publish Pipeline failed
2026-08-18 05:57:06 +00:00
jarvis 7c7dab3898 feat(ri-050): RI-5-001 typed freshness states and stale-safe Mission Control surfaces (#1275) (#1300)
ci/woodpecker/push/publish Pipeline was canceled
2026-08-18 05:56:58 +00:00
fargoandjarvis d92de53399 feat(prd): one transitional PRD authority — RI-4-001 (#1275) (#1294)
ci/woodpecker/push/publish Pipeline was canceled
Co-authored-by: fargo <[email protected]>
2026-08-18 05:56:51 +00:00
mos-dt-0andjarvis d7e303d3c0 fix(macp): fail-closed typed gate states — RI-2-002 (#1275) (#1293)
ci/woodpecker/push/publish Pipeline was canceled
Co-authored-by: mos-dt-0 <[email protected]>
2026-08-18 05:56:38 +00:00
jarvis 726d2ad3a2 fix(ri-050): forge fails closed without providers; explicit typed simulation (#1275) (#1278)
ci/woodpecker/push/publish Pipeline was canceled
2026-08-18 05:52:45 +00:00
jason.woltjeandjarvis e4ee1acf24 feat(doctor): brain-home fleet-state check (#1298 follow-up) (#1301)
ci/woodpecker/push/publish Pipeline failed
Co-authored-by: Jason Woltje <[email protected]>
2026-08-18 05:26:01 +00:00
jarvis 5c5a25e4de fix(ci): image pushes read the registry secrets that exist (#1275) (#1306)
ci/woodpecker/push/publish Pipeline failed
Squash-merged by topher (jarvis principal) via break-glass API path (wrapper main-only gap, documented). Gates: CI 2492 green at head e19013ed, review 181 APPROVED (fred) at pinned head. Diagnostic merge: the push pipeline now exercises kaniko with the REGISTRY_* secrets - valid creds yield the first fully green gated publish; invalid yield an explicit 401.

Co-authored-by: Jarvis <[email protected]>
2026-08-18 05:02:23 +00:00
jarvis 7669321ea2 test(gateway): cross-user-isolation cleanup honors dbAvailable — unblocks gated publish verify (#1275) (#1304)
ci/woodpecker/push/publish Pipeline failed
Squash-merged by topher (jarvis principal) via break-glass API path (wrapper main-only gap, documented in #1275 log). Gates: CI 2487 green at head 81f500bd, review 180 APPROVED (fred) at pinned head. Unblocks gated publish: verify's no-DB path now skips cleanly.

Co-authored-by: Jarvis <[email protected]>
2026-08-18 04:24:48 +00:00
coder2andMos 7102ccb93e docs(tools): index pull request edit wrapper (#1200)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
Co-authored-by: coder2 <[email protected]>
2026-08-13 14:50:38 +00:00
Mos afdaa6d0e6 framework: make tool discoverability, workspace placement and model tiering mechanical (#1174)
ci/woodpecker/push/ci Pipeline failed
ci/woodpecker/push/publish Pipeline was successful
2026-08-13 14:21:22 +00:00
coder3andMos 41749bbd33 fix(framework): detect installed tool drift (#1194) (#1195)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
Co-authored-by: coder3 <[email protected]>
2026-08-13 10:43:11 +00:00
Mos 120af4e193 feat(git-tools): add pull request edit wrapper (#1080) (#1173)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
2026-08-13 06:28:01 +00:00
mos-dt-0 ec260e678f Merge pull request 'fix(git): accept http/https as one scheme class in comment URL verification (#991)' (#1022) from fix/991-comment-url-scheme-normalise into main
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
2026-08-12 01:11:09 +00:00
Jason Woltje b590a5c3d8 fix(git): accept http/https as one scheme class in comment URL verification (#991)
ci/woodpecker/pr/ci Pipeline was successful
issue-comment.sh and pr-review.sh verify a durable write by pinning the
provider-returned object URL's origin and full path. The origin included the
SCHEME verbatim. On a Gitea whose ROOT_URL is configured `http://` while every
client reaches it over `https://`, the provider returns `http://` object URLs,
so the comparison rejects the provider's own truthful answer about a write that
LANDED. The failure is deterministic, not intermittent: every comment, every
time, on such a deployment.

The scheme was never what the check defends. The forgeries it exists to catch —
look-alike host, decoy path prefix, wrong owner/repo/kind/number — all vary the
HOST or the PATH. Both stay strict. `http` and `https` now collapse to one
scheme class; any other scheme (file:, ftp:, javascript:) stays distinguishing,
and an EXPLICIT non-default port still distinguishes, because a different port
is a different service on the same host.

Consequences of the bug, both observed:

- The wrapper reports failure on a comment that is durably on the issue/PR, and
  attributes it to #865 ("no durable comment created"). The write landed; the
  citation is wrong. Reproduced here: the harness's persisted state contains the
  record while the wrapper exits 1.
- pr-review.sh's comment path is worse. On a host where no seat can create a
  review OBJECT, comment-form is the only gate-16 review record obtainable, and
  this check refuses all of it.

Test gap this closes: every URL fixture in both harnesses was `https://`, and
every negative case varied only host or path. The one axis that fails in
production had zero coverage — the fixtures encoded the assumption that breaks.
Added, in both suites:

- scheme-downgrade (http vs https, otherwise correct) — must be ACCEPTED. Fails
  against the unmodified wrappers, passes against the fixed ones; verified in
  both directions, and the negative control's captured output is the #865
  misattribution above.
- explicit non-default port (`:8443`) — must stay REJECTED.
- non-web scheme (`ftp://`) — must stay REJECTED.

Also fixes test-issue-comment-readback.sh hermeticity (#1007), without which the
suite cannot run on any seat that has a per-agent Gitea token: detect-platform's
step-0 identity lookup reads ~/.config/mosaic/gitea-tokens/<identity>, outside
both XDG_CONFIG_HOME and MOSAIC_CREDENTIALS_FILE, so the suite resolved a
PRODUCTION credential and died at HTTP 401 before case 1. Same two-part fix
already merged for test-pr-review-gitea-comment.sh in #1006: a sandboxed HOME
plus an empty REPO-LOCAL mosaic.gitIdentity to shadow the global. Note the
env-var route does NOT work — detect-platform.sh reads `${MOSAIC_GIT_IDENTITY:-}`
and `:-` treats set-but-empty identically to unset.

The owner-side half of #991 (setting the deployment's Gitea ROOT_URL to https)
is not in scope here and is not made unnecessary by this change; this makes the
wrappers correct against a deployment that returns either scheme.
2026-08-11 19:00:31 -05:00
mos-dt-0 540ec5b6ef Merge pull request 'fix(git): #1007 suite hermeticity — pin repo-local mosaic.gitIdentity in five test suites' (#1024) from fix/1007-suite-hermeticity into main
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline failed
2026-08-11 23:19:49 +00:00
mos-dt-0 563d1ac053 Merge pull request 'fix(shell): remove wake validation pipe hazards' (#1107) from fix/1099-pipefail-wake into main
ci/woodpecker/push/publish Pipeline was canceled
ci/woodpecker/push/ci Pipeline was canceled
2026-08-11 23:19:45 +00:00
Mos 4d8ddb9a0a fix: quote SKILL.md descriptions containing colons (silent skill-load failure) (#3) 2026-08-11 22:53:47 +00:00
mos-dt-0 722163671f feat(pi): add persistent Mosaic /goal controller (#1152)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
2026-08-10 22:54:43 +00:00
f10-coder f158be8003 fix(shell): remove wake validation pipe hazards
ci/woodpecker/pr/ci Pipeline was successful
2026-08-07 06:13:47 -05:00
f10-coderandMos b0f7d26dd9 fix(shell): remove test harness pipe hazards (#1106)
ci/woodpecker/push/publish Pipeline failed
ci/woodpecker/push/ci Pipeline was successful
ci/woodpecker/manual/ci-image Pipeline failed
ci/woodpecker/manual/publish Pipeline failed
ci/woodpecker/manual/ci Pipeline was successful
Co-authored-by: f10-coder <[email protected]>
2026-08-07 11:12:17 +00:00
f10-coderandMos 3a1203b2f8 fix(shell): remove runtime early-exit pipe hazards (#1105)
ci/woodpecker/push/publish Pipeline failed
ci/woodpecker/push/ci Pipeline was successful
Co-authored-by: f10-coder <[email protected]>
2026-08-07 09:38:37 +00:00
be-coder-08 df4c591ab4 fix(fleet): make framework shell assertions SIGPIPE-safe (#1100)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
2026-08-07 08:21:20 +00:00
be-coder-06andMos 4fa2768962 fix(fleet): propagate roster git identity (#1073)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
Co-authored-by: be-coder-06 <[email protected]>
2026-08-07 07:07:45 +00:00
Mos aa0a7b5fa2 fix(tools/git): issue-close.sh silently dropped the closing comment (#1085)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
2026-08-07 05:59:12 +00:00
Mos 42ac19af48 test(gateway): size the enrollment clamp tolerance to CI jitter, not to a fast machine (closes #1090) (#1094)
ci/woodpecker/push/publish Pipeline failed
ci/woodpecker/push/ci Pipeline was successful
2026-08-07 05:33:38 +00:00
Mos f744f32214 feat(tools/git): explain tea's misleading user does not exist error (stale token, not a missing account) (#1086)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
2026-08-07 05:07:40 +00:00
Mos 8ff7aac0ca fix(tools/git): detect-platform died silently outside a repo, taking every wrapper with it (#1089)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
2026-08-07 04:26:36 +00:00
be-coder-08andMos 80a45b1e1c feat(pr-merge): preserve linked authors in squash messages (#1066)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
Co-authored-by: be-coder-08 <[email protected]>
2026-08-06 05:36:59 +00:00
be-coder-08andMos 85d2108e4e fix(ci): remove upgrade rollback signal race (#1060)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
Co-authored-by: be-coder-08 <[email protected]>
2026-08-05 22:14:15 +00:00
be-coder-08andMos 16f91157a1 test(ci): make queue guard harness deterministic (#1062)
ci/woodpecker/push/publish Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
Co-authored-by: be-coder-08 <[email protected]>
2026-08-05 21:49:44 +00:00
Jason Woltje 2fa6bcd576 fix(git): #1007 — test-issue-comment-readback is a FIFTH affected suite (second census correction)
ci/woodpecker/pr/ci Pipeline was successful
My previous commit said four. It is five. `test-issue-comment-readback.sh` has
the same defect and is fixed the same way, and I had already looked straight at
it and filed it as an *unrelated* silent failure. Correcting that here rather
than folding it in quietly.

WHY IT WAS MISSED — the general lesson, not the excuse. `run_comment()` sends
the wrapper's stdout AND stderr to `$OUTPUT_FILE`, and the `EXIT` trap deletes
`$WORK_DIR`. The suite therefore exits 1 with ZERO bytes on stdout and stderr,
and the one line that says what went wrong —

    Error: Gitea authenticated-identity read failed with HTTP 401

— lives only inside a directory that no longer exists when anyone looks. Every
oracle I had swept the family with greps for a SYMPTOM in surviving output, so
against this suite all of them returned "nothing found", which I read as "clean"
in the first sweep and as "unrelated pre-existing failure" in the second. A
suite that discards or deletes its own evidence converts a post-hoc assay into a
non-measurement, and I wrote that sentence into the previous commit while it was
already false about a file in the same directory.

HOW IT WAS ACTUALLY FOUND. Intercept the identity read at its SOURCE instead of
grepping for its consequence: a PATH shim over `git` that logs every
`mosaic.gitIdentity` read — args, rc, and resolved value — to a file OUTSIDE any
suite's work dir, then execs the real git. Deletion-proof by construction, and
it measures the defect's cause rather than one of its symptoms. Sweeping all 16
suites with it under an ordinary invocation:

  resolves a REAL identity (`mos-dt-0`) before the fix:
    test-issue-comment-readback          1 read   rc=1 (RED on every seat)
    test-pr-review-repo-host-override    6 reads  rc=0
    test-ci-queue-wait-branch-absent     3 reads  rc=0
  the four fixed in the previous commit now read empty; the rest never read at all.

The latter two are NOT affected and are deliberately left alone: under a seat
replica (identity set, no per-slot token) neither reaches `get_gitea_token`'s
fail-loud branch, and under a canary HOME neither carries the canary credential
into any surviving artifact. They read the identity and never enter a credential
path. That residual is structural and belongs to the wrapper half of #1007 —
scoping the read with `git -C "$repo"` removes it for everyone at once.

An earlier version of that sweep reported the four fixed suites as still
resolving a real identity. That was my grep, not the suites: `value=\[..*\]` is
satisfied by `value=[] args=[…]`, because `.*` runs past the empty pair and
matches the closing bracket of the NEXT one. `value=\[[^]]` is the correct test.
Recorded because the wrong pattern failed in the direction that would have sent
me re-fixing four already-correct files.

VERIFICATION of this suite, four HOME arms, all rc=0 with zero non-empty
identity reads and the pass line on stdout: real HOME, seat replica, canary
HOME, and an empty HOME with no identity at all. Full 16-suite sweep after the
change: every suite rc=0.

CONSEQUENCE FOR THE FINDING LIST IN THE PREVIOUS COMMIT: item 2 there — the
"silently red, unrelated to #1007" suite — is withdrawn. It was #1007 all along.
Item 1 (`pr-metadata.sh:89-92`, the anonymous fallback that reports an HTTP 200
carrying valid JSON as "unknown API error") stands and is still unfixed here.

Refs #1007
2026-07-31 07:27:10 -05:00
Jason Woltje 1afe2b36dc fix(git): #1007 suite hermeticity — pin repo-local mosaic.gitIdentity in four test suites
CENSUS CORRECTION: FOUR suites, not the three my own #1007 audit named. The
fourth (test-pr-metadata-gitea.sh) was outside the candidate set that audit
worked from and was found only by sweeping the discriminator across all 16
tools/git/test-*.sh suites. Recording that as a correction to my finding, not
as part of the original claim.

THE DEFECT. get_gitea_token() (detect-platform.sh:502-599) resolves a per-agent
identity at STEP 0, from `git config --get mosaic.gitIdentity`, BEFORE both the
Mosaic credential loader (step 1) and the GITEA_TOKEN env check (step 2). On a
provisioned agent seat that value is set GLOBALLY in ~/.gitconfig and is
inherited by any freshly-`git init`ed repo, so step 0 reads a REAL per-slot
token out of $HOME and returns it without ever consulting the suite's own
MOSAIC_CREDENTIALS_FILE / GITEA_TOKEN fixtures. The suites were running against
production credentials, and the fixture credential each one carefully
constructs was inert.

THE FIX: an empty repo-local `mosaic.gitIdentity`. An empty local value shadows
the global one and reads back empty at rc=0, so step 0 declines. The env route
does NOT work: detect-platform.sh reads "${MOSAIC_GIT_IDENTITY:-}", and `:-`
treats set-but-empty identically to unset.

OPERATIVE vs CONTAINMENT — the two mechanisms are not interchangeable and the
comment in each suite says so. The pin is operative: it prevents the resolution.
The sandboxed HOME each suite now also gets is containment: it bounds a failure
the pin should already have prevented. Conflating them is how this class stays
invisible, because a decoy HOME REMOVES the trigger (~/.gitconfig is where the
global identity lives), so any suite audited under one reads clean however
vulnerable it is. To MEASURE, replicate a seat: a decoy HOME whose .gitconfig
sets mosaic.gitIdentity with no per-slot token, so step 0 reaches its fail-loud
branch. That note is in each file for the next auditor.

SECOND, INDEPENDENT DEFECT in test-pr-metadata-gitea.sh. Applying the pin alone
turned that suite RED — and a control at baseline 826a8b3 under a plain HOME
reproduced the same failure, so it is pre-existing, not introduced. Its
`GITEA_TOKEN="stub-token"` / `GITEA_URL="https://git.example.test"` pair can
never satisfy step 2, because step 2 accepts GITEA_TOKEN only when GITEA_URL
matches the remote host and this repo's origin is git.uscllc.com. The suite had
therefore only ever passed by resolving a REAL credential — step 0 on a seat, or
step 1 from the operator's own credentials.json. A MOSAIC_CREDENTIALS_FILE
fixture is added rather than leaning on the sandboxed HOME making step 1 find
nothing: a test that passes because production configuration is ABSENT fails the
moment it is present. Shipping the pin without this would have moved the failure
rather than removed it.

NO CI ARM. .woodpecker/ci.yml does not run these suites; packages/mosaic/
package.json:28 (test:framework-shell) runs an ENUMERATED list that excludes all
four. They run only by hand — i.e. exclusively on a provisioned seat, the one
environment where the defect is live. "Passes in CI, fails on a seat" does not
apply here; there is no CI observation at all.

VERIFICATION (seat replica = decoy HOME with mosaic.gitIdentity set, no per-slot
token; canary = same plus a marked non-credential at both per-slot paths; plain
= empty HOME; real = ordinary invocation):
  - bash -n clean on all four.
  - Sweep of all 16 suites at baseline 826a8b3 under the seat replica:
    test-gitea-login-resolution rc=1 REACHES-STEP0; test-issue-create-
    interactive-auth rc=1 REACHES-STEP0; test-pr-merge-gitea-empty-uid rc=1
    REACHES-STEP0; test-pr-metadata-gitea rc=1 REACHES-STEP0.
  - Same sweep after: every row rc=0 with step0 absent.
  - test-gitea-token-identity flags REACHES-STEP0 in BOTH arms and is NOT a
    defect: it runs under `env -i HOME="$FAKE_HOME"` (line 77) and its hit is
    its own deliberate assert_failloud fixtures (lines 158-171). The fail-loud
    grep matches the intended behaviour as well as the defect, so it needs the
    second discriminator; recorded here so the next sweep does not re-file it.
  - Durable-argv assay (a PATH shim that tees argv out of each suite's own mock
    curl, because test-pr-merge-gitea-empty-uid truncates its log between phases
    and its EXIT trap removes the sandbox — a post-hoc read of that suite is a
    non-measurement, and "no trace" there is not a clearance):
      test-pr-merge-gitea-empty-uid  before: canary token in argv, fixture never
        used. after: fixture token in argv, canary absent. 5 curl calls both arms.
      test-pr-metadata-gitea         before: canary in argv. after: both calls
        carry the fixture token against git.uscllc.com.
  - test-pr-metadata-gitea across seat/canary/plain HOMEs after the fix: rc=0,
    rc=0, rc=0.
  - All four under the real HOME: rc=0. No regression to ordinary invocation.

The comment block is duplicated across the four files rather than pointing at a
shared note. Deliberate, and matching the merged #1006 precedent
(test-pr-review-gitea-comment.sh:87-95): the reader who needs it is auditing one
file.

TWO FINDINGS DELIBERATELY NOT FIXED HERE (out of this branch's scope, to be
filed):
  1. pr-metadata.sh:89-92 — the anonymous curl fallback does not check ^2, so an
     HTTP 200 carrying valid JSON is reported as "unknown API error" at rc=1.
  2. test-issue-comment-readback.sh exits 1 with ZERO bytes on stdout AND
     stderr, dying at its first seed_state python3 heredoc. Reproduces at
     baseline 826a8b3 under both a seat replica and the real HOME. Silently red
     at main for everyone; unrelated to #1007.

Refs #1007
2026-07-31 07:13:13 -05:00
jason.woltje 63e77887a8 feat: add mosaic-tools skill (fleet toolkit fast path) (#1) 2026-06-19 18:30:58 +00:00
Jarvis 809ca9a1d9 feat: add mosaic ops skills (portainer, gitea, woodpecker, deploy, orchestrator)
- mosaic-portainer: stack list/status/redeploy/logs via Portainer API scripts
- mosaic-gitea: PR/issue/milestone ops for git.mosaicstack.dev
- mosaic-woodpecker: pipeline status, trigger, CI wait
- mosaic-deploy: full end-to-end deploy flow (push → CI → merge → redeploy)
- mosaic-orchestrator: mission init/run/status + worker launch rules
2026-03-22 15:32:05 +00:00
Jason Woltje 57435bb879 switch skill docs to xdg mosaic config path 2026-02-17 14:12:07 -06:00
Jason Woltje 8c51bf7575 remove legacy nested skill link artifacts 2026-02-17 14:08:02 -06:00
Jason Woltje c3d2179ad8 add delegation mode fallback to matrix rail in kickstart 2026-02-17 14:05:43 -06:00
Jason Woltje b47c4024cc standardize skills to mosaic-first paths and docs 2026-02-17 13:09:06 -06:00
Jason Woltje 74d1cdc7c1 migrate kickstart skill to mosaic-first paths 2026-02-17 13:01:49 -06:00
Jason WoltjeandClaude Opus 4.6 8cae9e0883 feat: Add lint skill (zero-tolerance) + strengthen kickstart linting mandate
New skill: lint — zero-tolerance linting enforcement for all code changes.
Detects project linter, fixes ALL violations, never disables rules.

Updated kickstart: linting now explicit standing order #3 in worker template
with "NON-NEGOTIABLE" language and zero-tolerance enforcement.

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-02-16 17:13:00 -06:00
Jason WoltjeandClaude Opus 4.6 2c524b6da2 feat: Add kickstart skill — orchestrator launcher via /kickstart command
/kickstart [milestone|issue|task] replaces manual orchestrator boilerplate.
Auto-discovers project context, fetches issues from Gitea/GitHub, bootstraps
tracking files, and transforms the session into an orchestrator.

Modes: milestone, issue, task ID, resume, interactive (no args)
Built-in: quality gates, Two-Phase Completion, context handoff protocol

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-02-16 16:59:04 -06:00
Jason WoltjeandClaude Opus 4.6 b4f2019529 security: Remove vercel-deploy (data exfiltration), annotate LD_PRELOAD shims
Security audit findings:
- CRITICAL: vercel-deploy uploaded entire project to external endpoint — REMOVED
- ANNOTATED: docx/pptx/xlsx soffice.py LD_PRELOAD shims — security warnings added
- README updated to 93 skills with full security audit section and Vue/Vite ecosystem

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-02-16 16:39:04 -06:00
Jason WoltjeandClaude Opus 4.6 b1eb1fb2f9 feat: Complete fleet — 94 skills across 10+ domains
Pulled ALL skills from 15 source repositories:
- anthropics/skills: 16 (docs, design, MCP, testing)
- obra/superpowers: 14 (TDD, debugging, agents, planning)
- coreyhaines31/marketingskills: 25 (marketing, CRO, SEO, growth)
- better-auth/skills: 5 (auth patterns)
- vercel-labs/agent-skills: 5 (React, design, Vercel)
- antfu/skills: 16 (Vue, Vite, Vitest, pnpm, Turborepo)
- Plus 13 individual skills from various repos

Mosaic Stack is not limited to coding — the Orchestrator and
subagents serve coding, business, design, marketing, writing,
logistics, analysis, and more.

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-02-16 16:27:42 -06:00
Jason WoltjeandClaude Opus 4.6 dfeb4d9692 feat: Expand fleet to 23 skills across all domains
New skills (14):
- nestjs-best-practices: 40 priority-ranked rules (kadajett)
- fastapi: Pydantic v2, async SQLAlchemy, JWT auth (jezweb)
- architecture-patterns: Clean Architecture, Hexagonal, DDD (wshobson)
- python-performance-optimization: Profiling and optimization (wshobson)
- ai-sdk: Vercel AI SDK streaming and agent patterns (vercel)
- create-agent: Modular agent architecture with OpenRouter (openrouterteam)
- proactive-agent: WAL Protocol, compaction recovery, self-improvement (halthelobster)
- brand-guidelines: Brand identity enforcement (anthropics)
- ui-animation: Motion design with accessibility (mblode)
- marketing-ideas: 139 ideas across 14 categories (coreyhaines31)
- pricing-strategy: SaaS pricing and tier design (coreyhaines31)
- programmatic-seo: SEO at scale with playbooks (coreyhaines31)
- competitor-alternatives: Comparison page architecture (coreyhaines31)
- referral-program: Referral and affiliate programs (coreyhaines31)

README reorganized by domain: Code Quality, Frontend, Backend,
Auth, AI/Agent Building, Marketing, Design, Meta.

Mosaic Stack is not limited to coding — the Orchestrator serves
coding, business, design, marketing, writing, logistics, and analysis.

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-02-16 16:22:53 -06:00
Jason WoltjeandClaude Opus 4.6 7ea13332ed feat: Add 5 curated skills for Mosaic Stack
New skills:
- next-best-practices: Next.js 15+ RSC, async patterns, self-hosting (vercel-labs)
- better-auth-best-practices: Official Better-Auth with Drizzle adapter (better-auth)
- verification-before-completion: Evidence-based completion claims (obra/superpowers)
- shadcn-ui: Component patterns with Tailwind v4 adaptation note (developer-kit)
- writing-skills: TDD methodology for skill authoring (obra/superpowers)

README reorganized by category with Mosaic Stack alignment section.
Total: 9 skills (4 existing + 5 new).

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-02-16 16:17:40 -06:00
Jason WoltjeandClaude Opus 4.6 b032d23889 docs: Add npx install commands and clone instructions
- Add npx skills add commands for single, all, and non-interactive install
- Document .git suffix requirement for Gitea-hosted repos
- Add git clone step to manual installation
- Use ln -sf for idempotent symlinks

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-02-16 16:09:05 -06:00
Jason WoltjeandClaude Opus 4.6 ecde74439c feat: Initial agent-skills repo — 4 adapted skills for Mosaic Stack
Skills included:
- pr-reviewer: Adapted for Gitea/GitHub via platform-aware scripts
  (dropped fetch_pr_data.py and add_inline_comment.py, kept generate_review_files.py)
- code-review-excellence: Methodology and checklists (React, TS, Python, etc.)
- vercel-react-best-practices: 57 rules for React/Next.js performance
- tailwind-design-system: Tailwind CSS v4 patterns, CVA, design tokens

New shell scripts added to ~/.claude/scripts/git/:
- pr-diff.sh: Get PR diff (GitHub gh / Gitea API)
- pr-metadata.sh: Get PR metadata as normalized JSON

Co-Authored-By: Claude Opus 4.6 <[email protected]>
2026-02-16 16:03:39 -06:00
1585 changed files with 285647 additions and 1325 deletions
+2 -2
View File
@@ -22,9 +22,9 @@ steps:
image: gcr.io/kaniko-project/executor:debug
environment:
REGISTRY_USER:
from_secret: gitea_username
from_secret: REGISTRY_USERNAME
REGISTRY_PASS:
from_secret: gitea_password
from_secret: REGISTRY_PASSWORD
CI_COMMIT_BRANCH: ${CI_COMMIT_BRANCH}
CI_COMMIT_TAG: ${CI_COMMIT_TAG}
CI_COMMIT_SHA: ${CI_COMMIT_SHA}
+22
View File
@@ -59,6 +59,28 @@ steps:
# [0] of the pnpm chain, so severing that chain would silence it together
# with everything it guards; this direct line keeps one instrument running.
- bash packages/mosaic/framework/tools/quality/scripts/check-test-enumeration.sh
# Tool-index gate: a shipped wrapper that appears in no resident index doc
# is undiscoverable from inside a session, and an agent that cannot learn a
# wrapper exists reaches for raw curl instead — which is how a Gitea review
# got filed PENDING three times. Ships-and-documented is one commit, or red.
- bash packages/mosaic/framework/tools/quality/scripts/check-tools-index.sh --self-test
- bash packages/mosaic/framework/tools/quality/scripts/check-tools-index.sh
# Hermetic regression for issue-close.sh (#1081): mocks tea/curl onto PATH
# and sandboxes a throwaway git repo, so it resolves no real credentials and
# joins CI directly rather than the exclusions file.
- bash packages/mosaic/framework/tools/git/test-issue-close-fail-closed.sh
# Hermetic behavioural regression for the PreToolUse wrapper guard: proves
# it still blocks the three mistakes AND still lets reads, unwrapped
# endpoints and ordinary commands through. Both directions are asserted —
# a guard that over-blocks gets routed around, which fails just as hard.
- bash packages/mosaic/framework/tools/git/test-wrapper-guard.sh
# Hermetic regression for mosaic-worktree.sh at fleet scale: stubs git onto
# PATH so `list` faces ~450 KB of porcelain. The defect it pins is invisible
# at small size — `git … | awk '…exit'` gives the producer SIGPIPE, which
# under `set -euo pipefail` aborts the caller silently with rc=141 and no
# output. A repo only reaches that once it has enough worktrees, so the
# stub supplies the scale instead of the host's own checkout.
- bash packages/mosaic/framework/tools/git/test-mosaic-worktree-large-repo.sh
# Canonical verify:release stage `upgrade-guard`.
# Blocking gate (#791): a framework upgrade must never write or delete an
+6 -6
View File
@@ -270,9 +270,9 @@ steps:
when: *image_build_when
environment:
REGISTRY_USER:
from_secret: gitea_username
from_secret: REGISTRY_USERNAME
REGISTRY_PASS:
from_secret: gitea_password
from_secret: REGISTRY_PASSWORD
CI_COMMIT_BRANCH: ${CI_COMMIT_BRANCH}
CI_COMMIT_TAG: ${CI_COMMIT_TAG}
CI_COMMIT_SHA: ${CI_COMMIT_SHA}
@@ -306,9 +306,9 @@ steps:
when: *main_image_build_when
environment:
REGISTRY_USER:
from_secret: gitea_username
from_secret: REGISTRY_USERNAME
REGISTRY_PASS:
from_secret: gitea_password
from_secret: REGISTRY_PASSWORD
CI_COMMIT_BRANCH: ${CI_COMMIT_BRANCH}
CI_COMMIT_TAG: ${CI_COMMIT_TAG}
CI_COMMIT_SHA: ${CI_COMMIT_SHA}
@@ -333,9 +333,9 @@ steps:
when: *main_image_build_when
environment:
REGISTRY_USER:
from_secret: gitea_username
from_secret: REGISTRY_USERNAME
REGISTRY_PASS:
from_secret: gitea_password
from_secret: REGISTRY_PASSWORD
CI_COMMIT_BRANCH: ${CI_COMMIT_BRANCH}
CI_COMMIT_TAG: ${CI_COMMIT_TAG}
CI_COMMIT_SHA: ${CI_COMMIT_SHA}
+11 -3
View File
@@ -74,6 +74,14 @@ The launcher verifies your config, checks for `SOUL.md`, injects your `AGENTS.md
Pi launches default to a token-lean skill posture: `mosaic pi` passes `--no-skills` so Pi does not preload every global skill description into the system prompt. Use `MOSAIC_PI_SKILL_MODE=all mosaic pi` for the legacy all-skills catalog, or `MOSAIC_PI_SKILL_MODE=discover mosaic pi` to let Pi use its native settings/project skill discovery.
Mosaic also loads its Pi extensions from `~/.config/mosaic/runtime/pi/`. Inside Pi,
`/goal set <statement>` starts a bounded persistent loop that checks every turn and successful
compaction, requires two evidence-bearing completion reports, and can be inspected or stopped with
`/goal status`, `/goal pause`, `/goal resume`, and `/goal cancel`. Controller-owned goal-state
entries redact common credential shapes, but Pi's model/tool-call history is separate, so goals and
evidence must never contain secrets or raw sensitive output. Mosaic does not install this extension
into `~/.pi/agent/extensions/`.
### TUI & Gateway
```bash
@@ -138,9 +146,9 @@ mosaic brain tasks
mosaic brain conversations
# Agent forge pipeline
mosaic forge run
mosaic forge run [--simulate] # fails closed (FORGE_NO_EXECUTOR) with no executor wired; --simulate for typed simulated runs
mosaic forge status
mosaic forge resume
mosaic forge resume [--simulate] # same fail-closed rule as forge run
mosaic forge personas
# Structured logging
@@ -331,7 +339,7 @@ The framework is the bash-based standards layer installed to every developer mac
├── bin/mosaic ← Unified launcher (claude, codex, opencode, pi, yolo)
├── guides/ ← E2E delivery, orchestrator protocol, PRD, etc.
├── runtime/ ← Per-runtime configs (claude/, codex/, opencode/, pi/)
├── skills/ ← Universal skills (synced from agent-skills repo)
├── skills/ ← Universal skills (shipped with the framework package)
├── tools/ ← Tool suites (orchestrator, git, quality, prdy, etc.)
└── memory/ ← Persistent agent memory (preserved across upgrades)
```
@@ -190,7 +190,13 @@ beforeEach((ctx) => {
});
afterAll(async () => {
if (!handle) return;
// Cleanup only when the fixture actually installed rows. `handle` is set
// before the first query (createDb connects lazily), so on an unreachable
// database `handle` is truthy while nothing was inserted — cleanup must
// honor `dbAvailable` or the skip path fails the file with ECONNREFUSED in
// afterAll (caught live by the publish pipeline's no-DATABASE_URL verify
// step, pipeline 2486).
if (!handle || !dbAvailable) return;
const db = handle.db;
// Delete in dependency order (FK constraints)
@@ -245,9 +245,21 @@ describe('EnrollmentService.createToken', () => {
const after = Date.now();
const expiresMs = new Date(result.expiresAt).getTime();
// Should be at most 900s from now
expect(expiresMs - before).toBeLessThanOrEqual(900_000 + 100);
// The property under test is CLAMPING: a 9999s request must come back as 900s.
// The gap between clamped and unclamped is 9_099_000 ms, so the tolerance below
// only has to exceed CI scheduling jitter — it does not need to be tight to keep
// the assertion discriminating. A 5s allowance consumes 0.05% of that margin and
// an unclamped result still misses by three orders of magnitude.
//
// It was 100ms and failed on a loaded agent at 900_106 — 6ms over (#1090). A
// wall-clock budget sized to a fast machine is a flake, not a tighter test.
const CI_JITTER_MS = 5_000;
expect(expiresMs - before).toBeLessThanOrEqual(900_000 + CI_JITTER_MS);
expect(expiresMs - after).toBeGreaterThanOrEqual(0);
// Explicitly pin the clamp itself, independent of any timing allowance:
// unclamped (9999s) would exceed this by ~9_099_000 ms.
expect(expiresMs - before).toBeLessThan(1_000_000);
});
});
@@ -0,0 +1,110 @@
'use client';
import type { ReactElement } from 'react';
import { formatAge, type FreshnessLabel } from '@/lib/freshness/model';
/**
* Rendering rules for non-current freshness states (RI-5-001).
*
* - `unavailable` renders an explicit failure panel — never an empty
* healthy collection.
* - `stale` may render last-known data, but only under a visible label
* carrying source identity, snapshot version, and age.
* - `partial` renders the verified parts plus an explicit list of what is
* missing.
*/
interface RetryableNoticeProps {
readonly onRetry?: () => void;
readonly retryLabel?: string;
}
function RetryButton({ onRetry, retryLabel }: RetryableNoticeProps): ReactElement | null {
if (!onRetry) return null;
return (
<button
type="button"
onClick={onRetry}
className="mt-2 rounded-lg border border-surface-border px-3 py-1.5 text-xs transition-colors hover:border-gray-500"
>
{retryLabel ?? 'Retry'}
</button>
);
}
export interface UnavailableDataNoticeProps extends RetryableNoticeProps {
/** What is unavailable, e.g. "Tasks". */
readonly title: string;
/** Optional underlying failure detail (network message, invalidation reason). */
readonly detail?: string | null;
}
/** Explicit `unavailable` state. Never renders as an empty healthy collection. */
export function UnavailableDataNotice({
title,
detail,
onRetry,
retryLabel,
}: UnavailableDataNoticeProps): ReactElement {
return (
<div role="alert" className="rounded-lg border border-error/40 px-4 py-3 text-sm">
<p className="font-medium text-text-primary">{title} are unavailable</p>
<p className="mt-1 text-text-muted">
This is not an empty result the data could not be verified from the gateway.
{detail ? ` ${detail}` : ''}
</p>
<RetryButton onRetry={onRetry} retryLabel={retryLabel} />
</div>
);
}
export interface StaleDataNoticeProps extends RetryableNoticeProps {
/** Provenance of the last-known snapshot being displayed. */
readonly label: FreshnessLabel;
}
/**
* Situational-awareness banner for `stale` data: last-known data may render,
* but visibly labeled with source identity, snapshot version, and age.
*/
export function StaleDataNotice({
label,
onRetry,
retryLabel,
}: StaleDataNoticeProps): ReactElement {
return (
<div role="status" className="rounded-lg border border-warning/40 px-4 py-3 text-sm">
<p className="font-medium text-warning">Showing last-known data it may be out of date</p>
<p className="mt-1 text-xs text-text-muted">
Source {label.source} · snapshot v{label.version} · fetched{' '}
{formatAge(label.fetchedAt, Date.now())}. Verdicts derived from this data are unknown and
changes are disabled until it is revalidated.
</p>
<RetryButton onRetry={onRetry} retryLabel={retryLabel ?? 'Revalidate'} />
</div>
);
}
export interface PartialDataNoticeProps extends RetryableNoticeProps {
/** Display names of the sections whose collections are unavailable. */
readonly missing: readonly string[];
}
/** `partial` surface banner: verified parts render, missing parts are explicit. */
export function PartialDataNotice({
missing,
onRetry,
retryLabel,
}: PartialDataNoticeProps): ReactElement {
return (
<div role="status" className="rounded-lg border border-warning/40 px-4 py-3 text-sm">
<p className="font-medium text-warning">Some data could not be loaded</p>
<p className="mt-1 text-xs text-text-muted">
{missing.join(', ')} {missing.length === 1 ? 'is' : 'are'} unavailable sections below show
an explicit unavailable state instead of an empty list. Derived verdicts remain unknown
until every collection is revalidated.
</p>
<RetryButton onRetry={onRetry} retryLabel={retryLabel ?? 'Revalidate'} />
</div>
);
}
+324
View File
@@ -0,0 +1,324 @@
import { describe, expect, it } from 'vitest';
import type { Task } from '@/lib/types';
import {
acceptSnapshot,
assertMutable,
canMutate,
combineFreshness,
computeDigest,
computeFreshness,
DEFAULT_FRESHNESS_POLICY,
formatAge,
type FreshSnapshot,
invalidationReasonLabels,
StaleMutationError,
UNKNOWN_VERDICT,
verdictValue,
} from './model';
import { validateProjectCollection, validateTaskCollection } from './validators';
const NOW = 1_800_000_000_000;
const policy = { ...DEFAULT_FRESHNESS_POLICY, staleAfterMs: 60_000 };
const taskPayload: Task[] = [
{
id: 'task-1',
title: 'T1',
description: null,
status: 'not-started',
priority: 'high',
projectId: 'project-1',
missionId: null,
assignee: null,
tags: null,
dueDate: null,
metadata: null,
createdAt: '2026-08-01T00:00:00.000Z',
updatedAt: '2026-08-01T00:00:00.000Z',
},
];
function acceptedTaskSnapshot(
overrides: Partial<FreshSnapshot<typeof taskPayload>> = {},
): FreshSnapshot<typeof taskPayload> {
const result = acceptSnapshot({
value: taskPayload,
validate: validateTaskCollection,
previous: null,
policy,
source: 'gateway:/api/tasks',
now: NOW,
});
if (result.outcome !== 'accepted') {
throw new Error(`fixture setup failed: ${result.reason}`);
}
return { ...result.snapshot, ...overrides };
}
describe('computeFreshness', () => {
it('treats a missing snapshot as unavailable, never as an empty healthy collection', () => {
expect(computeFreshness({ snapshot: null, policy, now: NOW })).toBe('unavailable');
});
it('returns current for a fresh verified snapshot regardless of data emptiness', () => {
const empty = acceptSnapshot({
value: [],
validate: validateTaskCollection,
previous: null,
policy,
source: 'gateway:/api/tasks',
now: NOW,
});
if (empty.outcome !== 'accepted') throw new Error('expected acceptance');
expect(computeFreshness({ snapshot: empty.snapshot, policy, now: NOW })).toBe('current');
});
it('degrades to stale once the snapshot ages past staleAfterMs', () => {
const snapshot = acceptedTaskSnapshot();
expect(computeFreshness({ snapshot, policy, now: NOW + 60_001 })).toBe('stale');
expect(computeFreshness({ snapshot, policy, now: NOW + 59_999 })).toBe('current');
});
it('degrades to stale when the latest revalidation failed', () => {
const snapshot = acceptedTaskSnapshot();
expect(computeFreshness({ snapshot, policy, now: NOW, degraded: true })).toBe('stale');
});
});
describe('mutation guard', () => {
it('permits mutations only on current data', () => {
expect(canMutate('current')).toBe(true);
for (const state of ['stale', 'partial', 'unknown', 'unavailable'] as const) {
expect(canMutate(state)).toBe(false);
}
});
it('refuses mutations on non-current data via assertMutable', () => {
expect(() => assertMutable('current')).not.toThrow();
for (const state of ['stale', 'partial', 'unknown', 'unavailable'] as const) {
let thrown: unknown;
try {
assertMutable(state);
} catch (caught) {
thrown = caught;
}
expect(thrown).toBeInstanceOf(StaleMutationError);
expect(thrown).toBeInstanceOf(Error);
if (thrown instanceof StaleMutationError) {
expect(thrown.name).toBe('StaleMutationError');
expect(thrown.freshness).toBe(state);
expect(thrown.message).toContain(state);
expect(thrown.message).toContain('revalidat');
}
}
});
});
describe('acceptSnapshot', () => {
it('accepts a valid payload with provenance', () => {
const result = acceptSnapshot({
value: taskPayload,
validate: validateTaskCollection,
previous: null,
policy,
source: 'gateway:/api/tasks',
now: NOW,
});
expect(result.outcome).toBe('accepted');
if (result.outcome !== 'accepted') return;
expect(result.snapshot.source).toBe('gateway:/api/tasks');
expect(result.snapshot.version).toBe(1);
expect(result.snapshot.fetchedAt).toBe(NOW);
expect(result.snapshot.data).toEqual(taskPayload);
});
it('invalidates a schema-mismatched payload instead of rendering it', () => {
const result = acceptSnapshot({
value: { not: 'an array' },
validate: validateTaskCollection,
previous: acceptedTaskSnapshot(),
policy,
source: 'gateway:/api/tasks',
now: NOW,
});
expect(result).toEqual({ outcome: 'invalidated', reason: 'schema-mismatch' });
expect(invalidationReasonLabels['schema-mismatch']).toContain('schema');
});
it('invalidates cross-workspace payloads', () => {
const userOne = acceptSnapshot({
value: [
{
id: 'p1',
name: 'P1',
description: null,
status: 'active',
userId: 'user-1',
metadata: null,
createdAt: '2026-08-01T00:00:00.000Z',
updatedAt: '2026-08-01T00:00:00.000Z',
},
],
validate: validateProjectCollection,
previous: null,
policy,
source: 'gateway:/api/projects',
now: NOW,
});
if (userOne.outcome !== 'accepted') throw new Error('expected acceptance');
const switched = acceptSnapshot({
value: [
{
id: 'p9',
name: 'P9',
description: null,
status: 'active',
userId: 'user-2',
metadata: null,
createdAt: '2026-08-01T00:00:00.000Z',
updatedAt: '2026-08-01T00:00:00.000Z',
},
],
validate: validateProjectCollection,
previous: userOne.snapshot,
policy,
source: 'gateway:/api/projects',
now: NOW,
});
expect(switched).toEqual({ outcome: 'invalidated', reason: 'cross-workspace' });
});
it('keeps the previous workspace for collections with no intrinsic identity', () => {
const userOne = acceptSnapshot({
value: [
{
id: 'p1',
name: 'P1',
description: null,
status: 'active',
userId: 'user-1',
metadata: null,
createdAt: '2026-08-01T00:00:00.000Z',
updatedAt: '2026-08-01T00:00:00.000Z',
},
],
validate: validateProjectCollection,
previous: null,
policy,
source: 'gateway:/api/projects',
now: NOW,
});
if (userOne.outcome !== 'accepted') throw new Error('expected acceptance');
// Empty list after the user deleted every project: no identity to check,
// so the verified scope is retained and the empty state stays healthy.
const emptied = acceptSnapshot({
value: [],
validate: validateProjectCollection,
previous: userOne.snapshot,
policy,
source: 'gateway:/api/projects',
now: NOW,
});
expect(emptied.outcome).toBe('accepted');
if (emptied.outcome === 'accepted') {
expect(emptied.snapshot.data).toEqual([]);
expect(emptied.snapshot.workspace).toBe('user-1');
}
});
it('invalidates version regressions', () => {
const previous = acceptedTaskSnapshot({ version: 7 });
const regressed = acceptSnapshot({
value: taskPayload,
validate: validateTaskCollection,
previous,
policy,
source: 'gateway:/api/tasks',
now: NOW,
incomingVersion: 3,
});
expect(regressed).toEqual({ outcome: 'invalidated', reason: 'version-regression' });
const newerSchema = acceptedTaskSnapshot({ schemaVersion: 4 });
const downgradedClient = acceptSnapshot({
value: taskPayload,
validate: validateTaskCollection,
previous: newerSchema,
policy: { ...policy, schemaVersion: 2 },
source: 'gateway:/api/tasks',
now: NOW,
});
expect(downgradedClient).toEqual({ outcome: 'invalidated', reason: 'version-regression' });
});
it('increments the version monotonically across accepted snapshots', () => {
const first = acceptedTaskSnapshot();
const second = acceptSnapshot({
value: taskPayload,
validate: validateTaskCollection,
previous: first,
policy,
source: 'gateway:/api/tasks',
now: NOW,
});
expect(second.outcome).toBe('accepted');
if (second.outcome === 'accepted') {
expect(second.snapshot.version).toBe(first.version + 1);
}
});
});
describe('combineFreshness', () => {
it('gates the surface on the primary collection', () => {
expect(combineFreshness('unavailable', ['current'])).toBe('unavailable');
expect(combineFreshness('unknown', ['current'])).toBe('unknown');
expect(combineFreshness('current', [])).toBe('current');
});
it('degrades to partial when a secondary is unavailable', () => {
expect(combineFreshness('current', ['current', 'unavailable'])).toBe('partial');
});
it('degrades to unknown while a secondary is still loading', () => {
expect(combineFreshness('current', ['unknown'])).toBe('unknown');
});
it('degrades to stale when any collection is stale', () => {
expect(combineFreshness('current', ['stale'])).toBe('stale');
expect(combineFreshness('stale', ['current'])).toBe('stale');
});
it('propagates partial secondaries', () => {
expect(combineFreshness('current', ['partial'])).toBe('partial');
});
});
describe('computeDigest', () => {
it('is stable across key order and changes with data', () => {
const a = computeDigest({ x: 1, y: [1, 2] });
const b = computeDigest({ y: [1, 2], x: 1 });
expect(a).toBe(b);
expect(computeDigest({ x: 1, y: [1, 3] })).not.toBe(a);
});
});
describe('verdictValue', () => {
it('returns the value only for verified inputs', () => {
expect(verdictValue(true, '5')).toBe('5');
expect(verdictValue(false, '5')).toBe(UNKNOWN_VERDICT);
expect(verdictValue(false, '5')).not.toBe('5');
});
});
describe('formatAge', () => {
it('labels age in human terms', () => {
expect(formatAge(NOW, NOW)).toBe('just now');
expect(formatAge(NOW, NOW + 15_000)).toBe('under a minute ago');
expect(formatAge(NOW, NOW + 120_000)).toBe('2m ago');
expect(formatAge(NOW, NOW + 3 * 3_600_000)).toBe('3h ago');
expect(formatAge(NOW, NOW + 2 * 86_400_000)).toBe('2d ago');
});
});
+261
View File
@@ -0,0 +1,261 @@
/**
* Typed freshness model for gateway-fetched collections (RI-5-001).
*
* A failed or stale fetch must never be indistinguishable from an empty
* healthy collection. Every fetched surface carries an explicit freshness
* state, a verified snapshot identity (source, workspace, version, age), and
* a mutation guard that refuses state-changing operations unless the data is
* verified current.
*/
/** Freshness states for fetched data. Never inferred from emptiness. */
export type FreshnessState = 'current' | 'stale' | 'partial' | 'unknown' | 'unavailable';
/**
* Reasons a snapshot is invalidated. An invalidated snapshot is treated as
* unavailable and is never rendered as current.
*/
export type InvalidationReason =
| 'cache-corruption'
| 'cross-workspace'
| 'schema-mismatch'
| 'version-regression';
/** Human-readable labels for invalidation reasons (UI + error messages). */
export const invalidationReasonLabels: Record<InvalidationReason, string> = {
'cache-corruption': 'cached snapshot failed integrity checks',
'cross-workspace': 'data belongs to a different workspace',
'schema-mismatch': 'response did not match the expected schema',
'version-regression': 'snapshot version regressed below the accepted version',
};
/** A verified snapshot of fetched data with full provenance. */
export interface FreshSnapshot<T> {
readonly data: T;
/** Source identity of the fetch, e.g. `gateway:/api/tasks`. */
readonly source: string;
/** Workspace scope the data belongs to. */
readonly workspace: string;
/** Monotonic snapshot sequence number for this surface. */
readonly version: number;
/** Schema version of the validator that accepted this snapshot. */
readonly schemaVersion: number;
/** Epoch ms at which the data was verified. */
readonly fetchedAt: number;
/** Integrity digest of `data`, used to detect cache corruption. */
readonly digest: string;
}
/** Provenance label rendered next to last-known data. */
export interface FreshnessLabel {
readonly source: string;
readonly version: number;
readonly fetchedAt: number;
}
/** Policy governing freshness for a surface. */
export interface FreshnessPolicy {
/** Active workspace scope. Snapshots from other scopes are invalidated. */
readonly workspace: string;
/** Schema version of the current validator. */
readonly schemaVersion: number;
/** Age after which a verified snapshot degrades from current to stale. */
readonly staleAfterMs: number;
}
export const DEFAULT_FRESHNESS_POLICY: FreshnessPolicy = {
workspace: 'default',
schemaVersion: 1,
staleAfterMs: 60_000,
};
/** Payload returned by a successful schema validation. */
export interface FreshPayload<T> {
readonly data: T;
/**
* Workspace identity extracted from the payload itself when the collection
* carries one (e.g. a uniform `userId` on projects). `null` when the
* collection has no intrinsic workspace identity.
*/
readonly workspace: string | null;
}
/** Error thrown when a mutation is attempted on non-current data. */
export class StaleMutationError extends Error {
readonly freshness: FreshnessState;
constructor(freshness: FreshnessState) {
super(`Refused mutation on ${freshness} data: revalidation is required before mutating.`);
this.name = 'StaleMutationError';
this.freshness = freshness;
}
}
/** Stable JSON digest used for snapshot integrity checks. */
export function computeDigest(value: unknown): string {
// FNV-1a 32-bit over the stable JSON serialization. This is an integrity
// check against corruption, not a cryptographic guarantee.
let hash = 0x811c9dc5;
for (const byte of stableStringify(value)) {
hash ^= byte.charCodeAt(0);
hash = Math.imul(hash, 0x01000193) >>> 0;
}
return hash.toString(16).padStart(8, '0');
}
function stableStringify(value: unknown): string {
return serialize(value);
}
function serialize(value: unknown): string {
if (value === null || typeof value !== 'object') return JSON.stringify(value) ?? 'null';
if (Array.isArray(value)) return `[${value.map(serialize).join(',')}]`;
const entries = Object.entries(value as Record<string, unknown>)
.filter(([, item]) => item !== undefined)
.sort(([left], [right]) => (left < right ? -1 : left > right ? 1 : 0))
.map(([key, item]) => `${JSON.stringify(key)}:${serialize(item)}`);
return `{${entries.join(',')}}`;
}
export type AcceptSnapshotResult<T> =
| { readonly outcome: 'accepted'; readonly snapshot: FreshSnapshot<T> }
| { readonly outcome: 'invalidated'; readonly reason: InvalidationReason };
export interface AcceptSnapshotOptions<T> {
/** Raw fetched value (untrusted JSON). */
readonly value: unknown;
/** Schema validator; returns `null` when the value does not match. */
readonly validate: (value: unknown) => FreshPayload<T> | null;
/** Previously accepted snapshot for this surface, if any. */
readonly previous: FreshSnapshot<T> | null;
readonly policy: FreshnessPolicy;
readonly source: string;
/**
* Version carried by the incoming payload when the transport exposes one.
* Must not regress below the accepted snapshot's version.
*/
readonly incomingVersion?: number;
readonly now: number;
}
/**
* Validate and accept a fetched value as a snapshot, or invalidate it.
*
* Invalidation rules (each treated as unavailable, never rendered current):
* - schema mismatch: the payload fails validation
* - cross-workspace: the payload's workspace differs from the verified one
* - version regression: payload/schema version is below the accepted one
*/
export function acceptSnapshot<T>(options: AcceptSnapshotOptions<T>): AcceptSnapshotResult<T> {
const payload = options.validate(options.value);
if (payload === null) {
return { outcome: 'invalidated', reason: 'schema-mismatch' };
}
// Workspace identity: the payload's own scope wins; a collection with no
// intrinsic identity (e.g. an empty list after every project was deleted)
// keeps the previously verified scope rather than resetting to the policy
// default, so a legitimately empty response is not mistaken for a scope
// change.
const workspace = payload.workspace ?? options.previous?.workspace ?? options.policy.workspace;
if (options.previous !== null && options.previous.workspace !== workspace) {
return { outcome: 'invalidated', reason: 'cross-workspace' };
}
if (options.previous !== null && options.policy.schemaVersion < options.previous.schemaVersion) {
return { outcome: 'invalidated', reason: 'version-regression' };
}
if (
options.incomingVersion !== undefined &&
options.previous !== null &&
options.incomingVersion < options.previous.version
) {
return { outcome: 'invalidated', reason: 'version-regression' };
}
const snapshot: FreshSnapshot<T> = {
data: payload.data,
source: options.source,
workspace,
version: options.incomingVersion ?? (options.previous?.version ?? 0) + 1,
schemaVersion: options.policy.schemaVersion,
fetchedAt: options.now,
digest: computeDigest(payload.data),
};
return { outcome: 'accepted', snapshot };
}
export interface ComputeFreshnessOptions {
readonly snapshot: FreshSnapshot<unknown> | null;
readonly policy: FreshnessPolicy;
readonly now: number;
/**
* True when the snapshot cannot be trusted as current regardless of age:
* the latest revalidation failed, or the snapshot was restored from cache
* and has not been verified by a fetch in this session.
*/
readonly degraded?: boolean;
}
/**
* Compute the freshness state of a snapshot. A missing snapshot is
* `unavailable` (never "empty and healthy"); a degraded or aged snapshot is
* `stale` (situational awareness only).
*/
export function computeFreshness(options: ComputeFreshnessOptions): FreshnessState {
const { snapshot, policy, now, degraded = false } = options;
if (snapshot === null) return 'unavailable';
if (degraded) return 'stale';
if (now - snapshot.fetchedAt > policy.staleAfterMs) return 'stale';
return 'current';
}
/** Only verified-current data may back a state-changing action. */
export function canMutate(state: FreshnessState): boolean {
return state === 'current';
}
/** Defense in depth: reject the mutation call itself on non-current data. */
export function assertMutable(state: FreshnessState): void {
if (!canMutate(state)) {
throw new StaleMutationError(state);
}
}
/**
* Combine freshness across a multi-collection surface (primary + secondaries).
* The primary collection gates the surface: unknown while it loads,
* unavailable when it fails. Missing secondaries degrade the surface to
* `partial`; aged collections degrade it to `stale`.
*/
export function combineFreshness(
primary: FreshnessState,
secondaries: readonly FreshnessState[],
): FreshnessState {
if (primary === 'unavailable') return 'unavailable';
if (primary === 'unknown') return 'unknown';
if (secondaries.includes('unavailable')) return 'partial';
if (secondaries.includes('unknown')) return 'unknown';
if (secondaries.includes('stale') || primary === 'stale') return 'stale';
if (secondaries.includes('partial')) return 'partial';
return 'current';
}
/** Render-safe age label for snapshot provenance. */
export function formatAge(fetchedAt: number, now: number): string {
const ageMs = Math.max(0, now - fetchedAt);
if (ageMs < 10_000) return 'just now';
const minutes = Math.floor(ageMs / 60_000);
if (minutes < 1) return 'under a minute ago';
if (minutes < 60) return `${minutes}m ago`;
const hours = Math.floor(minutes / 60);
if (hours < 24) return `${hours}h ago`;
const days = Math.floor(hours / 24);
return `${days}d ago`;
}
/** Derived verdict placeholder for non-current inputs — never a green value. */
export const UNKNOWN_VERDICT = '?';
export function verdictValue(verified: boolean, value: string): string {
return verified ? value : UNKNOWN_VERDICT;
}
@@ -0,0 +1,197 @@
import { afterEach, beforeEach, describe, expect, it } from 'vitest';
import { acceptSnapshot, DEFAULT_FRESHNESS_POLICY } from './model';
import { clearSnapshotCache, readSnapshotCache, writeSnapshotCache } from './snapshot-cache';
import { validateProjectCollection, validateTaskCollection } from './validators';
import { projectFixtures, taskFixtures } from '@/spa/pages/page-fixtures';
import type { Project, Task } from '@/lib/types';
const KEY = 'test:tasks';
const NOW = 1_800_000_000_000;
const policy = { ...DEFAULT_FRESHNESS_POLICY, staleAfterMs: 60_000 };
function storedTaskSnapshot() {
const result = acceptSnapshot({
value: taskFixtures,
validate: validateTaskCollection,
previous: null,
policy,
source: 'gateway:/api/tasks',
now: NOW,
});
if (result.outcome !== 'accepted') throw new Error('fixture setup failed');
return result.snapshot;
}
function storedProjectSnapshot() {
const result = acceptSnapshot({
value: projectFixtures,
validate: validateProjectCollection,
previous: null,
policy,
source: 'gateway:/api/projects',
now: NOW,
});
if (result.outcome !== 'accepted') throw new Error('fixture setup failed');
return result.snapshot;
}
function readTasks() {
return readSnapshotCache({
key: KEY,
workspace: policy.workspace,
policy,
validate: validateTaskCollection,
});
}
/** Write an arbitrary value directly at the raw cache slot. */
function writeRaw(key: string, value: unknown): void {
sessionStorage.setItem(`mosaic:freshness:v1:${key}`, JSON.stringify(value));
}
/** Parse and re-write the stored entry (for tampering with internals). */
function tamperStored<T>(key: string, mutate: (stored: T) => void): void {
const parsed = JSON.parse(sessionStorage.getItem(`mosaic:freshness:v1:${key}`) ?? '{}') as T;
mutate(parsed);
writeRaw(key, parsed);
}
beforeEach(() => {
sessionStorage.clear();
});
afterEach(() => {
sessionStorage.clear();
});
describe('readSnapshotCache', () => {
it('misses when nothing is stored', () => {
expect(readTasks()).toEqual({ outcome: 'miss' });
});
it('hits for a well-formed entry and preserves provenance', () => {
const snapshot = storedTaskSnapshot();
writeSnapshotCache(KEY, snapshot);
const result = readTasks();
expect(result.outcome).toBe('hit');
if (result.outcome === 'hit') {
expect(result.snapshot.data).toEqual(taskFixtures);
expect(result.snapshot.source).toBe('gateway:/api/tasks');
expect(result.snapshot.version).toBe(snapshot.version);
expect(result.snapshot.fetchedAt).toBe(snapshot.fetchedAt);
expect(result.snapshot.workspace).toBe(snapshot.workspace);
}
});
it('invalidates unparsable entries as cache corruption', () => {
sessionStorage.setItem(`mosaic:freshness:v1:${KEY}`, '{not json');
expect(readTasks()).toEqual({ outcome: 'invalidated', reason: 'cache-corruption' });
});
it('invalidates structurally wrong entries as cache corruption', () => {
const malformed: unknown[] = [
'nested but not a snapshot',
{ data: taskFixtures }, // missing provenance fields
{
data: taskFixtures,
source: 1,
workspace: 'w',
version: 1,
schemaVersion: 1,
fetchedAt: 1,
digest: 'x',
},
null,
17,
];
for (const entry of malformed) {
writeRaw(KEY, entry);
expect(readTasks()).toEqual({ outcome: 'invalidated', reason: 'cache-corruption' });
}
});
it('invalidates digest mismatches as cache corruption (tampered data)', () => {
writeSnapshotCache(KEY, storedTaskSnapshot());
tamperStored<{ data: Task[] }>(KEY, (stored) => {
stored.data = [...stored.data, { ...stored.data[0]!, id: 'injected-task' }];
});
expect(readTasks()).toEqual({ outcome: 'invalidated', reason: 'cache-corruption' });
});
it('invalidates entries scoped to another workspace', () => {
const snapshot = storedTaskSnapshot();
writeSnapshotCache(KEY, { ...snapshot, workspace: 'someone-else' });
expect(readTasks()).toEqual({ outcome: 'invalidated', reason: 'cross-workspace' });
});
it('invalidates entries written by a newer schema as a version regression', () => {
const snapshot = storedTaskSnapshot();
writeSnapshotCache(KEY, { ...snapshot, schemaVersion: policy.schemaVersion + 1 });
expect(readTasks()).toEqual({ outcome: 'invalidated', reason: 'version-regression' });
});
it('invalidates entries whose data no longer validates (schema mismatch)', () => {
writeSnapshotCache(KEY, storedTaskSnapshot());
tamperStored<{ data: unknown }>(KEY, (stored) => {
stored.data = { malformed: true };
});
expect(readTasks()).toEqual({ outcome: 'invalidated', reason: 'schema-mismatch' });
});
it('never reports a corrupted raw entry as a hit (negative control)', () => {
for (const raw of ['{oops', 'null', '"string"', '[]', '12']) {
sessionStorage.setItem(`mosaic:freshness:v1:${KEY}`, raw);
const result = readTasks();
expect(result.outcome).not.toBe('hit');
expect(result.outcome).toBe('invalidated');
}
});
it('scopes project collections by their workspace identity', () => {
const snapshot = storedProjectSnapshot();
writeSnapshotCache('test:projects', snapshot);
const sameScope = readSnapshotCache({
key: 'test:projects',
workspace: 'user-1',
policy,
validate: validateProjectCollection,
});
expect(sameScope.outcome).toBe('hit');
const foreignScope = readSnapshotCache({
key: 'test:projects',
workspace: 'user-2',
policy,
validate: validateProjectCollection,
});
expect(foreignScope).toEqual({ outcome: 'invalidated', reason: 'cross-workspace' });
});
});
describe('writeSnapshotCache round-trip', () => {
it('round-trips an accepted project snapshot', () => {
const snapshot = storedProjectSnapshot();
writeSnapshotCache('test:projects', snapshot);
const result = readSnapshotCache({
key: 'test:projects',
workspace: snapshot.workspace,
policy,
validate: validateProjectCollection,
});
expect(result.outcome).toBe('hit');
if (result.outcome === 'hit') {
expect(result.snapshot.data).toEqual(projectFixtures as Project[]);
}
});
});
describe('clearSnapshotCache', () => {
it('drops the entry so the next read misses', () => {
writeSnapshotCache(KEY, storedTaskSnapshot());
expect(readTasks().outcome).toBe('hit');
clearSnapshotCache(KEY);
expect(readTasks()).toEqual({ outcome: 'miss' });
});
});
@@ -0,0 +1,154 @@
import {
computeDigest,
type FreshPayload,
type FreshSnapshot,
type FreshnessPolicy,
type InvalidationReason,
} from './model';
/**
* Session-scoped last-known snapshot cache (RI-5-001).
*
* Restored snapshots are situational awareness only: they surface as `stale`
* until a fetch re-verifies them. A cache entry that is corrupted, belongs to
* another workspace, was written by a newer schema, or no longer validates is
* invalidated (treated as unavailable, never rendered as current).
*/
const CACHE_PREFIX = 'mosaic:freshness:v1';
interface StoredSnapshot {
data: unknown;
source: string;
workspace: string;
version: number;
schemaVersion: number;
fetchedAt: number;
digest: string;
}
export type SnapshotCacheRead<T> =
| { readonly outcome: 'hit'; readonly snapshot: FreshSnapshot<T> }
| { readonly outcome: 'miss' }
| { readonly outcome: 'invalidated'; readonly reason: InvalidationReason };
export interface ReadSnapshotCacheOptions<T> {
readonly key: string;
readonly workspace: string;
readonly policy: FreshnessPolicy;
readonly validate: (value: unknown) => FreshPayload<T> | null;
}
function cacheKey(key: string): string {
return `${CACHE_PREFIX}:${key}`;
}
function isStoredSnapshot(value: unknown): value is StoredSnapshot {
if (typeof value !== 'object' || value === null) return false;
const candidate = value as Record<string, unknown>;
return (
typeof candidate['data'] === 'object' &&
candidate['data'] !== null &&
typeof candidate['source'] === 'string' &&
typeof candidate['workspace'] === 'string' &&
typeof candidate['version'] === 'number' &&
typeof candidate['schemaVersion'] === 'number' &&
typeof candidate['fetchedAt'] === 'number' &&
typeof candidate['digest'] === 'string'
);
}
function getStorage(): Storage | null {
try {
return globalThis.sessionStorage ?? null;
} catch {
return null;
}
}
/**
* Restore a cached snapshot under the active workspace scope. Every failure
* mode maps to an explicit invalidation reason or a miss — never to data
* that renders as current.
*/
export function readSnapshotCache<T>(options: ReadSnapshotCacheOptions<T>): SnapshotCacheRead<T> {
const storage = getStorage();
if (storage === null) return { outcome: 'miss' };
let raw: string | null;
try {
raw = storage.getItem(cacheKey(options.key));
} catch {
return { outcome: 'miss' };
}
if (raw === null) return { outcome: 'miss' };
let parsed: unknown;
try {
parsed = JSON.parse(raw);
} catch {
return { outcome: 'invalidated', reason: 'cache-corruption' };
}
if (!isStoredSnapshot(parsed)) {
return { outcome: 'invalidated', reason: 'cache-corruption' };
}
if (parsed.workspace !== options.workspace) {
return { outcome: 'invalidated', reason: 'cross-workspace' };
}
if (parsed.schemaVersion > options.policy.schemaVersion) {
// Written by a newer build than the running client: version regression.
return { outcome: 'invalidated', reason: 'version-regression' };
}
const payload = options.validate(parsed.data);
if (payload === null) {
return { outcome: 'invalidated', reason: 'schema-mismatch' };
}
if (computeDigest(payload.data) !== parsed.digest) {
return { outcome: 'invalidated', reason: 'cache-corruption' };
}
return {
outcome: 'hit',
snapshot: {
data: payload.data,
source: parsed.source,
workspace: parsed.workspace,
version: parsed.version,
schemaVersion: parsed.schemaVersion,
fetchedAt: parsed.fetchedAt,
digest: parsed.digest,
},
};
}
/** Persist a verified snapshot. Failures are non-fatal (cache is best-effort). */
export function writeSnapshotCache<T>(key: string, snapshot: FreshSnapshot<T>): void {
const storage = getStorage();
if (storage === null) return;
const stored: StoredSnapshot = {
data: snapshot.data,
source: snapshot.source,
workspace: snapshot.workspace,
version: snapshot.version,
schemaVersion: snapshot.schemaVersion,
fetchedAt: snapshot.fetchedAt,
digest: snapshot.digest,
};
try {
storage.setItem(cacheKey(key), JSON.stringify(stored));
} catch {
// Quota or serialization failures simply skip caching.
}
}
/** Drop a cached snapshot (used when a surface invalidates its cache entry). */
export function clearSnapshotCache(key: string): void {
const storage = getStorage();
if (storage === null) return;
try {
storage.removeItem(cacheKey(key));
} catch {
// Ignorable: a wedged storage entry is detected as corruption on read.
}
}
@@ -0,0 +1,372 @@
import { act } from 'react';
import { createRoot, type Root } from 'react-dom/client';
import { afterEach, beforeAll, beforeEach, describe, expect, it, vi } from 'vitest';
import type { Task } from '@/lib/types';
import { acceptSnapshot, StaleMutationError, DEFAULT_FRESHNESS_POLICY } from './model';
import type { FreshnessFailure } from './use-fresh-collection';
import {
describeFailure,
useFreshCollection,
type FreshCollection,
type UseFreshCollectionOptions,
} from './use-fresh-collection';
import { validateProjectCollection, validateTaskCollection } from './validators';
import { projectFixtures, taskFixtures } from '@/spa/pages/page-fixtures';
/**
* Failure-matrix coverage for the freshness seam (RI-5-001): network failure,
* auth failure, malformed response, cache corruption, stale age, schema
* mismatch, cross-workspace, recovery, and stale-action rejection — with
* negative controls proving no case yields current data or an enabled
* mutation.
*/
const NOW = 1_800_000_000_000;
interface Deferred<T> {
promise: Promise<T>;
resolve: (value: T) => void;
reject: (reason?: unknown) => void;
}
function createDeferred<T>(): Deferred<T> {
let resolve!: (value: T) => void;
let reject!: (reason?: unknown) => void;
const promise = new Promise<T>((res, rej) => {
resolve = res;
reject = rej;
});
return { promise, resolve, reject };
}
let root: Root | null = null;
let container: HTMLDivElement;
let latest: FreshCollection<Task[]> | null = null;
function Probe({
options,
}: {
options: UseFreshCollectionOptions<Task[]>;
}): React.ReactElement | null {
latest = useFreshCollection<Task[]>(options);
return null;
}
beforeAll(() => {
Object.defineProperty(globalThis, 'IS_REACT_ACT_ENVIRONMENT', {
configurable: true,
value: true,
});
});
beforeEach(() => {
sessionStorage.clear();
});
afterEach(async () => {
await act(async () => {
root?.unmount();
});
document.body.replaceChildren();
root = null;
latest = null;
sessionStorage.clear();
vi.restoreAllMocks();
});
async function renderCollection(
options: UseFreshCollectionOptions<Task[]>,
): Promise<FreshCollection<Task[]>> {
container = document.createElement('div');
document.body.append(container);
root = createRoot(container);
await act(async () => {
root?.render(<Probe options={options} />);
});
if (latest === null) throw new Error('hook did not run');
return latest;
}
function taskOptions(
overrides: Partial<UseFreshCollectionOptions<Task[]>> = {},
): UseFreshCollectionOptions<Task[]> {
return {
source: 'gateway:/api/tasks',
fetcher: () => Promise.resolve(taskFixtures),
validate: validateTaskCollection,
cacheKey: 'tasks',
clock: () => NOW,
...overrides,
};
}
function authError(statusCode: number): Error & { statusCode: number } {
return Object.assign(new Error(`Request failed with ${statusCode}`), { statusCode });
}
function seedCache(key: string): number {
const result = acceptSnapshot({
value: taskFixtures,
validate: validateTaskCollection,
previous: null,
policy: DEFAULT_FRESHNESS_POLICY,
source: 'gateway:/api/tasks',
now: NOW,
});
if (result.outcome !== 'accepted') throw new Error('fixture setup failed');
sessionStorage.setItem(`mosaic:freshness:v1:${key}`, JSON.stringify({ ...result.snapshot }));
return result.snapshot.version;
}
describe('useFreshCollection failure matrix', () => {
it('is unknown (not empty) while the first validation is in flight', async () => {
const deferred = createDeferred<Task[]>();
const collection = await renderCollection(taskOptions({ fetcher: () => deferred.promise }));
expect(collection.freshness).toBe('unknown');
expect(collection.validating).toBe(true);
expect(collection.data).toBeNull();
expect(collection.canMutate).toBe(false);
await act(async () => {
deferred.resolve(taskFixtures);
await deferred.promise;
});
});
it('becomes current with provenance after a verified fetch', async () => {
const collection = await renderCollection(taskOptions());
expect(collection.freshness).toBe('current');
expect(collection.data).toEqual(taskFixtures);
expect(collection.snapshot?.source).toBe('gateway:/api/tasks');
expect(collection.snapshot?.version).toBe(1);
expect(collection.failure).toBeNull();
expect(collection.canMutate).toBe(true);
// Verified snapshot is persisted for last-known restore.
expect(sessionStorage.getItem('mosaic:freshness:v1:tasks')).toBeTruthy();
});
it('treats a network failure as unavailable — never an empty healthy collection', async () => {
const collection = await renderCollection(
taskOptions({ fetcher: () => Promise.reject(new Error('network down')) }),
);
expect(collection.freshness).toBe('unavailable');
expect(collection.data).toBeNull();
expect(collection.failure).toEqual({ kind: 'fetch', message: 'network down' });
expect(collection.canMutate).toBe(false);
expect(describeFailure(collection.failure)).toBe('network down');
});
it('treats an auth failure as unavailable and drops the last-known snapshot', async () => {
let call = 0;
const collection = await renderCollection(
taskOptions({
fetcher: () => {
call += 1;
return call === 1 ? Promise.resolve(taskFixtures) : Promise.reject(authError(401));
},
}),
);
expect(collection.freshness).toBe('current');
await act(async () => {
await collection.revalidate();
});
expect(latest?.freshness).toBe('unavailable');
expect(latest?.data).toBeNull();
expect(latest?.failure?.kind).toBe('fetch');
// The previous user's data must not linger in the session cache.
expect(sessionStorage.getItem('mosaic:freshness:v1:tasks')).toBeNull();
});
it('invalidates a malformed response as a schema mismatch', async () => {
const collection = await renderCollection(
taskOptions({ fetcher: () => Promise.resolve({ malformed: true }) }),
);
expect(collection.freshness).toBe('unavailable');
expect(collection.data).toBeNull();
expect(collection.failure).toEqual({ kind: 'invalidated', reason: 'schema-mismatch' });
expect(collection.canMutate).toBe(false);
});
it('keeps the previous snapshot as labeled stale when a later payload mismatches', async () => {
let call = 0;
const collection = await renderCollection(
taskOptions({
fetcher: () => {
call += 1;
return call === 1 ? Promise.resolve(taskFixtures) : Promise.resolve('garbage');
},
}),
);
expect(collection.freshness).toBe('current');
await act(async () => {
await collection.revalidate();
});
expect(latest?.freshness).toBe('stale');
expect(latest?.data).toEqual(taskFixtures);
expect(latest?.failure).toEqual({ kind: 'invalidated', reason: 'schema-mismatch' });
expect(latest?.canMutate).toBe(false);
});
it('drops the snapshot when the workspace changes under it (cross-workspace)', async () => {
let call = 0;
const collection = await renderCollection(
taskOptions({
fetcher: () => {
call += 1;
return Promise.resolve(
call === 1 ? projectFixtures : [{ ...projectFixtures[0], userId: 'user-2' }],
);
},
validate: validateProjectCollection as unknown as (value: unknown) => {
data: Task[];
workspace: string | null;
},
source: 'gateway:/api/projects',
}),
);
expect(collection.freshness).toBe('current');
await act(async () => {
await collection.revalidate();
});
expect(latest?.freshness).toBe('unavailable');
expect(latest?.data).toBeNull();
expect(latest?.failure).toEqual({ kind: 'invalidated', reason: 'cross-workspace' });
});
it('ages from current to stale and refuses mutations on stale data', async () => {
let fakeNow = NOW;
const collection = await renderCollection(
taskOptions({
clock: () => fakeNow,
policy: { staleAfterMs: 40 },
tickMs: 10,
}),
);
expect(collection.freshness).toBe('current');
// Age the snapshot past the policy and let the tick recompute.
fakeNow = NOW + 60;
await act(async () => {
await new Promise((resolve) => setTimeout(resolve, 25));
});
expect(latest?.freshness).toBe('stale');
expect(latest?.data).toEqual(taskFixtures);
expect(latest?.canMutate).toBe(false);
const operation = vi.fn(async () => 'result');
await expect(latest?.mutate(operation)).rejects.toBeInstanceOf(StaleMutationError);
expect(operation).not.toHaveBeenCalled();
});
it('recovers to current after a successful revalidation', async () => {
let call = 0;
const collection = await renderCollection(
taskOptions({
fetcher: () => {
call += 1;
return call === 1
? Promise.reject(new Error('first attempt failed'))
: Promise.resolve(taskFixtures);
},
}),
);
expect(collection.freshness).toBe('unavailable');
await act(async () => {
await collection.revalidate();
});
expect(latest?.freshness).toBe('current');
expect(latest?.failure).toBeNull();
const operation = vi.fn(async (data: Task[]) => data.length);
await expect(latest?.mutate(operation)).resolves.toBe(taskFixtures.length);
expect(operation).toHaveBeenCalledOnce();
});
it('restores a cached snapshot as unverified stale data, then verifies it', async () => {
const seededVersion = seedCache('tasks');
const deferred = createDeferred<Task[]>();
const collection = await renderCollection(taskOptions({ fetcher: () => deferred.promise }));
// Restored data is situational awareness only: labeled stale, never
// current, and mutations are refused before verification.
expect(collection.freshness).toBe('stale');
expect(collection.data).toEqual(taskFixtures);
expect(collection.canMutate).toBe(false);
await expect(collection.mutate(vi.fn())).rejects.toBeInstanceOf(StaleMutationError);
await act(async () => {
deferred.resolve(taskFixtures);
await deferred.promise;
});
expect(latest?.freshness).toBe('current');
expect(latest?.snapshot?.version).toBe(seededVersion + 1);
});
it('never promotes corrupted cache data to current (cache corruption)', async () => {
sessionStorage.setItem('mosaic:freshness:v1:tasks', '{"data":');
const collection = await renderCollection(
taskOptions({ fetcher: () => Promise.reject(new Error('still down')) }),
);
expect(collection.freshness).toBe('unavailable');
expect(collection.data).toBeNull();
expect(collection.canMutate).toBe(false);
// The corrupted entry is dropped so it cannot come back.
expect(sessionStorage.getItem('mosaic:freshness:v1:tasks')).toBeNull();
});
it('refuses mutations while unknown or unavailable — the call itself, not just the button', async () => {
const deferred = createDeferred<Task[]>();
const unknown = await renderCollection(taskOptions({ fetcher: () => deferred.promise }));
const operation = vi.fn(async () => 'result');
await expect(unknown.mutate(operation)).rejects.toBeInstanceOf(StaleMutationError);
expect(operation).not.toHaveBeenCalled();
await act(async () => {
deferred.reject(new Error('failed'));
await deferred.promise.catch(() => undefined);
});
const unavailable = latest!;
await expect(unavailable.mutate(operation)).rejects.toBeInstanceOf(StaleMutationError);
expect(operation).not.toHaveBeenCalled();
expect(unavailable.canMutate).toBe(false);
});
it('degrades to stale with last-known data when a revalidation fails after success', async () => {
let call = 0;
const collection = await renderCollection(
taskOptions({
fetcher: () => {
call += 1;
return call === 1
? Promise.resolve(taskFixtures)
: Promise.reject(new Error('connection lost'));
},
}),
);
expect(collection.freshness).toBe('current');
await act(async () => {
await collection.revalidate();
});
expect(latest?.freshness).toBe('stale');
expect(latest?.data).toEqual(taskFixtures);
const failure: FreshnessFailure | null = latest?.failure ?? null;
expect(failure).toEqual({ kind: 'fetch', message: 'connection lost' });
});
});
@@ -0,0 +1,281 @@
import { useCallback, useEffect, useMemo, useRef, useState } from 'react';
import {
acceptSnapshot,
assertMutable,
computeFreshness,
DEFAULT_FRESHNESS_POLICY,
invalidationReasonLabels,
type FreshPayload,
type FreshSnapshot,
type FreshnessPolicy,
type FreshnessState,
type InvalidationReason,
StaleMutationError,
} from './model';
import { clearSnapshotCache, readSnapshotCache, writeSnapshotCache } from './snapshot-cache';
/**
* Freshness-aware collection fetch hook (RI-5-001).
*
* One hook owns one gateway collection end to end: fetch, schema validation,
* snapshot acceptance with provenance, session-scoped last-known caching,
* aging, and the mutation guard. Pages consume `freshness` and never infer
* health from emptiness.
*/
/** Why the latest validation did not produce a current snapshot. */
export type FreshnessFailure =
| { readonly kind: 'fetch'; readonly message: string }
| { readonly kind: 'invalidated'; readonly reason: InvalidationReason };
export interface UseFreshCollectionOptions<T> {
/** Source identity for provenance labels, e.g. `gateway:/api/tasks`. */
readonly source: string;
/** Performs the unvalidated fetch. The hook owns abort and verification. */
readonly fetcher: (signal: AbortSignal) => Promise<unknown>;
/**
* Runtime schema validator. Returning `null` invalidates the payload
* (`schema-mismatch`) instead of letting malformed JSON flow into render.
*/
readonly validate: (value: unknown) => FreshPayload<T> | null;
/** Overrides of the default freshness policy. */
readonly policy?: Partial<FreshnessPolicy>;
/**
* Session cache key for last-known snapshots. `null`/omitted disables
* restore. Restored snapshots are unverified: they render only as
* labeled `stale` data until a fetch re-verifies them.
*/
readonly cacheKey?: string | null;
/** Injectable clock for deterministic age transitions in tests. */
readonly clock?: () => number;
/** Aging tick interval override (default derived from `staleAfterMs`). */
readonly tickMs?: number;
/** When false, no fetch runs (surfaces stay `unavailable`/`unknown`). */
readonly enabled?: boolean;
}
export interface FreshCollection<T> {
/** Last verified (or restored-unverified) snapshot, or `null`. */
readonly snapshot: FreshSnapshot<T> | null;
/** Snapshot data or `null` — never a fabricated empty collection. */
readonly data: T | null;
readonly freshness: FreshnessState;
/** True while a validation request is in flight. */
readonly validating: boolean;
/** Outcome of the latest failed validation, `null` when healthy. */
readonly failure: FreshnessFailure | null;
/** False unless freshness is `current`; drives disabled UI affordances. */
readonly canMutate: boolean;
/** Re-run the fetch and re-verify. Always allowed (it is a read). */
readonly revalidate: () => Promise<void>;
/**
* Run a state-changing operation against verified-current data only.
* Rejects with `StaleMutationError` on any other state — the guard fires
* even if a disabled button was bypassed (defense in depth).
*/
readonly mutate: <R>(operation: (data: T) => Promise<R>) => Promise<R>;
}
const defaultClock = (): number => Date.now();
function resolveTickMs(policy: FreshnessPolicy, override?: number): number {
if (override !== undefined && override > 0) return override;
return Math.min(5_000, Math.max(250, Math.floor(policy.staleAfterMs / 4)));
}
function isAuthFailure(caught: unknown): boolean {
return (
typeof caught === 'object' &&
caught !== null &&
'statusCode' in caught &&
((caught as { statusCode?: unknown }).statusCode === 401 ||
(caught as { statusCode?: unknown }).statusCode === 403)
);
}
function fetchFailureMessage(caught: unknown): string {
if (caught instanceof Error && caught.message.trim().length > 0) return caught.message;
return 'The request failed.';
}
/** Human-readable summary of a failure for unavailable/stale notices. */
export function describeFailure(failure: FreshnessFailure | null): string | null {
if (failure === null) return null;
if (failure.kind === 'fetch') return failure.message;
return `The snapshot was invalidated: ${invalidationReasonLabels[failure.reason]}.`;
}
export function useFreshCollection<T>(options: UseFreshCollectionOptions<T>): FreshCollection<T> {
const optionsRef = useRef(options);
optionsRef.current = options;
const policy = useMemo<FreshnessPolicy>(
() => ({ ...DEFAULT_FRESHNESS_POLICY, ...options.policy }),
[options.policy],
);
const policyRef = useRef(policy);
policyRef.current = policy;
const clockRef = useRef(options.clock ?? defaultClock);
clockRef.current = options.clock ?? defaultClock;
const [snapshot, setSnapshot] = useState<FreshSnapshot<T> | null>(null);
const [failure, setFailure] = useState<FreshnessFailure | null>(null);
const [unverified, setUnverified] = useState(false);
const [validating, setValidating] = useState(options.enabled !== false);
const [now, setNow] = useState(() => (options.clock ?? defaultClock)());
const snapshotRef = useRef(snapshot);
snapshotRef.current = snapshot;
const failureRef = useRef(failure);
failureRef.current = failure;
const unverifiedRef = useRef(unverified);
unverifiedRef.current = unverified;
const runRef = useRef(0);
const abortRef = useRef<AbortController | null>(null);
const revalidate = useCallback(async (): Promise<void> => {
const current = optionsRef.current;
if (current.enabled === false) {
setValidating(false);
return;
}
const runId = ++runRef.current;
abortRef.current?.abort();
const controller = new AbortController();
abortRef.current = controller;
setValidating(true);
let value: unknown;
try {
value = await current.fetcher(controller.signal);
} catch (caught) {
if (runRef.current !== runId || controller.signal.aborted) return;
if (isAuthFailure(caught)) {
// An unauthenticated viewer must not keep (or be served) the
// previous user's last-known data.
setSnapshot(null);
setUnverified(false);
if (current.cacheKey) clearSnapshotCache(current.cacheKey);
}
setFailure({ kind: 'fetch', message: fetchFailureMessage(caught) });
setValidating(false);
return;
}
if (runRef.current !== runId) return;
const result = acceptSnapshot({
value,
validate: current.validate,
previous: snapshotRef.current,
policy: policyRef.current,
source: current.source,
now: clockRef.current(),
});
if (result.outcome === 'accepted') {
setSnapshot(result.snapshot);
setUnverified(false);
setFailure(null);
if (current.cacheKey) writeSnapshotCache(current.cacheKey, result.snapshot);
} else {
if (result.reason === 'cross-workspace') {
// Data verified for a different workspace must not linger as
// last-known situational awareness either.
setSnapshot(null);
setUnverified(false);
}
if (current.cacheKey) clearSnapshotCache(current.cacheKey);
setFailure({ kind: 'invalidated', reason: result.reason });
}
setValidating(false);
}, []);
// Restore the last-known snapshot (unverified) and run the first fetch.
useEffect(() => {
if (optionsRef.current.enabled === false) {
setValidating(false);
return;
}
const cacheKey = optionsRef.current.cacheKey;
if (cacheKey) {
const restored = readSnapshotCache<T>({
key: cacheKey,
workspace: policyRef.current.workspace,
policy: policyRef.current,
validate: optionsRef.current.validate,
});
if (restored.outcome === 'hit') {
setSnapshot(restored.snapshot);
setUnverified(true);
} else if (restored.outcome === 'invalidated') {
// A corrupted/foreign/regressed entry is dropped immediately; it must
// never surface as data. The fetch decides the visible state.
clearSnapshotCache(cacheKey);
}
}
void revalidate();
return () => {
abortRef.current?.abort();
};
// Mount-once by design: `revalidate` is stable and reads live options
// through refs, so it never needs to re-run when options change.
// Route-param pages remount this hook via an identity `key` instead.
}, [revalidate]);
// Aging tick: recomputes freshness as the snapshot ages past the policy.
useEffect(() => {
const interval = setInterval(
() => {
setNow(clockRef.current());
},
resolveTickMs(policyRef.current, optionsRef.current.tickMs),
);
return () => clearInterval(interval);
}, []);
const freshness = useMemo<FreshnessState>(() => {
if (snapshot === null) return validating ? 'unknown' : 'unavailable';
return computeFreshness({
snapshot,
policy,
now,
degraded: failure !== null || unverified,
});
// `now` from state covers age; refs inside computeFreshness are pure.
}, [snapshot, validating, failure, unverified, now, policy]);
const canMutate = freshness === 'current';
const mutate = useCallback(async <R>(operation: (data: T) => Promise<R>): Promise<R> => {
const currentSnapshot = snapshotRef.current;
// No verified snapshot at all: with nothing verified there is nothing
// current to mutate, regardless of the recorded failure.
if (currentSnapshot === null) throw new StaleMutationError('unavailable');
const state = computeFreshness({
snapshot: currentSnapshot,
policy: policyRef.current,
now: clockRef.current(),
degraded: failureRef.current !== null || unverifiedRef.current,
});
assertMutable(state);
return operation(currentSnapshot.data);
}, []);
return {
snapshot,
data: snapshot === null ? null : snapshot.data,
freshness,
validating,
failure,
canMutate,
revalidate,
mutate,
};
}
@@ -0,0 +1,103 @@
import { describe, expect, it } from 'vitest';
import type { Mission, Project, Task } from '@/lib/types';
import {
validateMissionCollection,
validateProjectCollection,
validateProjectEntity,
validateTaskCollection,
} from './validators';
import { missionFixtures, projectFixtures, taskFixtures } from '@/spa/pages/page-fixtures';
describe('validateTaskCollection', () => {
it('accepts a well-formed task collection', () => {
expect(validateTaskCollection(taskFixtures)).toEqual({
data: taskFixtures,
workspace: null,
});
});
it('accepts an empty collection (a healthy empty state is a valid payload)', () => {
expect(validateTaskCollection([])).toEqual({ data: [], workspace: null });
});
it.each([
['not an array', { items: [] }],
['item is not an object', ['nope']],
['missing id', [{ ...(taskFixtures[0] as Task), id: undefined }]],
['missing title', [{ ...(taskFixtures[0] as Task), title: undefined }]],
['unknown status enum', [{ ...(taskFixtures[0] as Task), status: 'finished' }]],
['unknown priority enum', [{ ...(taskFixtures[0] as Task), priority: 'urgent' }]],
['tags of the wrong type', [{ ...(taskFixtures[0] as Task), tags: 'spa' }]],
['metadata of the wrong type', [{ ...(taskFixtures[0] as Task), metadata: 'notes' }]],
['createdAt of the wrong type', [{ ...(taskFixtures[0] as Task), createdAt: 1234 }]],
['null sneaks past a required string', [{ ...(taskFixtures[0] as Task), title: null }]],
])('rejects a malformed payload: %s', (_label, value) => {
expect(validateTaskCollection(value)).toBeNull();
});
});
describe('validateMissionCollection', () => {
it('accepts a well-formed mission collection', () => {
expect(validateMissionCollection(missionFixtures)).toEqual({
data: missionFixtures,
workspace: null,
});
});
it.each([
['not an array', null],
['item missing name', [{ ...(missionFixtures[0] as Mission), name: 42 }]],
['unknown status enum', [{ ...(missionFixtures[0] as Mission), status: 'canceled' }]],
['projectId of the wrong type', [{ ...(missionFixtures[0] as Mission), projectId: 7 }]],
])('rejects a malformed payload: %s', (_label, value) => {
expect(validateMissionCollection(value)).toBeNull();
});
});
describe('validateProjectCollection', () => {
it('accepts a uniform workspace-scoped collection and reports its workspace', () => {
expect(validateProjectCollection(projectFixtures)).toEqual({
data: projectFixtures,
workspace: 'user-1',
});
});
it('accepts an empty collection with no workspace identity', () => {
expect(validateProjectCollection([])).toEqual({ data: [], workspace: null });
});
it.each([
['not an array', 42],
['item missing userId', [{ ...(projectFixtures[0] as Project), userId: undefined }]],
['unknown status enum', [{ ...(projectFixtures[0] as Project), status: 'live' }]],
['description of the wrong type', [{ ...(projectFixtures[0] as Project), description: 1 }]],
])('rejects a malformed payload: %s', (_label, value) => {
expect(validateProjectCollection(value)).toBeNull();
});
it('rejects a collection mixing workspace identities (cross-workspace leak)', () => {
const mixed = [
projectFixtures[0] as Project,
{ ...(projectFixtures[1] as Project), userId: 'user-2' },
];
expect(validateProjectCollection(mixed)).toBeNull();
});
});
describe('validateProjectEntity', () => {
it('accepts a well-formed project and reports its workspace', () => {
expect(validateProjectEntity(projectFixtures[0])).toEqual({
data: projectFixtures[0],
workspace: 'user-1',
});
});
it.each([
['not an object', 'project-1'],
['null', null],
['array', [projectFixtures[0]]],
['missing userId', [{ ...(projectFixtures[0] as Project), userId: null }]],
])('rejects a malformed entity: %s', (_label, value) => {
expect(validateProjectEntity(value)).toBeNull();
});
});
+135
View File
@@ -0,0 +1,135 @@
import type { Mission, Project, Task, MissionStatus, TaskPriority, TaskStatus } from '@/lib/types';
import type { FreshPayload } from './model';
/**
* Runtime schema validators for gateway collections (RI-5-001).
*
* `api<T>()` returns untrusted JSON cast to `T`; these validators are the
* seam where a malformed response becomes an explicit schema mismatch
* instead of flowing into the render path as if it were healthy data.
*/
const taskStatuses: readonly TaskStatus[] = [
'not-started',
'in-progress',
'blocked',
'done',
'cancelled',
];
const taskPriorities: readonly TaskPriority[] = ['critical', 'high', 'medium', 'low'];
const missionStatuses: readonly MissionStatus[] = [
'planning',
'active',
'paused',
'completed',
'failed',
];
const projectStatuses: readonly Project['status'][] = ['active', 'paused', 'completed', 'archived'];
function isRecord(value: unknown): value is Record<string, unknown> {
return typeof value === 'object' && value !== null && !Array.isArray(value);
}
function isString(value: unknown): value is string {
return typeof value === 'string';
}
function isNullableString(value: unknown): value is string | null {
return value === null || typeof value === 'string';
}
function isOneOf<T extends string>(value: unknown, allowed: readonly T[]): value is T {
return typeof value === 'string' && (allowed as readonly string[]).includes(value);
}
function isNullableRecord(value: unknown): value is Record<string, unknown> | null {
return value === null || isRecord(value);
}
function isNullableStringArray(value: unknown): value is string[] | null {
if (value === null) return true;
if (!Array.isArray(value)) return false;
return value.every((item) => typeof item === 'string');
}
function isIsoLike(value: unknown): value is string {
return typeof value === 'string' && value.length > 0;
}
function isTask(value: unknown): value is Task {
if (!isRecord(value)) return false;
return (
isString(value['id']) &&
isString(value['title']) &&
isOneOf(value['status'], taskStatuses) &&
isOneOf(value['priority'], taskPriorities) &&
isNullableString(value['projectId']) &&
isNullableString(value['missionId']) &&
isNullableString(value['assignee']) &&
isNullableStringArray(value['tags']) &&
isNullableRecord(value['metadata']) &&
isNullableString(value['dueDate']) &&
isIsoLike(value['createdAt']) &&
isIsoLike(value['updatedAt'])
);
}
/** Tasks carry no workspace identity; scope falls back to the policy. */
export function validateTaskCollection(value: unknown): FreshPayload<Task[]> | null {
if (!Array.isArray(value) || !value.every(isTask)) return null;
return { data: value as Task[], workspace: null };
}
function isMission(value: unknown): value is Mission {
if (!isRecord(value)) return false;
return (
isString(value['id']) &&
isString(value['name']) &&
isOneOf(value['status'], missionStatuses) &&
isNullableString(value['projectId']) &&
isNullableString(value['description']) &&
isNullableRecord(value['metadata']) &&
isIsoLike(value['createdAt']) &&
isIsoLike(value['updatedAt'])
);
}
/** Missions carry no workspace identity; scope falls back to the policy. */
export function validateMissionCollection(value: unknown): FreshPayload<Mission[]> | null {
if (!Array.isArray(value) || !value.every(isMission)) return null;
return { data: value as Mission[], workspace: null };
}
function isProject(value: unknown): value is Project {
if (!isRecord(value)) return false;
return (
isString(value['id']) &&
isString(value['name']) &&
isOneOf(value['status'], projectStatuses) &&
isString(value['userId']) &&
isNullableString(value['description']) &&
isNullableRecord(value['metadata']) &&
isIsoLike(value['createdAt']) &&
isIsoLike(value['updatedAt'])
);
}
/**
* Projects are workspace-scoped: every item must carry the same `userId`.
* A collection mixing identities (cross-workspace leak) is a schema
* mismatch; the uniform `userId` becomes the snapshot workspace.
*/
export function validateProjectCollection(value: unknown): FreshPayload<Project[]> | null {
if (!Array.isArray(value) || !value.every(isProject)) return null;
const projects = value as Project[];
const workspaces = new Set(projects.map((project) => project.userId));
if (workspaces.size > 1) return null;
return { data: projects, workspace: projects.length > 0 ? projects[0]!.userId : null };
}
/** Single project entity (project detail primary collection). */
export function validateProjectEntity(value: unknown): FreshPayload<Project> | null {
if (!isProject(value)) return null;
const project = value as Project;
return { data: project, workspace: project.userId };
}
+171 -27
View File
@@ -35,6 +35,7 @@ afterEach(async () => {
document.body.replaceChildren();
root = null;
apiMock.mockReset();
sessionStorage.clear();
});
async function renderProjectDetailPage(): Promise<ReturnType<typeof createMemoryRouter>> {
@@ -64,21 +65,49 @@ function clickButtonByText(text: string): void {
button.dispatchEvent(new MouseEvent('click', { bubbles: true }));
}
async function flushAct(): Promise<void> {
await act(async () => {
await Promise.resolve();
});
}
interface Deferred<T> {
promise: Promise<T>;
resolve: (value: T) => void;
}
function createDeferred<T>(): Deferred<T> {
let resolve!: (value: T) => void;
const promise = new Promise<T>((res) => {
resolve = res;
});
return { promise, resolve };
}
const projectOneTasks = taskFixtures.filter((task) => task.projectId === 'project-1');
function mockHealthyLoad(): void {
apiMock
.mockResolvedValueOnce(projectFixtures[0])
.mockResolvedValueOnce(missionFixtures)
.mockResolvedValueOnce(projectOneTasks);
}
describe('ProjectDetailPage', () => {
it('loads the project, tasks, missions, and optional PRD content for the active project', async () => {
apiMock
.mockResolvedValueOnce(projectFixtures[0])
.mockResolvedValueOnce(missionFixtures)
.mockResolvedValueOnce(taskFixtures.filter((task) => task.projectId === 'project-1'));
mockHealthyLoad();
await renderProjectDetailPage();
expect(apiMock.mock.calls).toEqual([
['/api/projects/project-1'],
['/api/missions'],
['/api/tasks?projectId=project-1'],
expect(apiMock.mock.calls.map((call) => call[0])).toEqual([
'/api/projects/project-1',
'/api/missions',
'/api/tasks?projectId=project-1',
]);
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'current',
);
expect(container.textContent).toContain('Mosaic Stack');
expect(container.textContent).toContain('Route /projects/:id');
expect(container.textContent).toContain('Tasks');
@@ -101,10 +130,7 @@ describe('ProjectDetailPage', () => {
});
it('opens and closes the existing read-only task modal from the tasks tab', async () => {
apiMock
.mockResolvedValueOnce(projectFixtures[0])
.mockResolvedValueOnce(missionFixtures)
.mockResolvedValueOnce(taskFixtures.filter((task) => task.projectId === 'project-1'));
mockHealthyLoad();
await renderProjectDetailPage();
@@ -134,35 +160,153 @@ describe('ProjectDetailPage', () => {
expect(container.querySelector('[role="dialog"]')).toBeNull();
});
it('renders the project with an empty missions tab when the missions request fails', async () => {
apiMock
.mockResolvedValueOnce(projectFixtures[0])
.mockRejectedValueOnce(new Error('Missions request failed'))
.mockResolvedValueOnce(taskFixtures.filter((task) => task.projectId === 'project-1'));
it('shows verified completion verdicts when the task collection is current', async () => {
mockHealthyLoad();
await renderProjectDetailPage();
expect(container.textContent).toContain('Mosaic Stack');
expect(container.querySelector('[role="alert"]')).toBeNull();
await act(async () => {
clickButtonByText('Missions (0)');
});
expect(container.textContent).toContain('No missions for this project');
const doneCard = [...container.querySelectorAll('div')].find(
(candidate) => candidate.textContent === 'Done1',
);
expect(doneCard).toBeTruthy();
const inProgressCard = [...container.querySelectorAll('div')].find(
(candidate) => candidate.textContent === 'In Progress1',
);
expect(inProgressCard).toBeTruthy();
});
it('renders a visible alert when the project request fails and lets the user navigate back', async () => {
it('renders an explicit unavailable missions tab when the missions request fails (partial, not empty)', async () => {
apiMock
.mockResolvedValueOnce(projectFixtures[0])
.mockRejectedValueOnce(new Error('Missions request failed'))
.mockResolvedValueOnce(projectOneTasks);
await renderProjectDetailPage();
// Secondary failure degrades the surface to partial; the project itself
// still renders.
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'partial',
);
expect(container.textContent).toContain('Mosaic Stack');
const partial = container.querySelector('[role="status"]');
expect(partial?.textContent).toContain('Missions');
expect(partial?.textContent).toContain('unavailable');
await act(async () => {
clickButtonByText('Missions (?)');
});
const alert = container.querySelector('[role="alert"]');
expect(alert?.textContent).toContain('Missions request failed');
// Negative control: a failed fetch must not look like an empty list.
expect(container.textContent).not.toContain('No missions for this project');
});
it('marks derived verdicts unknown when the tasks collection is unavailable', async () => {
apiMock
.mockResolvedValueOnce(projectFixtures[0])
.mockResolvedValueOnce(missionFixtures)
.mockRejectedValueOnce(new Error('Tasks request failed'));
await renderProjectDetailPage();
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'partial',
);
// Completion verdicts become unknown ('?') — never green counts.
for (const label of ['Done', 'In Progress', 'Blocked', 'Tasks']) {
const unknownCard = [...container.querySelectorAll('div')].find(
(candidate) => candidate.textContent === `${label}?`,
);
expect(unknownCard, `expected ${label} card to render ?`).toBeTruthy();
}
// Negative control: no green "Done 1" verdict anywhere.
expect(
[...container.querySelectorAll('div')].some((candidate) => candidate.textContent === 'Done1'),
).toBe(false);
await act(async () => {
clickButtonByText('Tasks (?)');
});
const alert = container.querySelector('[role="alert"]');
expect(alert?.textContent).toContain('Tasks request failed');
// Negative control: no healthy empty task list from a failed fetch.
expect(container.textContent).not.toContain('No tasks found');
expect(container.querySelector('table')).toBeNull();
});
it('recovers a partial surface to current after revalidation', async () => {
apiMock
.mockResolvedValueOnce(projectFixtures[0])
.mockResolvedValueOnce(missionFixtures)
.mockRejectedValueOnce(new Error('Tasks request failed'))
.mockResolvedValueOnce(projectFixtures[0])
.mockResolvedValueOnce(missionFixtures)
.mockResolvedValueOnce(projectOneTasks);
await renderProjectDetailPage();
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'partial',
);
await act(async () => {
clickButtonByText('Revalidate');
});
await flushAct();
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'current',
);
expect(
[...container.querySelectorAll('div')].some((candidate) => candidate.textContent === 'Done1'),
).toBe(true);
});
it("never shows one project's data on another project's route after navigation", async () => {
mockHealthyLoad();
const router = await renderProjectDetailPage();
expect(container.textContent).toContain('Mosaic Stack');
const deferred = createDeferred<(typeof projectFixtures)[number]>();
apiMock
.mockResolvedValueOnce(deferred.promise)
.mockResolvedValueOnce([])
.mockResolvedValueOnce([]);
await act(async () => {
await router.navigate('/projects/project-2');
});
// While project-2 loads, nothing from project-1 may render on its route.
expect(container.textContent).toContain('Loading project...');
expect(container.textContent).not.toContain('Mosaic Stack');
expect(container.textContent).not.toContain('Route /projects/:id');
await act(async () => {
deferred.resolve(projectFixtures[1]!);
await deferred.promise;
});
expect(container.textContent).toContain('Agent Runtime');
expect(apiMock.mock.calls[3]?.[0]).toBe('/api/projects/project-2');
});
it('renders a visible unavailable state when the project request fails and lets the user navigate back', async () => {
apiMock
.mockRejectedValueOnce(new Error('Project request failed'))
.mockResolvedValueOnce(missionFixtures)
.mockResolvedValueOnce(taskFixtures.filter((task) => task.projectId === 'project-1'));
.mockResolvedValueOnce(projectOneTasks);
const router = await renderProjectDetailPage();
const alert = container.querySelector('[role="alert"]');
expect(alert).toBeTruthy();
expect(alert?.textContent).toContain('Project request failed');
expect(alert?.textContent).toContain('not an empty result');
expect(container.textContent).not.toContain('Mosaic Stack');
await act(async () => {
+203 -90
View File
@@ -1,14 +1,30 @@
import { useEffect, useState, type ReactElement } from 'react';
import { useState, type ReactElement } from 'react';
import { useNavigate, useParams } from 'react-router-dom';
import { MissionTimeline } from '@/components/projects/mission-timeline';
import { PrdViewer } from '@/components/projects/prd-viewer';
import { TaskDetailModal } from '@/components/tasks/task-detail-modal';
import { TaskListView } from '@/components/tasks/task-list-view';
import { TaskStatusSummary } from '@/components/tasks/task-status-summary';
import {
PartialDataNotice,
StaleDataNotice,
UnavailableDataNotice,
} from '@/components/freshness/freshness-notices';
import { api } from '@/lib/api';
import { cn } from '@/lib/cn';
import type { Mission, Project, Task, TaskStatus } from '@/lib/types';
import { getErrorMessage } from './page-errors';
import {
combineFreshness,
UNKNOWN_VERDICT,
verdictValue,
type FreshSnapshot,
} from '@/lib/freshness/model';
import { describeFailure, useFreshCollection } from '@/lib/freshness/use-fresh-collection';
import {
validateMissionCollection,
validateProjectEntity,
validateTaskCollection,
} from '@/lib/freshness/validators';
type Tab = 'overview' | 'tasks' | 'missions' | 'prd';
@@ -51,73 +67,62 @@ function TabButton({ id, label, activeTab, onClick }: TabButtonProps): ReactElem
);
}
/** Remounts per project id so no state from one project renders for another. */
export function ProjectDetailPage(): ReactElement {
const { id = '' } = useParams();
return <ProjectDetail id={id} key={id} />;
}
function ProjectDetail({ id }: { id: string }): ReactElement {
const navigate = useNavigate();
const [project, setProject] = useState<Project | null>(null);
const [missions, setMissions] = useState<Mission[]>([]);
const [tasks, setTasks] = useState<Task[]>([]);
const [loading, setLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
const enabled = id.length > 0;
// Primary collection gates the surface; missions and tasks are secondaries
// whose failures degrade the surface to `partial` instead of rendering
// empty healthy lists.
const project = useFreshCollection<Project>({
source: `gateway:/api/projects/${id}`,
fetcher: (signal) => api<unknown>(`/api/projects/${id}`, { signal }),
validate: validateProjectEntity,
// No last-known restore: the entity carries workspace identity that
// cannot be scope-checked before display (see ProjectsPage note).
enabled,
});
const missions = useFreshCollection<Mission[]>({
source: 'gateway:/api/missions',
fetcher: (signal) => api<unknown>('/api/missions', { signal }),
validate: validateMissionCollection,
cacheKey: enabled ? 'missions' : null,
enabled,
});
const tasks = useFreshCollection<Task[]>({
source: `gateway:/api/tasks?projectId=${id}`,
fetcher: (signal) => api<unknown>(`/api/tasks?projectId=${id}`, { signal }),
validate: validateTaskCollection,
cacheKey: enabled ? `project-tasks:${id}` : null,
enabled,
});
const [activeTab, setActiveTab] = useState<Tab>('overview');
const [taskFilter, setTaskFilter] = useState<TaskStatus | 'all'>('all');
const [selectedTask, setSelectedTask] = useState<Task | null>(null);
useEffect(() => {
if (!id) {
setError('Project id is missing.');
setLoading(false);
return;
}
const surface = combineFreshness(project.freshness, [missions.freshness, tasks.freshness]);
const tasksVerified = tasks.freshness === 'current';
const projectMissions = missions.data?.filter((mission) => mission.projectId === id) ?? null;
let cancelled = false;
setLoading(true);
setError(null);
const retryAll = (): void => {
void Promise.all([project.revalidate(), missions.revalidate(), tasks.revalidate()]);
};
void Promise.all([
api<Project>('/api/projects/' + id),
api<Mission[]>('/api/missions').catch(() => [] as Mission[]),
api<Task[]>('/api/tasks?projectId=' + id).catch(() => [] as Task[]),
])
.then(([loadedProject, allMissions, loadedTasks]) => {
if (cancelled) return;
setProject(loadedProject);
setMissions(allMissions.filter((mission) => mission.projectId === id));
setTasks(loadedTasks);
})
.catch((caught: unknown) => {
if (cancelled) return;
setError(getErrorMessage(caught, 'Failed to load project.'));
})
.finally(() => {
if (cancelled) return;
setLoading(false);
});
return () => {
cancelled = true;
};
}, [id]);
if (loading) {
return (
<div className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<header className="mb-6 border-b px-1 pb-3">
<h1 className="text-2xl font-semibold">Project</h1>
</header>
<p className="py-16 text-center text-sm text-text-muted">Loading project...</p>
</div>
);
}
if (error || !project) {
if (!enabled) {
return (
<div className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<header className="mb-6 border-b px-1 pb-3">
<h1 className="text-2xl font-semibold">Project</h1>
</header>
<div role="alert" className="rounded-lg border border-error/40 px-4 py-3 text-sm">
{error ?? 'Project not found.'}
Project id is missing.
</div>
<button
type="button"
@@ -130,18 +135,81 @@ export function ProjectDetailPage(): ReactElement {
);
}
if (project.freshness === 'unknown') {
return (
<div className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<header className="mb-6 border-b px-1 pb-3">
<h1 className="text-2xl font-semibold">Project</h1>
</header>
<p className="py-16 text-center text-sm text-text-muted">Loading project...</p>
</div>
);
}
if (project.freshness === 'unavailable' || project.data === null) {
return (
<div className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<header className="mb-6 border-b px-1 pb-3">
<h1 className="text-2xl font-semibold">Project</h1>
</header>
<UnavailableDataNotice
title="This project"
detail={describeFailure(project.failure)}
onRetry={retryAll}
/>
<button
type="button"
onClick={() => navigate('/projects')}
className="mt-4 w-fit text-sm underline"
>
Back to projects
</button>
</div>
);
}
const projectTasks = tasks.data ?? null;
const filteredTasks =
taskFilter === 'all' ? tasks : tasks.filter((task) => task.status === taskFilter);
const prdContent = getPrdContent(project);
projectTasks === null
? []
: taskFilter === 'all'
? projectTasks
: projectTasks.filter((task) => task.status === taskFilter);
// Derived completion verdicts: unknown (never green) unless the task
// collection is verified current.
const doneCount = projectTasks?.filter((task) => task.status === 'done').length ?? 0;
const inProgressCount = projectTasks?.filter((task) => task.status === 'in-progress').length ?? 0;
const blockedCount = projectTasks?.filter((task) => task.status === 'blocked').length ?? 0;
const prdContent = getPrdContent(project.data);
const tabs: Array<{ id: Tab; label: string }> = [
{ id: 'overview', label: 'Overview' },
{ id: 'tasks', label: `Tasks (${tasks.length})` },
{ id: 'missions', label: `Missions (${missions.length})` },
{
id: 'tasks',
label: `Tasks (${projectTasks === null ? UNKNOWN_VERDICT : projectTasks.length})`,
},
{
id: 'missions',
label: `Missions (${projectMissions === null ? UNKNOWN_VERDICT : projectMissions.length})`,
},
...(prdContent ? [{ id: 'prd' as const, label: 'PRD' }] : []),
];
const staleSnapshot: FreshSnapshot<unknown> | null =
project.freshness === 'stale'
? project.snapshot
: missions.freshness === 'stale'
? missions.snapshot
: tasks.freshness === 'stale'
? tasks.snapshot
: null;
const missingSections: string[] = [];
if (missions.freshness === 'unavailable') missingSections.push('Missions');
if (tasks.freshness === 'unavailable') missingSections.push('Tasks');
return (
<div className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<div data-freshness={surface} className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<header className="mb-6 border-b px-1 pb-3">
<nav className="mb-4 flex items-center gap-2 text-sm text-text-muted">
<button
@@ -152,49 +220,64 @@ export function ProjectDetailPage(): ReactElement {
Projects
</button>
<span>/</span>
<span className="text-text-primary">{project.name}</span>
<span className="text-text-primary">{project.data.name}</span>
</nav>
<div className="flex items-start justify-between gap-4">
<div>
<div className="flex items-center gap-3">
<h1 className="text-2xl font-semibold text-text-primary">{project.name}</h1>
<h1 className="text-2xl font-semibold text-text-primary">{project.data.name}</h1>
<span
className={cn(
'rounded-full px-2 py-0.5 text-xs',
projectStatusColors[project.status] ?? 'bg-gray-600/20 text-gray-400',
projectStatusColors[project.data.status] ?? 'bg-gray-600/20 text-gray-400',
)}
>
{project.status}
{project.data.status}
</span>
</div>
{project.description ? (
<p className="mt-1 text-sm text-text-muted">{project.description}</p>
{project.data.description ? (
<p className="mt-1 text-sm text-text-muted">{project.data.description}</p>
) : null}
<p className="mt-2 text-xs text-text-muted">
Created {new Date(project.createdAt).toLocaleDateString()} · Updated{' '}
{new Date(project.updatedAt).toLocaleDateString()}
Created {new Date(project.data.createdAt).toLocaleDateString()} · Updated{' '}
{new Date(project.data.updatedAt).toLocaleDateString()}
</p>
</div>
</div>
</header>
{staleSnapshot !== null ? (
<div className="mb-6">
<StaleDataNotice label={staleSnapshot} onRetry={retryAll} />
</div>
) : null}
{missingSections.length > 0 ? (
<div className="mb-6">
<PartialDataNotice missing={missingSections} onRetry={retryAll} />
</div>
) : null}
<div className="mb-6 grid grid-cols-2 gap-3 sm:grid-cols-4">
<StatCard label="Tasks" value={String(tasks.length)} />
<StatCard
label="Tasks"
value={projectTasks === null ? UNKNOWN_VERDICT : String(projectTasks.length)}
/>
<StatCard
label="Done"
value={String(tasks.filter((task) => task.status === 'done').length)}
valueClass="text-success"
value={verdictValue(tasksVerified, String(doneCount))}
valueClass={tasksVerified ? 'text-success' : undefined}
/>
<StatCard
label="In Progress"
value={String(tasks.filter((task) => task.status === 'in-progress').length)}
valueClass="text-blue-400"
value={verdictValue(tasksVerified, String(inProgressCount))}
valueClass={tasksVerified ? 'text-blue-400' : undefined}
/>
<StatCard
label="Blocked"
value={String(tasks.filter((task) => task.status === 'blocked').length)}
valueClass={tasks.some((task) => task.status === 'blocked') ? 'text-error' : undefined}
value={verdictValue(tasksVerified, String(blockedCount))}
valueClass={tasksVerified && blockedCount > 0 ? 'text-error' : undefined}
/>
</div>
@@ -211,23 +294,43 @@ export function ProjectDetailPage(): ReactElement {
</div>
{activeTab === 'overview' ? (
<OverviewTab project={project} missions={missions} tasks={tasks} />
<OverviewTab project={project.data} missions={projectMissions} tasks={projectTasks} />
) : null}
{activeTab === 'tasks' ? (
<div>
<div className="mb-4">
<TaskStatusSummary
tasks={tasks}
activeFilter={taskFilter}
onFilterChange={setTaskFilter}
{projectTasks === null ? (
<UnavailableDataNotice
title="Tasks"
detail={describeFailure(tasks.failure)}
onRetry={retryAll}
/>
</div>
<TaskListView tasks={filteredTasks} onTaskClick={setSelectedTask} />
) : (
<>
<div className="mb-4">
<TaskStatusSummary
tasks={projectTasks}
activeFilter={taskFilter}
onFilterChange={setTaskFilter}
/>
</div>
<TaskListView tasks={filteredTasks} onTaskClick={setSelectedTask} />
</>
)}
</div>
) : null}
{activeTab === 'missions' ? <MissionTimeline missions={missions} /> : null}
{activeTab === 'missions' ? (
projectMissions === null ? (
<UnavailableDataNotice
title="Missions"
detail={describeFailure(missions.failure)}
onRetry={retryAll}
/>
) : (
<MissionTimeline missions={projectMissions} />
)
) : null}
{activeTab === 'prd' && prdContent ? (
<div className="rounded-lg border border-surface-border bg-surface-card p-6">
@@ -248,18 +351,26 @@ function OverviewTab({
tasks,
}: {
project: Project;
missions: Mission[];
tasks: Task[];
missions: Mission[] | null;
tasks: Task[] | null;
}): ReactElement {
const recentTasks = [...tasks]
.sort((left, right) => new Date(right.updatedAt).getTime() - new Date(left.updatedAt).getTime())
.slice(0, 5);
const recentTasks =
tasks === null
? null
: [...tasks]
.sort(
(left, right) =>
new Date(right.updatedAt).getTime() - new Date(left.updatedAt).getTime(),
)
.slice(0, 5);
return (
<div className="grid gap-6 lg:grid-cols-2">
<section>
<h2 className="mb-3 text-sm font-semibold text-text-secondary">Recent Tasks</h2>
{recentTasks.length === 0 ? (
{recentTasks === null ? (
<UnavailableDataNotice title="Tasks" />
) : recentTasks.length === 0 ? (
<div className="rounded-lg border border-surface-border bg-surface-card p-4 text-center">
<p className="text-sm text-text-muted">No tasks yet</p>
</div>
@@ -287,7 +398,9 @@ function OverviewTab({
<section>
<h2 className="mb-3 text-sm font-semibold text-text-secondary">Missions</h2>
{missions.length === 0 ? (
{missions === null ? (
<UnavailableDataNotice title="Missions" />
) : missions.length === 0 ? (
<div className="rounded-lg border border-surface-border bg-surface-card p-4 text-center">
<p className="text-sm text-text-muted">No missions yet</p>
</div>
+69 -3
View File
@@ -51,6 +51,7 @@ afterEach(async () => {
document.body.replaceChildren();
root = null;
apiMock.mockReset();
sessionStorage.clear();
});
async function renderProjectsPage(): Promise<ReturnType<typeof createMemoryRouter>> {
@@ -71,6 +72,22 @@ async function renderProjectsPage(): Promise<ReturnType<typeof createMemoryRoute
return router;
}
function clickButtonByText(text: string): void {
const button = [...container.querySelectorAll('button')].find((candidate) =>
candidate.textContent?.includes(text),
);
if (!button) {
throw new Error(`Button containing "${text}" not found`);
}
button.dispatchEvent(new MouseEvent('click', { bubbles: true }));
}
async function flushAct(): Promise<void> {
await act(async () => {
await Promise.resolve();
});
}
describe('ProjectsPage', () => {
it('shows a visible loading state while the project request is in flight', async () => {
const deferred = createDeferred<typeof projectFixtures>();
@@ -91,7 +108,7 @@ describe('ProjectsPage', () => {
const router = await renderProjectsPage();
expect(apiMock).toHaveBeenCalledWith('/api/projects');
expect(apiMock.mock.calls[0]?.[0]).toBe('/api/projects');
expect(container.textContent).toContain('Mosaic Stack');
expect(container.textContent).toContain('Agent Runtime');
@@ -108,7 +125,7 @@ describe('ProjectsPage', () => {
expect(container.textContent).toContain('Project detail target');
});
it('renders the empty state when the API returns no projects', async () => {
it('renders the empty state only for a verified empty collection', async () => {
apiMock.mockResolvedValueOnce([]);
await renderProjectsPage();
@@ -117,9 +134,12 @@ describe('ProjectsPage', () => {
expect(container.textContent).toContain(
'Projects will appear here when created via the gateway API',
);
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'current',
);
});
it('renders a visible alert when the projects request fails', async () => {
it('renders a failed fetch as an explicit unavailable state, never an empty collection', async () => {
apiMock.mockRejectedValueOnce(new Error('Projects are unavailable'));
await renderProjectsPage();
@@ -127,5 +147,51 @@ describe('ProjectsPage', () => {
const alert = container.querySelector('[role="alert"]');
expect(alert).toBeTruthy();
expect(alert?.textContent).toContain('Projects are unavailable');
expect(alert?.textContent).toContain('not an empty result');
// Negative controls: no healthy empty state and no project cards render
// from a failed fetch.
expect(container.textContent).not.toContain('No projects yet');
expect(container.textContent).not.toContain('Mosaic Stack');
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'unavailable',
);
});
it('renders an auth failure as unavailable and recovers after retry', async () => {
apiMock
.mockRejectedValueOnce(Object.assign(new Error('Unauthorized'), { statusCode: 401 }))
.mockResolvedValueOnce(projectFixtures);
await renderProjectsPage();
const alert = container.querySelector('[role="alert"]');
expect(alert?.textContent).toContain('Unauthorized');
expect(container.textContent).not.toContain('No projects yet');
await act(async () => {
clickButtonByText('Retry');
});
await flushAct();
expect(container.querySelector('[role="alert"]')).toBeNull();
expect(container.textContent).toContain('Mosaic Stack');
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'current',
);
});
it('renders a schema-mismatched response as unavailable, never as data', async () => {
apiMock.mockResolvedValueOnce({ results: projectFixtures });
await renderProjectsPage();
const alert = container.querySelector('[role="alert"]');
expect(alert?.textContent).toContain('not an empty result');
expect(container.textContent).not.toContain('Mosaic Stack');
expect(container.textContent).not.toContain('No projects yet');
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'unavailable',
);
});
});
+32 -34
View File
@@ -1,53 +1,51 @@
import { useEffect, useState, type ReactElement } from 'react';
import { type ReactElement } from 'react';
import { useNavigate } from 'react-router-dom';
import { ProjectCard } from '@/components/projects/project-card';
import { StaleDataNotice, UnavailableDataNotice } from '@/components/freshness/freshness-notices';
import { api } from '@/lib/api';
import type { Project } from '@/lib/types';
import { getErrorMessage } from './page-errors';
import { useFreshCollection, describeFailure } from '@/lib/freshness/use-fresh-collection';
import { validateProjectCollection } from '@/lib/freshness/validators';
export function ProjectsPage(): ReactElement {
const navigate = useNavigate();
const [projects, setProjects] = useState<Project[]>([]);
const [loading, setLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
useEffect(() => {
let cancelled = false;
void api<Project[]>('/api/projects')
.then((response) => {
if (cancelled) return;
setProjects(response);
})
.catch((caught: unknown) => {
if (cancelled) return;
setError(getErrorMessage(caught, 'Failed to load projects.'));
})
.finally(() => {
if (cancelled) return;
setLoading(false);
});
return () => {
cancelled = true;
};
}, []);
const projects = useFreshCollection<Project[]>({
source: 'gateway:/api/projects',
fetcher: (signal) => api<unknown>('/api/projects', { signal }),
validate: validateProjectCollection,
// Projects carry workspace identity (userId) that is only knowable from
// the payload itself, so a restored entry cannot be scope-checked before
// display. Conservative choice: no last-known restore for this surface;
// cross-workspace switching is still invalidated at verification time.
});
const retry = (): void => {
void projects.revalidate();
};
return (
<div className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<div
data-freshness={projects.freshness}
className="flex min-h-screen flex-col px-4 py-6 sm:px-6"
>
<header className="mb-6 border-b px-1 pb-3">
<h1 className="text-2xl font-semibold">Projects</h1>
</header>
{error ? (
<div role="alert" className="mb-6 rounded-lg border border-error/40 px-4 py-3 text-sm">
{error}
{projects.freshness === 'stale' && projects.snapshot ? (
<div className="mb-6">
<StaleDataNotice label={projects.snapshot} onRetry={retry} />
</div>
) : null}
{loading ? (
{projects.freshness === 'unknown' ? (
<p className="py-8 text-center text-sm text-text-muted">Loading projects...</p>
) : projects.length === 0 ? (
) : projects.freshness === 'unavailable' ? (
<UnavailableDataNotice
title="Projects"
detail={describeFailure(projects.failure)}
onRetry={retry}
/>
) : projects.data !== null && projects.data.length === 0 ? (
<div className="py-12 text-center">
<h2 className="text-lg font-medium text-text-secondary">No projects yet</h2>
<p className="mt-1 text-sm text-text-muted">
@@ -56,7 +54,7 @@ export function ProjectsPage(): ReactElement {
</div>
) : (
<div className="grid gap-4 sm:grid-cols-2 lg:grid-cols-3">
{projects.map((project) => (
{(projects.data ?? []).map((project) => (
<ProjectCard
key={project.id}
project={project}
+87 -1
View File
@@ -3,6 +3,9 @@ import { createRoot, type Root } from 'react-dom/client';
import { createMemoryRouter, RouterProvider, type RouteObject } from 'react-router-dom';
import { afterAll, afterEach, beforeAll, describe, expect, it, vi } from 'vitest';
import { taskFixtures } from './page-fixtures';
import { acceptSnapshot, DEFAULT_FRESHNESS_POLICY } from '@/lib/freshness/model';
import { writeSnapshotCache } from '@/lib/freshness/snapshot-cache';
import { validateTaskCollection } from '@/lib/freshness/validators';
const { apiMock } = vi.hoisted(() => ({
apiMock: vi.fn(),
@@ -48,6 +51,7 @@ afterEach(async () => {
document.body.replaceChildren();
root = null;
apiMock.mockReset();
sessionStorage.clear();
});
async function renderTasksPage(): Promise<void> {
@@ -72,6 +76,13 @@ function clickButtonByText(text: string): void {
button.dispatchEvent(new MouseEvent('click', { bubbles: true }));
}
/** Flush pending promise callbacks inside the act environment. */
async function flushAct(): Promise<void> {
await act(async () => {
await Promise.resolve();
});
}
describe('TasksPage', () => {
it('shows a visible loading state before the tasks request settles', async () => {
const deferred = createDeferred<typeof taskFixtures>();
@@ -132,7 +143,7 @@ describe('TasksPage', () => {
expect(container.textContent).toContain('Wire list and kanban modal interactions');
});
it('renders a visible alert when the tasks request fails', async () => {
it('renders a failed fetch as an explicit unavailable state, never an empty healthy board', async () => {
apiMock.mockRejectedValueOnce(new Error('Tasks request failed'));
await renderTasksPage();
@@ -140,5 +151,80 @@ describe('TasksPage', () => {
const alert = container.querySelector('[role="alert"]');
expect(alert).toBeTruthy();
expect(alert?.textContent).toContain('Tasks request failed');
expect(alert?.textContent).toContain('not an empty result');
// Negative controls: no board, no healthy empty-state markers, and the
// surface is marked unavailable rather than current.
expect(container.textContent).not.toContain('Not Started');
expect(container.textContent).not.toContain('No tasks');
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'unavailable',
);
});
it('recovers to a current board after retrying a failed fetch', async () => {
apiMock
.mockRejectedValueOnce(new Error('Tasks request failed'))
.mockResolvedValueOnce(taskFixtures);
await renderTasksPage();
expect(container.querySelector('[role="alert"]')).toBeTruthy();
await act(async () => {
clickButtonByText('Retry');
});
await flushAct();
expect(container.querySelector('[role="alert"]')).toBeNull();
expect(container.textContent).toContain('Not Started');
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'current',
);
});
it('labels restored last-known data as stale with source, version, and age until verified', async () => {
// Seed a last-known snapshot fetched five minutes ago; the page must
// render it only under an explicit staleness label while the fetch is
// still in flight.
const restored = acceptSnapshot({
value: taskFixtures,
validate: validateTaskCollection,
previous: null,
policy: DEFAULT_FRESHNESS_POLICY,
source: 'gateway:/api/tasks',
now: Date.now() - 5 * 60_000,
});
if (restored.outcome !== 'accepted') throw new Error('fixture setup failed');
writeSnapshotCache('tasks', restored.snapshot);
const deferred = createDeferred<typeof taskFixtures>();
apiMock.mockReturnValueOnce(deferred.promise);
await renderTasksPage();
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'stale',
);
const banner = container.querySelector('[role="status"]');
expect(banner?.textContent).toContain('last-known');
expect(banner?.textContent).toContain('may be out of date');
expect(banner?.textContent).toContain('gateway:/api/tasks');
expect(banner?.textContent).toContain('snapshot v1');
expect(banner?.textContent).toContain('5m ago');
// Last-known data still renders as situational awareness under the label.
expect(container.textContent).toContain('Route /tasks');
expect(container.textContent).not.toContain('Loading tasks...');
// Verification lands: the banner clears and the surface becomes current.
await act(async () => {
deferred.resolve(taskFixtures);
await deferred.promise;
});
expect(container.querySelector('[role="status"]')).toBeNull();
expect(container.querySelector('[data-freshness]')?.getAttribute('data-freshness')).toBe(
'current',
);
});
});
+26 -33
View File
@@ -1,45 +1,32 @@
import { useEffect, useState, type ReactElement } from 'react';
import { useState, type ReactElement } from 'react';
import { KanbanBoard } from '@/components/tasks/kanban-board';
import { TaskDetailModal } from '@/components/tasks/task-detail-modal';
import { TaskListView } from '@/components/tasks/task-list-view';
import { StaleDataNotice, UnavailableDataNotice } from '@/components/freshness/freshness-notices';
import { api } from '@/lib/api';
import { cn } from '@/lib/cn';
import type { Task } from '@/lib/types';
import { getErrorMessage } from './page-errors';
import { useFreshCollection, describeFailure } from '@/lib/freshness/use-fresh-collection';
import { validateTaskCollection } from '@/lib/freshness/validators';
type ViewMode = 'list' | 'kanban';
export function TasksPage(): ReactElement {
const [tasks, setTasks] = useState<Task[]>([]);
const tasks = useFreshCollection<Task[]>({
source: 'gateway:/api/tasks',
fetcher: (signal) => api<unknown>('/api/tasks', { signal }),
validate: validateTaskCollection,
cacheKey: 'tasks',
});
const [view, setView] = useState<ViewMode>('kanban');
const [loading, setLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
const [selectedTask, setSelectedTask] = useState<Task | null>(null);
useEffect(() => {
let cancelled = false;
void api<Task[]>('/api/tasks')
.then((response) => {
if (cancelled) return;
setTasks(response);
})
.catch((caught: unknown) => {
if (cancelled) return;
setError(getErrorMessage(caught, 'Failed to load tasks.'));
})
.finally(() => {
if (cancelled) return;
setLoading(false);
});
return () => {
cancelled = true;
};
}, []);
const retry = (): void => {
void tasks.revalidate();
};
return (
<div className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<div data-freshness={tasks.freshness} className="flex min-h-screen flex-col px-4 py-6 sm:px-6">
<header className="mb-6 flex items-center justify-between gap-4 border-b px-1 pb-3">
<h1 className="text-2xl font-semibold">Tasks</h1>
<div className="flex rounded-lg border border-surface-border">
@@ -70,18 +57,24 @@ export function TasksPage(): ReactElement {
</div>
</header>
{error ? (
<div role="alert" className="mb-6 rounded-lg border border-error/40 px-4 py-3 text-sm">
{error}
{tasks.freshness === 'stale' && tasks.snapshot ? (
<div className="mb-6">
<StaleDataNotice label={tasks.snapshot} onRetry={retry} />
</div>
) : null}
{loading ? (
{tasks.freshness === 'unknown' ? (
<p className="py-8 text-center text-sm text-text-muted">Loading tasks...</p>
) : tasks.freshness === 'unavailable' ? (
<UnavailableDataNotice
title="Tasks"
detail={describeFailure(tasks.failure)}
onRetry={retry}
/>
) : view === 'kanban' ? (
<KanbanBoard tasks={tasks} onTaskClick={setSelectedTask} />
<KanbanBoard tasks={tasks.data ?? []} onTaskClick={setSelectedTask} />
) : (
<TaskListView tasks={tasks} onTaskClick={setSelectedTask} />
<TaskListView tasks={tasks.data ?? []} onTaskClick={setSelectedTask} />
)}
{selectedTask ? (
+246
View File
@@ -1,5 +1,14 @@
# PRD: Mosaic Stack v0.1.0
## Current addendum: #1194 — Installed framework-tool drift detection
- Compare the framework tools shipped with the executing Mosaic package against the deployed `$MOSAIC_HOME/tools` tree by content hash.
- Treat every shipped `tools/**` file as framework-owned/required according to `framework-manifest.txt`, while excluding the explicit operator-owned credential carve-out and preserving installed-only operator/unknown files.
- Distinguish and count `IN_SYNC`, `STALE`, `NOT_INSTALLED`, and installed-only classifications; fail non-zero when shipped tools are stale or absent and refuse self-comparison that would make drift unobservable.
- Surface the observational check through `mosaic doctor`; do not refresh files, restart seats, or mutate live tooling.
- Document identity/messaging/gate behavior changes in the current stale set, the reviewed quiet-window keep-mode refresh command, and post-refresh probes against the installed path.
- Prove by construction that a stale and missing deployed tool are detected; that regression must fail before this checker exists.
## Metadata
- **Owner:** Jason Woltje
@@ -102,6 +111,128 @@ Context compaction, session replacement, and same-PID runtime reloads can leave
---
## Pi Persistent Goal Loop (#1150)
### Problem and objective
A Pi agent can stop after a plausible-looking answer even when the operator's broader objective is
not complete, and ordinary compaction can weaken or omit the original objective. Mosaic needs an
optional, operator-controlled goal loop that keeps a Pi session oriented, checks progress at native
lifecycle boundaries, and resumes work until completion is verified or a bounded safety state is
reached.
The objective is a Mosaic-owned Pi extension deployed from the framework into
`~/.config/mosaic/runtime/pi/`. It must not install into or depend on `~/.pi/agent/extensions/`.
### Scope
#### In scope
1. `PGL-REQ-01`: The framework SHALL ship a dedicated Pi goal extension under
`packages/mosaic/framework/runtime/pi/`, seed it under `$MOSAIC_HOME/runtime/pi/`, and make
`mosaic pi` load it alongside the core Mosaic extension when present.
2. `PGL-REQ-02`: `/goal` SHALL support setting a goal plus status, pause, resume, cancel, and help
operations without silently replacing an active goal.
3. `PGL-REQ-03`: Active branch-specific goal state SHALL be persisted in Pi custom session entries,
restored on session start and tree navigation, and never rely on a compaction summary as its
source of truth.
4. `PGL-REQ-04`: A hidden goal contract SHALL be injected through Pi's `context` event before every
model request so it remains effective across tool turns, retries, and post-compaction requests.
5. `PGL-REQ-05`: The harness SHALL inspect every `turn_end` and successful `session_compact` event.
A structured terminating goal-report tool SHALL capture `continue`, evidence-bearing `achieved`,
or `blocked` status without requiring a redundant model turn.
6. `PGL-REQ-06`: An achievement claim SHALL remain provisional until a second consecutive
evidence-bearing verification report. Any continuation report or successful compaction during
verification SHALL reset the verification sequence.
7. `PGL-REQ-07`: Continuation SHALL be initiated at safe lifecycle boundaries, primarily
`agent_settled`; manual compaction and restored active sessions may schedule a deferred idle
continuation without re-entering compaction handlers.
8. `PGL-REQ-08`: The loop SHALL have operator cancellation plus bounded turn and repeated-no-progress
limits. Exhausted or blocked goals pause rather than continuing indefinitely.
9. `PGL-REQ-09`: Framework installation and update SHALL preserve normal manifest ownership: the
goal extension is framework-owned under `runtime/**`, while no goal extension or configuration
asset is created or modified under the operator's main Pi configuration. Pi remains the owner of
its native session files used by `appendEntry()`.
#### Out of scope
1. A mathematical guarantee that an arbitrary natural-language goal is semantically complete.
2. Automatically executing user-supplied shell predicates or accepting executable validation code in
`/goal` arguments.
3. Restarting Pi after process, host, or supervisor failure; the existing Mosaic fleet/runtime
supervisor owns process durability.
4. Gateway, database, web UI, Discord, or cross-harness goal orchestration in this slice.
### User and stakeholder requirements
- An operator can start a goal from Pi and see its current phase, evidence, limits, and latest report.
- The agent remains oriented after each turn and compaction until verified, paused, blocked,
exhausted, or cancelled.
- Local testing uses a file under `~/.config/mosaic/runtime/pi/`; the feature never writes an
extension asset to `~/.pi/agent/extensions/`.
- Framework updates deploy the same reviewed extension source through Mosaic's existing manifest
sync path.
### Non-functional requirements
1. **Safety:** bounded continuation, explicit cancellation, no arbitrary command execution, and no
completion without non-empty reported evidence.
2. **Reliability:** serialized continuation scheduling, branch-aware restoration, compaction-safe
context injection, and stale-timer cancellation on session shutdown.
3. **Performance:** no extra nested judge-model request on every turn; structured reporting uses the
active agent's final terminating tool call.
4. **Observability:** Pi status/notifications expose phase and bounded counters without recording
credentials or hidden model reasoning.
5. **Maintainability:** the state machine is deterministic and behavior-tested independently from Pi
provider/network access.
### Acceptance criteria
1. `AC-PGL-01`: A framework-sync fixture installs the extension at
`$MOSAIC_HOME/runtime/pi/goal-extension.ts`, and launcher tests prove both Mosaic Pi extensions are
emitted in deterministic order while absent optional files remain backward-compatible.
2. `AC-PGL-02`: Command tests prove set/status/pause/resume/cancel behavior, active-goal replacement
refusal, and bounded input handling.
3. `AC-PGL-03`: Lifecycle tests prove every turn is recorded, active context is injected on every
request, two evidence-bearing achievement reports are required, and `agent_settled` continues an
unmet goal without duplicate scheduling.
4. `AC-PGL-04`: Compaction and restoration tests prove goal state survives, verification is reset and
rechecked after compaction, manual compaction continuation is deferred until idle, and tree/session
branch state is reconstructed correctly.
5. `AC-PGL-05`: Limit tests prove max-turn and repeated-no-progress exhaustion stop autonomous
continuation, while pause/cancel/blocked states do not restart.
6. `AC-PGL-06`: Focused tests, package typecheck/lint/test, repository quality gates, a local Pi load
smoke test from `~/.config/mosaic/runtime/pi/`, independent review, and terminal-green CI pass before
issue #1150 closes.
### Constraints, risks, and assumptions
- Dependency: Pi's extension API must continue to provide `registerCommand`, `registerTool`,
`context`, `turn_end`, `agent_settled`, `session_compact`, session custom entries, and terminating
tool results.
- Risk: the working agent can overstate completion. Mitigation: structured evidence, a mandatory
second verification pass, explicit semantic limitations, and operator-visible reports.
- Risk: an impossible goal can consume unbounded resources. Mitigation: hard turn/no-progress bounds
and paused terminal states.
- Risk: automatic continuation can race compaction or session replacement. Mitigation: drive from
`agent_settled`, defer idle restarts, generation-check timers, and clear timers on shutdown.
- `ASSUMPTION:` Two consecutive evidence-bearing reports are the initial local verification policy;
rationale: it provides a real recheck without doubling every turn's model cost. Future policy may
add independent or deterministic validators.
- `ASSUMPTION:` Default limits are 40 turns and 6 repeated no-progress reports, configurable only by
bounded Mosaic environment settings; rationale: useful persistence with a finite autonomous budget.
- `ASSUMPTION:` Documentation remains canonical in-repo for this slice; no external docs publication
is requested.
### Testing and delivery intent
Use TDD for the deterministic controller and lifecycle invariants. Test with fake Pi lifecycle
objects first, then run a local load/smoke test from the deployed Mosaic path. Deliver source, tests,
launcher wiring, framework/runtime documentation, user/developer guides, and sitemap updates in one
reviewed squash PR to `main` with terminal-green CI.
---
## Fleet Declarative Configuration Management Workstream (FCM, #758)
### Problem and objective
@@ -146,6 +277,68 @@ lands. M0 consists only of these normative requirements, the complete task DAG,
documentation IA checklist, and the legacy example/profile disposition inventory. Subsequent cards
are defined in [docs/TASKS.md](./TASKS.md) and must remain one card/one PR.
### Fleet git identity launch propagation (#1043)
#### Problem and objective
A fleet seat can have a registered per-agent Git credential while its launched runtime process lacks
`MOSAIC_GIT_IDENTITY`. The credential resolver then cannot select the seat identity reliably, which
blocks repository operations on fail-closed estates and can fall through to an unrelated identity on
estates where that refusal is not active. The objective is to make Git identity a deterministic,
roster-derived part of the generated launch projection and prove it reaches the launched process.
#### Normative requirements
1. `FGI-REQ-01`: Every generated fleet agent projection SHALL declare
`MOSAIC_GIT_IDENTITY=<MOSAIC_AGENT_NAME>`; a differing or unsafe identity SHALL fail closed before
tmux launch.
2. `FGI-REQ-02`: The clean `/usr/bin/env -i` pane boundary SHALL pass every variable declared by the
generated projection, including `MOSAIC_GIT_IDENTITY`, to the launched runtime process.
3. `FGI-REQ-03`: A behavioral integration test SHALL set-compare the complete generated projection
against the launched process environment. Source-text/string-presence assertions are insufficient.
4. `FGI-REQ-04`: Verification SHALL include RED-first evidence and a delete-the-subject mutation that
removes Git-identity pane propagation and makes the behavioral test fail.
#### Acceptance criteria
1. `AC-FGI-01`: A launched seat process contains every key/value pair declared by its generated
environment projection, including the roster-derived Git identity.
2. `AC-FGI-02`: Missing, unsafe, or split Git identity is rejected before a tmux session is created.
3. `AC-FGI-03`: Focused launcher and generated-environment tests, repository quality gates,
independent review, and the required RED/green/R7 evidence are recorded before push.
### Framework shell assertion portability (#1098)
#### Problem and objective
The blocking framework-shell chain can report that a pane command omitted `/usr/bin/env -i` even when
`-i` matched successfully. A short-circuiting `grep -q` under `set -o pipefail` may close its pipe after
the match and cause an upstream producer to exit with SIGPIPE, turning a valid semantic result into a
nonzero aggregate pipeline. The objective is to inspect the captured NUL-delimited argv directly and
make failures carry the observed records needed for diagnosis.
#### Normative requirements
1. `FSP-REQ-01`: The pane-boundary test SHALL validate an adjacent `/usr/bin/env`, `-i` argv pair from
the authoritative NUL-delimited tmux capture without a short-circuit pipeline whose upstream status
can override a successful match.
2. `FSP-REQ-02`: Missing, reversed, or non-adjacent boundary tokens SHALL fail, while valid boundaries
SHALL remain valid regardless of trailing argv size, pipe capacity, process scheduling, or host/CI
utility implementation.
3. `FSP-REQ-03`: A failed boundary check SHALL print stable indexed, shell-escaped observed argv records
before exiting nonzero; the fixture SHALL continue to contain generated non-secret launch data only.
4. `FSP-REQ-04`: Verification SHALL include RED-first large-payload evidence, negative token-order
controls, the complete focused launcher suite, canonical Woodpecker CI, and independent review.
#### Acceptance criteria
1. `AC-FSP-01`: A large captured argv with adjacent `/usr/bin/env`, `-i` passes even when the former
`grep -q` pipeline returns nonzero from an upstream SIGPIPE.
2. `AC-FSP-02`: Missing executable, missing flag, and detached/reversed flag fixtures return nonzero and
emit the indexed observed argv.
3. `AC-FSP-03`: The focused suite passes on the development host and CI image, and the merged-main
Woodpecker pipeline is terminal green before #1098 closes.
---
## Exact Cross-Harness Fleet Communications Contract (#766)
@@ -1345,6 +1538,59 @@ All work is **alpha** (< 0.1.0) until Jason approves 0.1.0 beta release.
---
## Workspace placement guard hardening (#1174)
### Problem and objective
The Bash pre-tool guard must prevent Git checkouts and repository state from being placed under
`$HOME` without refusing ordinary Git commands merely because a source, option value, branch name,
or metadata mentions `$HOME`. A guard that over-blocks routine work is unsafe because operators
will route around it.
### Scope and requirements
1. `WPG-REQ-01`: `git clone` and `git worktree add` placement SHALL be judged from their placement
operands, not from every HOME-shaped word in the command.
2. `WPG-REQ-02`: Clone sources, references, templates, environment assignments, and non-placement
worktree metadata MAY resolve under HOME when all placement operands resolve elsewhere.
3. `WPG-REQ-03`: Both attached and separate-value `--separate-git-dir` forms SHALL remain placement
operands and SHALL be refused when they resolve under HOME.
4. `WPG-REQ-04`: Option classification SHALL account for Git's rule-generated boolean negations
without relying on an enumerable allowlist of flag spellings.
5. `WPG-REQ-05`: Quote removal, escapes, shell command boundaries, redirections, and end-of-options
handling SHALL preserve existing fail-closed checkout coverage.
6. `WPG-REQ-06`: Absolute placement aliases SHALL resolve shell-known HOME spellings, dot segments,
repeated separators, and existing symlink parents before the HOME boundary comparison.
7. Relative targets whose effective path depends on the shell cwd are out of scope and tracked by
#1197.
### Acceptance and verification
1. Git's own option parser accepts each tested flag, including generated `--no-*` forms, while the
guard allows a HOME-valued source with an explicit safe destination.
2. Equivalent clone and worktree fixtures cover rule-generated negations and remain discriminating
against the prior head where the defect existed.
3. Real HOME destinations and both `--separate-git-dir` forms remain blocked, including placements
after shell command boundaries.
4. The full hermetic guard suite, syntax/static checks, adversarial probes, independent review, and
terminal-green CI pass before merge.
5. Any option-classification residual is documented with its deliberate failure direction.
### Constraints, risks, and assumptions
- Security and usability are co-equal: neither a placement bypass nor routine over-block is an
acceptable repair.
- `ASSUMPTION:` The value-taking option surface exposed by the installed Git version is closed and
measurable through Git's own parser/help output; rationale: boolean flags are rule-generated,
while separate-value options have explicit grammar and must be classified as such.
- Risk: a future Git release may add a new value-taking placement option. Mitigation: document the
chosen residual direction and pin every currently supported placement option in behavior tests.
- Risk: a symlink can be replaced after pre-execution canonicalization. Mitigation: resolve every
existing parent physically and document the remaining inherent TOCTOU window; the worktree helper
remains the authoritative path-derivation mechanism, with atomic closure tracked by #1199.
---
## Assumptions
1. RESOLVED: **pgvector is sufficient** for semantic search at v0.1.0 scale (personal/family/team = thousands to low hundreds-of-thousands of vectors). `@mosaicstack/memory` defines a `VectorStore` interface with pgvector as the default adapter. The interface boundary makes Qdrant a drop-in migration if PG resource contention or scale demands it later. Zero additional infrastructure for v0.1.0. Rationale: Reduces ops burden; pgvector HNSW indexes are fast at this scale; interface abstraction costs almost nothing now.
+38 -1
View File
@@ -7,7 +7,8 @@
3. [Provider Configuration](#provider-configuration)
4. [MCP Server Configuration](#mcp-server-configuration)
5. [Environment Variables Reference](#environment-variables-reference)
6. [Local Fleet Canary](./fleet-local-canary.md)
6. [Pi Goal Loop Operations](#pi-goal-loop-operations)
7. [Local Fleet Canary](./fleet-local-canary.md)
---
@@ -264,6 +265,16 @@ Each OIDC provider requires its client ID, client secret, and issuer URL togethe
| `AGENT_SYSTEM_PROMPT` | — | Platform-level system prompt injected into all sessions |
| `AGENT_USER_TOOLS` | all tools | Comma-separated allowlist of tools for non-admin users |
### Mosaic Pi goal loop
| Variable | Default | Description |
| ----------------------------- | ------- | -------------------------------------------------------------------- |
| `MOSAIC_GOAL_MAX_TURNS` | `40` | Per-goal autonomous turn limit; accepted range `1..500` |
| `MOSAIC_GOAL_MAX_NO_PROGRESS` | `6` | Consecutive identical progress-report limit; accepted range `1..100` |
These variables are consumed by the framework-owned Pi goal extension at goal creation. Invalid or
out-of-range values fall back to the defaults; they do not disable the bounds.
### Providers
| Variable | Default | Description |
@@ -374,3 +385,29 @@ Session cleanup is scoped to one session identifier and only removes that sessio
| Variable | Default | Description |
| ----------------------- | ----------------------------- | ------------------------------------------ |
| `MOSAIC_WORKSPACE_ROOT` | monorepo root (auto-detected) | Root path for mission workspace operations |
---
## Pi Goal Loop Operations
The reviewed runtime asset is deployed at
`~/.config/mosaic/runtime/pi/goal-extension.ts` by framework install/update. Do not install another
copy under `~/.pi/agent/extensions/`; duplicate registration can create suffixed commands and two
competing lifecycle controllers.
Operational checks:
1. Run `mosaic pi` and verify `/goal help` is available.
2. Use `/goal status` to inspect phase, turn/no-progress limits, compaction checks, and evidence.
Reports persist in Pi session data; controller-owned state redacts common credential shapes, but
Pi's model/tool-call history is separate. Operators must not place secrets or raw sensitive output
in goals, pause reasons, or evidence.
3. Use `/goal pause <reason>` before planned maintenance or manual investigation. Pause and cancel
abort the current goal-driven run when Pi is busy.
4. Use `/goal resume` only after addressing a blocker; counters restart with the configured bounds.
5. Use `/goal cancel` before replacing an unfinished goal.
A blocked or exhausted goal remains stopped and visible; Mosaic does not automatically raise its
limits or restart the process. Framework sync owns file deployment, while Pi's native session file
owns branch replay. Process/host restart remains the responsibility of the existing runtime or fleet
supervisor.
+55 -3
View File
@@ -8,9 +8,10 @@
4. [Tasks](#tasks)
5. [Settings](#settings)
6. [CLI Usage](#cli-usage)
7. [Sub-package Commands](#sub-package-commands)
8. [Telemetry](#telemetry)
9. [Local Fleet Canary](./fleet-local-canary.md)
7. [Pi Persistent Goals](#pi-persistent-goals)
8. [Sub-package Commands](#sub-package-commands)
9. [Telemetry](#telemetry)
10. [Local Fleet Canary](./fleet-local-canary.md)
---
@@ -317,6 +318,57 @@ mosaic prdy
mosaic quality-rails
```
## Pi Persistent Goals
`mosaic pi` loads a Mosaic-owned goal extension from
`~/.config/mosaic/runtime/pi/goal-extension.ts`. It is deliberately not installed in
`~/.pi/agent/extensions/`; framework installation and updates manage it with the rest of the Mosaic
runtime assets.
Start Pi, then set a goal:
```text
/goal set Deliver the feature, tests, documentation, and verification evidence
# Shorthand:
/goal Deliver the feature, tests, documentation, and verification evidence
```
Control and inspect the loop with:
| Command | Behavior |
| ---------------------- | ------------------------------------------------------------------ |
| `/goal status` | Show phase, limits, compaction checks, latest report, and evidence |
| `/goal pause [reason]` | Stop autonomous continuation while preserving the goal |
| `/goal resume` | Resume with fresh turn and no-progress counters |
| `/goal cancel` | Cancel the goal and remove its active status |
| `/goal help` | Show command help |
While a goal is active, Mosaic injects its contract before every Pi model request and checks every
completed model/tool turn. The agent ends each work cycle with the structured
`mosaic_goal_report` tool. `achieved` is provisional until a second consecutive report rechecks the
whole goal with evidence. A continuation report or a successful compaction resets provisional
verification.
Goal statements and reports are stored in Pi session data. Mosaic redacts common credential shapes
before appending its goal-state entries and before goal tool output or `/goal status`, but
pattern-based redaction is not a secret store. Pi's own model-message and tool-call records are
outside that redactor. Never put tokens, passwords, private keys, connection strings, or raw
sensitive output in a goal or report; cite the command, artifact, and pass/fail result instead.
The loop stops instead of running forever when it is paused, blocked, cancelled, verified, reaches
its turn limit, or repeats the same no-progress report too many times. Defaults are 40 turns and 6
repeated no-progress reports. Operators may lower or raise them within enforced bounds before
launching Pi:
```bash
MOSAIC_GOAL_MAX_TURNS=60 MOSAIC_GOAL_MAX_NO_PROGRESS=8 mosaic pi
```
Goal state is branch-specific Pi session data. It survives compaction and session resume, but Pi's
process still must be relaunched or supervised after a process/host failure. This initial verifier
checks structured evidence twice; it cannot mathematically prove every arbitrary natural-language
goal. Use explicit acceptance criteria and inspect `/goal status` for consequential work.
---
### Claude Code Skill Registration
+13 -10
View File
@@ -5,14 +5,14 @@ Generated environment files are rebuildable projections, not an operator-editabl
## Launch chain
| Layer | Responsibility |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| Roster | `fleet/roster.yaml` supplies the agent name, class, supported runtime, model, reasoning, tool policy, workdir, and tmux socket. |
| Projection writer | Renders deterministic fleet/agents/<name>.env.generated from the roster. |
| Optional local data | Reads a strict, data-only fleet/agents/<name>.env.local; it cannot shadow generated keys. |
| systemd | Starts the launcher with env -i and fixed bootstrap data. It does not preload either environment file. |
| session launcher | Validates generated and local data before it queries, creates, or stops an exact tmux session. |
| runtime launch | Derives the fixed mosaic yolo <runtime> argument array from validated roster data, then seeds the runtime contract. |
| Layer | Responsibility |
| ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Roster | `fleet/roster.yaml` supplies the agent name, class, supported runtime, model, reasoning, tool policy, workdir, and tmux socket; Git identity is derived from the exact agent name. |
| Projection writer | Renders deterministic fleet/agents/<name>.env.generated from the roster. |
| Optional local data | Reads a strict, data-only fleet/agents/<name>.env.local; it cannot shadow generated keys. |
| systemd | Starts the launcher with env -i and fixed bootstrap data. It does not preload either environment file. |
| session launcher | Validates generated and local data before it queries, creates, or stops an exact tmux session. |
| runtime launch | Derives the fixed mosaic yolo <runtime> argument array from validated roster data, then seeds the runtime contract. |
The launcher never `source`s or `eval`s an environment file and never accepts an environment-supplied
command. `MOSAIC_AGENT_COMMAND`, command/channel overrides, unknown keys, generated-key shadowing,
@@ -24,6 +24,7 @@ secret-like key names, duplicate keys, comments, quoted/export syntax, and unsaf
```dotenv
MOSAIC_AGENT_NAME=<roster name>
MOSAIC_GIT_IDENTITY=<roster name>
MOSAIC_AGENT_CLASS=<roster class>
MOSAIC_AGENT_RUNTIME=<roster runtime>
MOSAIC_AGENT_MODEL=<roster model hint>
@@ -33,8 +34,10 @@ MOSAIC_AGENT_WORKDIR=<absolute roster work directory>
MOSAIC_TMUX_SOCKET=<roster socket or empty>
```
The generated launch contract supports `claude`, `codex`, `opencode`, and `pi`. mosaic fleet add
rejects another runtime before it writes the roster or modifies generated, local, or quarantine state.
`MOSAIC_GIT_IDENTITY` is not independently configurable: it must equal `MOSAIC_AGENT_NAME`, preventing
split runtime and repository identity authority. The generated launch contract supports `claude`,
`codex`, `opencode`, and `pi`. mosaic fleet add rejects another runtime before it writes the roster or
modifies generated, local, or quarantine state.
The legacy dogfood stub remains an observability-only canary on its separate `mosaic-factory` socket;
it has no generated-launch adapter and cannot be added through this path.
@@ -3,11 +3,12 @@
The launcher consumes validated data, not shell configuration.
1. Read and validate the canonical roster.
2. Render deterministic <name>.env.generated data from that roster.
2. Render deterministic <name>.env.generated data from that roster, including `MOSAIC_GIT_IDENTITY` derived exactly from the roster agent name.
3. Parse optional <name>.env.local through a strict allowlist.
4. Reject generated-key shadowing, unknown or sensitive-looking keys, unsafe paths/values, duplicates, malformed lines, shell syntax, and command overrides.
5. Derive the runtime command from validated runtime/model/reasoning data.
6. Target only the exact configured tmux socket and roster session after ownership checks.
5. Reject a Git identity that is unsafe or differs from the generated agent name.
6. Derive the runtime command from validated runtime/model/reasoning data and pass every generated projection entry through the clean process environment boundary.
7. Target only the exact configured tmux socket and roster session after ownership checks.
## File precedence and ownership
@@ -35,6 +35,7 @@ values, credential material, or command text.
```dotenv
MOSAIC_AGENT_NAME=<roster name>
MOSAIC_GIT_IDENTITY=<roster name>
MOSAIC_AGENT_CLASS=<roster class>
MOSAIC_AGENT_RUNTIME=<roster runtime>
MOSAIC_AGENT_MODEL=<roster model hint>
@@ -44,8 +45,9 @@ MOSAIC_AGENT_WORKDIR=<absolute roster work directory>
MOSAIC_TMUX_SOCKET=<roster socket or empty>
```
The generated launch contract supports only `claude`, `codex`, `opencode`, and `pi`. fleet add
uses that same runtime authority and rejects any other runtime before it writes the roster or changes
`MOSAIC_GIT_IDENTITY` is derived from and must equal `MOSAIC_AGENT_NAME`; it is not a separate
operator-controlled identity authority. The generated launch contract supports only `claude`, `codex`,
`opencode`, and `pi`. fleet add uses that same runtime authority and rejects any other runtime before it writes the roster or changes
projection, local, or quarantine files. The legacy dogfood stub on its separate `mosaic-factory`
socket remains an observability canary; it has no generated-launch adapter and cannot be added through
this projection path.
+82 -2
View File
@@ -9,8 +9,9 @@
5. [Adding New MCP Tools](#adding-new-mcp-tools)
6. [Database Schema and Migrations](#database-schema-and-migrations)
7. [Claude Code Skill Bridge](#claude-code-skill-bridge)
8. [API Endpoint Reference](#api-endpoint-reference)
9. [Local Fleet Canary](./fleet-local-canary.md)
8. [Pi Persistent Goal Extension](#pi-persistent-goal-extension)
9. [API Endpoint Reference](#api-endpoint-reference)
10. [Local Fleet Canary](./fleet-local-canary.md)
---
@@ -396,6 +397,85 @@ M1 intentionally manages Claude Code only. Pi's Mosaic launcher can discover the
canonical root directly. Codex still relies on the existing full skill-sync
linker and needs separate parity analysis before this lifecycle API is extended.
## Pi Persistent Goal Extension
The source of the Mosaic-owned Pi goal controller is:
```text
packages/mosaic/framework/runtime/pi/goal-extension.ts
```
The framework manifest classifies `runtime/**` as framework-owned. Both the bash installer and the
TypeScript file adapter therefore deploy the same reviewed source to:
```text
$MOSAIC_HOME/runtime/pi/goal-extension.ts
# default: ~/.config/mosaic/runtime/pi/goal-extension.ts
```
Do not copy or link this extension into `~/.pi/agent/extensions/`. The launcher function
`discoverPiExtensionArgs()` emits the core `mosaic-extension.ts` first and the optional
`goal-extension.ts` second, preserving compatibility with an older installed framework that does
not have the goal file yet.
### Lifecycle design
| Pi API | Goal-controller responsibility |
| ------------------------------ | --------------------------------------------------------------------------------- |
| `registerCommand('goal')` | Set, inspect, pause, resume, or cancel one branch-specific goal |
| `registerTool(...)` | Record a terminating structured progress report with evidence |
| `context` | Inject the active goal contract before every provider request |
| `turn_end` | Record every turn, reject mixed final reports, and enforce the turn bound |
| `agent_settled` | Start one deduplicated continuation only after Pi has no retry/compact/queue work |
| `session_compact` | Record the compact check, reset provisional verification, and defer idle work |
| `session_start`/`session_tree` | Rebuild state from custom entries on the active branch |
| `session_shutdown` | Invalidate deferred callbacks and clear UI state |
State is appended as `mosaic-goal-state` custom entries, which do not enter model context. The
`context` hook creates a fresh hidden `mosaic-goal-context` message for each request instead of
trusting compaction summaries. The `mosaic_goal_report` result uses `terminate: true`; when it is the
sole final tool call, Pi avoids an unnecessary model response before the controller decides whether
to verify, continue, or stop.
Before state is appended or displayed, the controller applies bounded credential-pattern redaction
to the goal statement, report summary/evidence/next step, and stop reason. Fingerprints are computed
over redacted report content. Pi session entries are append-only, so a credential-bearing legacy
entry cannot honestly be erased by the extension: restoration fails closed, emits a warning, and
requires removal of the affected session before setting a new goal. This is defense-in-depth rather
than a secret-storage contract, and it does not rewrite Pi's separate model-message/tool-call
history. Goal prompts tell the agent not to submit credentials or raw sensitive output, and tests use
canaries to prove known forms do not reach new custom entries, status text, context, or tool details
while ordinary typed fields such as `token: string` remain intact.
Completion remains evidence-gated but semantic: two consecutive `achieved` reports are required,
and the second run is explicitly a verification pass. This avoids an extra judge-model request after
every turn. Deterministic validator commands are intentionally not accepted as `/goal` input in this
slice, so never describe this mechanism as proof of arbitrary natural-language completion.
### Tests and local smoke workflow
```bash
pnpm --filter @mosaicstack/mosaic exec vitest run \
src/runtime/pi-goal-extension.spec.ts \
src/commands/launch.spec.ts \
src/config/file-adapter.test.ts
bash packages/mosaic/framework/tools/quality/scripts/test-install-migration.sh
```
For an additive local smoke test without reseeding unrelated live framework files:
```bash
install -D -m 0644 \
packages/mosaic/framework/runtime/pi/goal-extension.ts \
~/.config/mosaic/runtime/pi/goal-extension.ts
pi --extension ~/.config/mosaic/runtime/pi/goal-extension.ts
```
Use `/goal help`, `/goal set ...`, and `/goal status` in that test session. A released framework
sync installs the file, and a released Mosaic CLI loads it automatically through `mosaic pi`.
## API Endpoint Reference
All endpoints are served by the gateway at `http://localhost:14242` by default.
+12 -12
View File
@@ -15,18 +15,18 @@
> `done` requires: repo quality gates green, independent review recorded, terminal-green CI on
> the PR head, squash merge to `next`, and acceptance evidence in notes.
| id | status | description | issue | agent | repo | branch | depends_on | estimate | notes |
| -------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----- | ---------- | ----------------- | --------------------------------- | ---------------------------------------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| RI-0-001 | in-progress | Bootstrap: issue #1275, PRD section, this DAG, scratchpad (docs only) | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-mission-bootstrap | — | 6K | |
| RI-1-001 | in-progress | RI-N1: canonical terminal verification command + publish-pipeline exact-commit gate (every publish step depends on verify; commit identity check; fail closed) | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-publish-gate | RI-0-001 | 25K | |
| RI-1-002 | not-started | RI-N1 negative control: checked-in tests proving a broken mandatory check blocks every publish step and that DAG edges cannot be bypassed | #1275 | pi-glm-5.3 | mosaicstack/stack | test/ri-050-publish-gate-negative | RI-1-001 | 12K | |
| RI-2-001 | in-progress | RI-N2 (Forge): remove stub-executor false success; `--simulate` typed `simulated` results that satisfy nothing; literal-`true` gates and echo-review replaced with real gates or typed waiting-for-authority | #1275 | pi-glm-5.3 | mosaicstack/stack | fix/ri-050-forge-fail-closed | RI-0-001 | 20K | Independent review APPROVED 2026-08-17 (Gitea review 172 on PR #1278, head 99b8f6ea; reviewing seat fargo — recorded under shared host principal mos-dt-0, provenance correction posted by fred; wrapper gap filed by fred). Executed at head: forge tests 116/116, lint green, typecheck green after building macp dist (minimal-install artifact, not a defect), workspace typecheck 45/45, no external type consumers of the changed interfaces. CI red = known lane-wide fleet-test failure only, carries no information about this change (fred, log-content analysis, pipelines 2456-2458). Non-blocking finding: README L141-143 + skills/mosaic-forge/SKILL.md document bare forge run/resume, which now fails closed — fast-follow docs touch. Merge queued behind #1270. |
| RI-2-002 | in-progress | RI-N2 (MACP): gate runner fails closed on empty commands, stub executors, and unimplemented CI-provider gates unless explicit simulate; typed capability failures | #1275 | pi-glm-5.3 | mosaicstack/stack | fix/ri-050-macp-fail-closed | RI-0-001 | 15K | |
| RI-3-001 | not-started | RI-N4: complete probe inventory mapping every TS and shell quality-rail check to one canonical check with disposition (preserve/strengthen/retire, each named) | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-qr-probe-inventory | RI-0-001 | 12K | |
| RI-3-002 | not-started | RI-N4: TS evaluator absorbs effective shell probes; typed results (passed/failed/blocked/error/not-applicable) with versioned digested check definitions; shell commands become thin adapters; contract/parity/negative-control tests | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-qr-evaluator | RI-3-001 | 30K | |
| RI-4-001 | in-progress | RI-N3: one PRD application service — `mission --plan` persists mission↔PRD linkage (ids/versions/selected requirements); `mosaic prdy` routes through the service or becomes a named import/export adapter; Markdown is a labeled generated view; explicit conflict-aware import | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-prd-authority | RI-0-001 | 35K | |
| RI-5-001 | not-started | RI-N5: typed freshness states (current/stale/partial/unknown/unavailable); no failed-fetch-renders-empty; stale derived verdicts → unknown; mutations disabled when stale; failure-matrix tests | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-web-stale-safety | RI-0-001 | 25K | |
| RI-V-001 | not-started | Final verification + release evidence: all cards verified merged, negative controls demonstrated, real `next` publish run green on exact commit, evidence pack recorded | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-release-evidence | RI-1-002, RI-2-001, RI-2-002, RI-3-002, RI-4-001, RI-5-001 | 10K | |
| id | status | description | issue | agent | repo | branch | depends_on | estimate | notes |
| -------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----- | ---------- | ----------------- | --------------------------------- | ---------------------------------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| RI-0-001 | done | Bootstrap: issue #1275, PRD section, this DAG, scratchpad (docs only) | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-mission-bootstrap | — | 6K | PR #1276 (head 758659dd): docs-only, CI green (2475). Review requested from fargo. Merges first (no publish run). |
| RI-1-001 | done | RI-N1: canonical terminal verification command + publish-pipeline exact-commit gate (every publish step depends on verify; commit identity check; fail closed) | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-publish-gate | RI-0-001 | 25K | PR #1277 (head 46784c8d): CI GREEN at head after serialized retry (pipeline 2476, 2026-08-18) - earlier red was CI-agent contention (web SPA timeouts under concurrent pipelines), not code. Review requested from fargo at pinned head (comms 20260818T021025Z). |
| RI-1-002 | done | RI-N1 negative control: checked-in tests proving a broken mandatory check blocks every publish step and that DAG edges cannot be bypassed | #1275 | pi-glm-5.3 | mosaicstack/stack | test/ri-050-publish-gate-negative | RI-1-001 | 12K | |
| RI-2-001 | done | RI-N2 (Forge): remove stub-executor false success; `--simulate` typed `simulated` results that satisfy nothing; literal-`true` gates and echo-review replaced with real gates or typed waiting-for-authority | #1275 | pi-glm-5.3 | mosaicstack/stack | fix/ri-050-forge-fail-closed | RI-0-001 | 20K | Independent review APPROVED 2026-08-17 (Gitea review 172 on PR #1278, head 99b8f6ea; reviewing seat fargo — recorded under shared host principal mos-dt-0, provenance correction posted by fred; wrapper gap filed by fred). Executed at head: forge tests 116/116, lint green, typecheck green after building macp dist (minimal-install artifact, not a defect), workspace typecheck 45/45, no external type consumers of the changed interfaces. CI red = known lane-wide fleet-test failure only, carries no information about this change (fred, log-content analysis, pipelines 2456-2458). Non-blocking finding: README L141-143 + skills/mosaic-forge/SKILL.md document bare forge run/resume, which now fails closed — fast-follow docs touch. Merge queued behind #1270. UPDATE 2026-08-18: #1270 merged; CI GREEN at head 4917df1f via serialized retry (pipeline 2477) - root cause of prior reds was CI-agent contention (web SPA timeouts under concurrent pipelines), superseding the fleet-test-failure theory. |
| RI-2-002 | done | RI-N2 (MACP): gate runner fails closed on empty commands, stub executors, and unimplemented CI-provider gates unless explicit simulate; typed capability failures | #1275 | pi-glm-5.3 | mosaicstack/stack | fix/ri-050-macp-fail-closed | RI-0-001 | 15K | PR #1293 (head 2097379e): CI green (pipeline 2465), independent review APPROVED (Gitea review 173, jarvis seat, 2026-08-17) - macp 109/109 verified at head. Merge queued behind #1276/#1277/#1278. |
| RI-3-001 | done | RI-N4: complete probe inventory mapping every TS and shell quality-rail check to one canonical check with disposition (preserve/strengthen/retire, each named) | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-qr-probe-inventory | RI-0-001 | 12K | PR #1302 (head e06a47fac591): CI green (2484), independent review APPROVED (Gitea review 187, fargo seat, 2026-08-18) — 54 rows / 21 canonical checks / dispositions 43-2-9-0 verified by row-count and code spot-checks. Merged by fargo at pinned head. |
| RI-3-002 | not-started | RI-N4: TS evaluator absorbs effective shell probes; typed results (passed/failed/blocked/error/not-applicable) with versioned digested check definitions; shell commands become thin adapters; contract/parity/negative-control tests | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-qr-evaluator | RI-3-001 | 30K | |
| RI-4-001 | in-progress | RI-N3: one PRD application service — `mission --plan` persists mission↔PRD linkage (ids/versions/selected requirements); `mosaic prdy` routes through the service or becomes a named import/export adapter; Markdown is a labeled generated view; explicit conflict-aware import | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-prd-authority | RI-0-001 | 35K | PR #1294 (head 8d258e1d): CI green (pipeline 2466), independent review APPROVED (Gitea review 174, jarvis seat, 2026-08-17) - prdy 20/20 + command specs 9/9 at head. Merge queued behind #1276/#1277/#1278. |
| RI-5-001 | done | RI-N5: typed freshness states (current/stale/partial/unknown/unavailable); no failed-fetch-renders-empty; stale derived verdicts → unknown; mutations disabled when stale; failure-matrix tests | #1275 | pi-glm-5.3 | mosaicstack/stack | feat/ri-050-web-stale-safety | RI-0-001 | 25K | |
| RI-V-001 | not-started | Final verification + release evidence: all cards verified merged, negative controls demonstrated, real `next` publish run green on exact commit, evidence pack recorded | #1275 | pi-glm-5.3 | mosaicstack/stack | docs/ri-050-release-evidence | RI-1-002, RI-2-001, RI-2-002, RI-3-002, RI-4-001, RI-5-001 | 10K | |
## Dispatch waves (max 2 parallel workers)
+186
View File
@@ -0,0 +1,186 @@
# Quality-Rails Probe Inventory — RI-3-001
- **Task:** RI-3-001 (SDLC-D-037 first half; PRD § Release Integrity Workstream, RI-N4)
- **Date:** 2026-08-18
- **Base:** `origin/next` @ `8199261c` (branch `docs/ri-050-qr-probe-inventory`)
- **Follow-up:** RI-3-002 consumes the dispositions here when building the single TS evaluator.
## 0. Scope and method
Every mechanism in this repository that verifies a quality, integrity, safety, or release
property — TypeScript checks, shell probes, pipeline steps, git hooks, and installer-side
assertions — gets one row. Each row's "what it actually verifies" was written from the
probe's **code**, not its name or docs. Framework tool unit/regression suites (git wrappers,
wake, tmux, orchestrator, …) are treated as one enforcement surface (`test:framework-shell`)
because they test tool behavior rather than repo quality; their wiring integrity is itself
guarded by `check-test-enumeration.sh`, and the quality-relevant members are rowed
individually.
**Kinds:** `ts` (TypeScript/Node check), `shell` (bash/python probe), `pipeline-step`
(exists only inside a Woodpecker pipeline).
**Enforcement points:** `local` (operator-invoked), `pre-commit`, `pre-push`,
`CI ci.yml#<step>`, `publish.yml#<step>` (CI on push to main/next), `turbo <task>`,
`agent-runtime` (framework hooks on an agent host), `installer` (host install path),
`unwired`.
**Dispositions** (recommendations for RI-3-002): `preserve` (keep as-is; already the
canonical or a correct guard-of-the-guard), `strengthen` (keep, but a concrete gap must
close — usually absorption into the TS evaluator), `strengthen (review)` (viable retirement
candidate once the evaluator absorbs it; do not retire yet). Note: RI-N4 requires that
effective shell probes be **absorbed before** their independent paths retire — no row here
is marked `retire` because no absorption exists yet.
## 1. Inventory
### 1.1 Repo-level gate tasks (pnpm / turbo)
| check | location | kind | what it actually verifies | enforcement point | canonical check | disposition | rationale |
| ------------------------------------- | ------------------------------------------------------------------------------------ | ---- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | ------------------------- | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `pnpm preflight` (checkout preflight) | `scripts/preflight.mjs` | ts | Six gate binaries (eslint, husky, prettier, tsc, turbo, vitest) exist and are executable in `node_modules/.bin` (exit 42 if not); no stale `.mosaic-test-work/web-build.lock` (exit 43); `apps/web/.next` is a real directory (not a symlink), every entry owned by the current uid, and its `.mosaic-source-hash` fingerprint + `.mosaic-symlink-manifest` hash match the certified build written by `scripts/build-web.mjs` | `pre-push`; inside `pnpm typecheck` (→ `CI ci.yml#typecheck`, verify-release `typecheck` stage) | QC-1 Checkout integrity | preserve | Blocks a poisoned/stale generated `.next` from faking a green typecheck (the five-month-stale-`.next` class); trust chain is self-contained per-checkout. |
| `pnpm typecheck` | root `package.json``turbo run typecheck` | ts | Per-package `tsc --noEmit` (all 20 packages); turbo `typecheck` depends on `^build`, so package builds must succeed first; prefixed by checkout preflight | `CI ci.yml#typecheck`; `pre-push`; verify-release `typecheck` stage; `turbo typecheck` | QC-2 Workspace typecheck | preserve | The single workspace-wide type gate; CI and hooks invoke the same task, no divergent checklist. |
| `pnpm lint` | root `package.json``turbo run lint` | ts | Per-package `eslint src` under root `eslint.config.mjs` (ignores `dist`, `.next`, `framework/**`, etc.) | `CI ci.yml#lint`; `pre-push`; verify-release `lint` stage; `turbo lint` | QC-3 Workspace lint | preserve | Same-task invocation from every surface; no second lint definition. |
| `pnpm format:check` | root `package.json``prettier --check` | ts | Prettier parse/format equality over `**/*.{ts,tsx,js,jsx,json,md}` minus `.prettierignore` (generated trees, `docs/scratchpads/`, venvs, …) | `CI ci.yml#format`; `pre-push`; verify-release `format` stage | QC-4 Format check | preserve | Single formatter, single ignore list, enforced identically everywhere. |
| `pnpm test` | root `package.json` `test` = `test:checkout` && `turbo run test` && `test:installer` | ts | (a) `node --test scripts/*.test.mjs` — checkout-tool units; (b) per-package `vitest run` (mosaic appends the 47-command `test:framework-shell` chain); (c) `tools/install-next-lane.test.sh`; turbo `test` declares DB env vars and depends on `^build` | `CI ci.yml#test` (with `DATABASE_URL` + `db:migrate` first); verify-release `test` stage; `turbo test` | QC-5 Test suite execution | preserve | One composed test command; the chain property (any link red ⇒ step red) is the gate. |
| `pnpm build` | root `package.json``turbo run build` | ts | Per-package build (`tsc`/Next) with `^build` dependency and `dist/**` outputs | `publish.yml#build`; verify-release `build` stage; `turbo build` | QC-6 Workspace build | preserve | Publish artifacts derive from the same build task CI verifies. |
### 1.2 Framework quality shell probes (`packages/mosaic/framework/tools/quality/`)
| check | location | kind | what it actually verifies | enforcement point | canonical check | disposition | rationale |
| ------------------------------------------- | ----------------------------------------------------------------------- | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- |
| Sanitization gate | `scripts/verify-sanitized.sh` | shell | Built-in self-test first (planted identity/structural/YAML+service fixtures; exit 2 if the regexes or extension coverage break), then: (1) identity denylist grep (`jarvis\|jason\|woltje\|brain.woltje.com\|/home/jwoltje\|\bPDA\b`) over all shipped text files **including** `examples/`; (2) structural grep for private `$HOME/src` defaults in shipped scripts **excluding** `examples/`. Any hit ⇒ exit 1 | `CI ci.yml#sanitization`; verify-release `sanitization` stage | QC-7 Framework sanitization | preserve | Labeled one-time regression guard with a self-test that prevents silent no-op; correctly scoped (identity vs structural) and documented as not a general PII detector. |
| Resident-context budget | `scripts/check-resident-budget.sh` (+ `--self-test`) | shell | Self-test of the comparator, then `wc -l` vs per-file ceilings (CONSTITUTION 120, AGENTS 120, each RUNTIME.md 90); missing file ⇒ fail; over ceiling ⇒ exit 1 | `CI ci.yml#sanitization` (both modes); verify-release `sanitization` stage | QC-8 Resident-context budget | preserve | Caps the container (lines), never the wording — the deliberate anti-drift design (DESIGN §7); CI-enforceable half only, by design. |
| Test-membership enumeration guard (#1017) | `scripts/check-test-enumeration.sh` + `test-enumeration-exclusions.txt` | shell | Parses surface S1 (`packages/mosaic` `test:framework-shell` via JSON+shlex) and S2 (every `framework/tools/\*.sh | .py`token in`ci.yml`, comment lines stripped); population = `_test_.sh`under`framework/tools`; FAILS on: suite-shaped file on disk neither enumerated nor signed-excluded; surface naming a path missing on disk (both directions); exclusion without reason / stale / outside population / contradicting enumeration. Proves **naming, not reachability** (stated in-file) | `CI ci.yml#sanitization` (direct line); link [0] of `test:framework-shell` (thus `CI ci.yml#test`); verify-release `sanitization` stage | QC-9 Test-membership enumeration | preserve | Makes silent under-run impossible; invoked from both surfaces it audits so severing the chain cannot silence it. |
| Enumeration-guard needles | `scripts/test-check-test-enumeration.sh` | shell | Needle/control fixtures driven through `--root`: every promised failure mode must trip the guard **on its own words**, plus controls that must pass (null-case defense); covers commented-out ci.yml lines (F1) and line-range parsing (n2b) | `test:framework-shell``CI ci.yml#test`; verify-release `test` stage | QC-9 Test-membership enumeration | preserve | Guard-of-the-guard with both polarities; same canonical check by design. |
| Upgrade manifest guard (#791 HARD GATE) | `scripts/test-upgrade-manifest-guard.sh` | shell | Keep-mode `install.sh` upgrade against seeded throwaway `MOSAIC_HOME`: every operator sentinel — including an **unanticipated** one — survives byte-identical with unchanged mtime; framework files still update; retired framework files pruned; matrix run with rsync present AND absent (keep path must be rsync-independent); fail-closed matrix (empty/operator-only/malformed/missing manifest aborts loudly, operator files untouched); operator secret never appears in installer output | `CI ci.yml#upgrade-guard`; verify-release `upgrade-guard` stage | QC-10 Upgrade/install safety | preserve | The operator-data hard gate for the `mosaic update` path; negative controls are load-bearing and documented. |
| Upgrade rollback gate (#791 B1) | `scripts/test-upgrade-rollback.sh` | shell | Mid-sync failure (PATH-shadowing `cp` shim) must trigger snapshot restore: restore message fires, corrupted file restored, target byte-identical to pre-upgrade; control installer with `set -E` stripped must NOT roll back (proves errtrace is load-bearing); plus signal/exit-guard controls | `CI ci.yml#upgrade-guard`; verify-release `upgrade-guard` stage | QC-10 Upgrade/install safety | preserve | Proves the rollback trap actually fires; the `-E`-stripped control keeps Part A honest. |
| Durable-snapshot gate (#791 PR2) | `scripts/test-upgrade-durable-snapshot.sh` | shell | Pre-update snapshot taken before any mutation (0700/0600 perms, secret never logged, retention-pruned); post-sync verify net restores operator files a manifest bug lets the sync touch; CWE-59 symlink-leaf guard proven with a portable cp shim in both polarities (write-through-link must not happen); v1→v2 migration semantics (intended `bin/` removal not healed) | `CI ci.yml#upgrade-guard`; verify-release `upgrade-guard` stage | QC-10 Upgrade/install safety | preserve | Covers tampering and leak vectors the manifest guard cannot see; the shim rationale (busybox vs GNU cp) is documented in-file. |
| Install migration matrix (v2→v3) | `scripts/test-install-migration.sh` | shell | Fixture matrix running the real installer with `MOSAIC_SYNC_ONLY=1`: fresh install seeds + stamps version 3; legacy user-edited AGENTS overwritten with `.pre-constitution.bak` preserved (and idempotent); tuned STANDARDS overwritten; operator files (SOUL, credentials) preserved. Mirrors the TS suite `packages/mosaic/src/config/file-adapter.test.ts` — both installers must behave identically | `CI ci.yml#upgrade-guard`; verify-release `upgrade-guard` stage | QC-10 Upgrade/install safety | preserve | Pins the shell/TS installer parity contract; removal would orphan that parity requirement. |
| Enforcement verification probe (bash) | `scripts/verify.sh` | shell | Attempts **real commits** in the target repo: planted type error must produce a commit blocked with `error`; planted `any` must trip `no-explicit-any`; planted lint error must trip `prettier`; gitleaks binary must exist (3a) and detect a planted AWS key via `gitleaks git --pre-commit --staged --redact` (3b). Verdicts are output-grep matches on hook stderr | `local` via installed `mosaic-quality-verify` on scaffolded target projects; **not run in this repo's CI** | QC-20 Downstream enforcement verification | strengthen (review) | Mechanism is genuinely behavioral (stronger than file presence) but verdict logic is grep-on-output and it is unwired here; absorb as the evaluator's enforcement-probe check (the RI-N4 evaluator invokes it or reimplements it) before retiring the shell path. |
| Enforcement verification probe (PowerShell) | `scripts/verify.ps1` | shell | Windows port of `verify.sh`: same planted-commit tests with `$output -match` matching; no gitleaks self-test parity beyond the same checks | `local` (Windows operator); no Windows CI runner exists | QC-20 Downstream enforcement verification | strengthen (review) | A hand-maintained twin of `verify.sh` with no CI coverage — exactly the drift shape the single evaluator removes; retire after the TS evaluator owns the probe. |
| Quality template installer (bash) | `scripts/install.sh` | shell | Copies template files (`.husky/pre-commit` incl. mandatory gitleaks, `.lintstagedrc.js`, `.eslintrc.js`, `tsconfig.json`, `.woodpecker.yml`, `.gitleaks.toml`) into a target project; **warns** (does not verify) about `package.json` snippet merge; no post-condition check | `local` / via `mosaic-quality-apply` | QC-21 Downstream rails scaffolding | strengthen (review) | Duplicates the TS `quality-rails init` scaffolder for a different template set; converging on one scaffolder (with post-scaffold verification) is prerequisite to retiring this path. |
| Quality template installer (PowerShell) | `scripts/install.ps1` | shell | Windows twin of the template copy above | `local` (Windows operator) | QC-21 Downstream rails scaffolding | strengthen (review) | Same twin-drift risk as `verify.ps1`; no runner exercises it. |
| `mosaic-quality-verify` adapter | `framework/tools/_scripts/mosaic-quality-verify` | shell | Thin adapter: validates target dir exists, asserts `verify.sh` present+executable, `cd` target, exec it. No verdict logic of its own | `local` (installed framework bin) | QC-20 Downstream enforcement verification | preserve | Already the thin-adapter shape RI-N4 prescribes for shell surfaces. |
| `mosaic-quality-apply` adapter | `framework/tools/_scripts/mosaic-quality-apply` | shell | Thin adapter: arg validation then exec of quality `install.sh --template … --target …` | `local` (installed framework bin) | QC-21 Downstream rails scaffolding | preserve | Thin adapter, no separate verdict; disposition follows its target script's convergence. |
| Roster schema regression | `scripts/test-roster-schema.py` | shell | jsonschema `Draft202012Validator` over `fleet/roster.schema.json` with valid/invalid connector-kind fixtures (tmux/discord/matrix conditional fields) | **unwired** — not on S1 or S2, not signed-excluded; also outside the enumeration guard's `*.sh` population, so the guard cannot see it | QC-5 Test suite execution | strengthen (review) | A real regression suite that currently runs nowhere; wire it into a CI surface or sign an exclusion — leaving it invisible re-arms the exact gap #1017 closed. |
| Framework shell chain (S1) | `packages/mosaic/package.json` `test:framework-shell` | shell | 47-command `&&` chain: enumeration guard + needles, 14 lease-broker/mutator-gate python unitests, `check-runtime-launches.py`, and ~30 framework-tool shell suites (git wrappers, wake, woodpecker, tmux, glpi, orchestrator, `_scripts`). Quality-relevant members rowed separately below | `turbo test``CI ci.yml#test`; verify-release `test` stage | QC-5 Test suite execution | preserve | The chain is the execution surface the enumeration guard audits; known residuals: a failing link stops later suites (measured in #1270 — suites after position 44 had not run), and the guard proves naming, not reachability. |
### 1.3 Framework runtime hooks and their harnesses (agent-host enforcement)
| check | location | kind | what it actually verifies | enforcement point | canonical check | disposition | rationale |
| ------------------------------------- | ----------------------------------------------------------------------------------------- | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------- | --------------------------------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| QA edit hook seam | `framework/tools/qa/qa-hook-stdin.sh` (+ `qa-hook-handler.sh`) | shell | PostToolUse stdin hook: extracts edited file from the tool JSON (jq or grep fallback), skips non-JS/TS, then the deps-preflight gate — exits 1 with the legible sentinel `deps not installed — run pnpm install` when `node_modules/.bin` is missing/empty (the #856 false-red class); the downstream handler only files QA remediation **report templates** (no verification logic) | `agent-runtime` (framework `runtime/claude/settings.json` PostToolUse); never CI | QC-16 Agent-runtime edit-time checks | strengthen (review) | The sentinel gate is real enforcement; the handler's report-filing adds no verdict and its name promises more than the code does — evaluator absorption should keep the sentinel, drop the report theater. |
| Typecheck-on-edit hook | `framework/tools/qa/typecheck-hook.sh` | shell | PostToolUse: for edited `.ts/.tsx`, finds nearest `tsconfig.json` and runs `tsc --noEmit`, surfacing errors nonzero to the agent immediately | `agent-runtime` (framework `runtime/claude/settings.json` PostToolUse) | QC-16 Agent-runtime edit-time checks | strengthen (review) | Edit-time duplicate of QC-2 with independent invocation logic; keep behavior, converge invocation through the evaluator adapter. |
| Deps-preflight harness | `framework/tools/qa/test-deps-preflight.sh` | shell | Five assertions against the seam incl. a documented RED control (raw `not found`), sentinel behavior for missing and empty `.bin`, and no-false-positive once populated | `test:framework-shell``CI ci.yml#test` | QC-16 Agent-runtime edit-time checks | preserve | Guard-of-the-check with a red control; keeps the sentinel from regressing. |
| Prompt-helper RCE regression | `framework/tools/_scripts/test-mosaic-init-rce.sh` | shell | Sources the prompt helpers and proves a literal `$(touch /tmp/pwned)` answer round-trips verbatim and never executes (no `/tmp/pwned` created) | `test:framework-shell``CI ci.yml#test` | QC-5 Test suite execution | preserve | Cheap, load-bearing security regression on the installer's input path. |
| Install-ordering harness (#869 C2) | `framework/tools/_scripts/test-install-ordering-guard.sh` | shell | Drives `mosaic-link-runtime-assets` with a fake `mosaic` on PATH: probe ok ⇒ settings copied + exit 0; probe fail ⇒ exit 1 with degraded outcome but all other runtime files still copied; `--allow-inactive-enforcement` forwarded; no-mosaic-on-PATH ⇒ python3 fallback strips enforcement hooks and exits 1; fallback + flag ⇒ wires as-is, exit 0 | `test:framework-shell``CI ci.yml#test` | QC-17 Lease-enforcement wiring safety | preserve | Exercises the shell wiring seam independently of the TS guard's own spec suite (complementary coverage, by design). |
| Fleet-transport harness (#1240) | `framework/tools/_scripts/test-fleet-transport-check.sh` | shell | Extracts the shipped `check_fleet_transport`/`fleet_declared_transport` functions **from the shipped scripts** (fails loud if extraction yields nothing) and drives both implementations (mosaic-doctor + `tools/install.sh`) from one case table | `test:framework-shell``CI ci.yml#test` | QC-18 Operator-host drift audit | preserve | The anti-drift harness for the one rule shipped twice; extraction-from-source keeps it from testing a stale copy. |
| Terminal-green contract (RM-61/#1000) | `framework/tools/woodpecker/test-terminal-green-contract.sh` + `verify-terminal-green.py` | shell | Red-first fixtures: pipeline JSON variants (service failure, step failure, cancelled, etc.) must produce the correct terminal-green verdict; controls must pass | `test:framework-shell``CI ci.yml#test` | QC-5 Test suite execution | preserve | Keeps the CI-wait wrapper's green-detection honest; a false green here would poison every merge gate that trusts `pr-ci-wait.sh`. |
| Lease-gate launch invariant | `framework/tools/lease-broker/check-runtime-launches.py` | shell | Scans production roots (`packages/`, `apps/`, `plugins/`, `tools/`) across sh/py/ts/yaml suffixes for Claude/Pi process launches **outside** the lease gate; allowlist-based; fails CI on violation | `test:framework-shell``CI ci.yml#test` | QC-15 Lease-gate architecture invariant | preserve | The only architectural "no ungated launches" rail; grep+allowlist is the right cost/benefit for this invariant. |
### 1.4 TypeScript quality logic (`@mosaicstack/quality-rails` + mosaic CLI)
| check | location | kind | what it actually verifies | enforcement point | canonical check | disposition | rationale |
| ---------------------------------------- | ---------------------------------------------------------------------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------- | ------------------------------------- | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `quality-rails check` | `packages/quality-rails/src/cli.ts` (`mosaic quality-rails check --project`) | ts | **Expected-file presence only**: loops `expectedFilesForKind` (node: `.eslintrc`, `biome.json`, `.githooks/pre-commit`, `PR-CHECKLIST.md`; python: `pyproject.toml`+hooks+checklist; rust: `rustfmt.toml`+…) and exits 1 listing missing paths. Does not execute any linter, formatter, hook, or scanner | `local` (operator CLI); **no CI wiring in this repo** | QC-19 Downstream rails presence check | strengthen | This is the RI-N4 evaluator seed. Today presence ≠ parity (explicitly called out by RI-N4): it must grow typed verdicts (`passed/failed/blocked/error/not-applicable`), check versioning/subject/reason, digested definitions, and absorb the effective shell probes (QC-20 first). |
| `quality-rails doctor` | `packages/quality-rails/src/cli.ts` | ts | Same presence data as `check`, printed with ok/missing lines; **cannot fail** (no nonzero exit on missing files) | `local` (operator CLI) | QC-19 Downstream rails presence check | strengthen | A doctor that cannot fail is advisory; fold into `check` (or return typed states) when the evaluator lands. |
| `quality-rails init` | `packages/quality-rails/src/cli.ts` + `scaffolder.ts`/`templates.ts` | ts | Scaffolds rails files per detected kind/profile (linters/formatters lists are advisory strings; hooks flag always true); writes files, prints follow-ups — no post-condition verification | `local` (operator CLI) | QC-21 Downstream rails scaffolding | strengthen (review) | Second scaffolding path alongside quality `install.sh` (§1.2); converge on one with post-scaffold verification before retiring either. |
| Lease activation probe (#869 C1, hidden) | `packages/mosaic/src/commands/lease-activation-probe.ts` | ts | Real capability probe, not file presence: resolves the installed mosaic CLI and requires it to advertise the exact `{name, version}` activation contract; all deps injectable; registered as hidden CLI command and consumed by C2/C5 | `local` (hidden CLI + consumed by C2/C5); spec-tested via `lease-activation-probe.spec.ts` in `turbo test` | QC-17 Lease-enforcement wiring safety | preserve | The versioned-contract probe is precisely the fail-closed capability check RI-N2 generalizes; already typed and injectable. |
| Install-ordering guard (#869 C2, hidden) | `packages/mosaic/src/commands/install-ordering-guard.ts` | ts | Decides whether enforcement hook entries are written into the `~/.claude/settings.json` the framework reseed ships: not activatable ⇒ strip hooks + nonzero loud outcome (default); explicit per-invocation `--allow-inactive-enforcement` opt-out wires-with-warning. Never touches the runtime gate's own fail-closed behavior | `installer` (framework reseed via `mosaic-link-runtime-assets`); spec + shell harness coverage in `turbo test` | QC-17 Lease-enforcement wiring safety | preserve | Correct default-deny with an explicit, non-env opt-out; test-locked from both the TS and shell sides. |
| Lease doctor check (#869 C5) | `packages/mosaic/src/commands/lease-doctor-check.ts` | ts | Combines hook-wiring detection in `~/.claude/settings.json` with C1 activatable and C3 broker-supervisor health: wired ∧ (¬activatable ¬healthy) ⇒ loud `[ERROR]` that forces `mosaic doctor` exit 1 regardless of the bash audit's own exit | `local` (inside `mosaic doctor`); spec coverage in `turbo test` | QC-17 Lease-enforcement wiring safety | preserve | Closes the "bricked host looks green" hole; cannot be masked by the bash script — that composition is the point. |
| `mosaic doctor` (framework drift audit) | `packages/mosaic/src/commands/launch.ts` (`doctor`) + `framework/tools/_scripts/mosaic-doctor` | shell+ts | Bash audit of the installed framework home: ~40 expected files/dirs present; runtime files are copies (not symlinks) matching source (`cmp`) or composed runtime-contract markers; hard-gates block present in AGENTS.md; sequential-thinking MCP configured; fleet transport binary present per roster (warn); legacy symlink trees gone; skills synced — **warn-based, exit 1 only with `--fail-on-warn`**, plus C5's forced error | `local` (operator audit) | QC-18 Operator-host drift audit | preserve | Host-state audit CI cannot see (user files by design, DESIGN §7); advisory exit is the documented contract — do not silently change it. |
| `mosaic gateway doctor` | `packages/mosaic/src/commands/gateway-doctor.ts` | ts | Probes per-service health (PostgreSQL, Valkey, pgvector) via `@mosaicstack/storage`, reports tier and JSON; exit 1 only when at least one **required** service fails (yellow stays 0) | `local` (operator) | QC-18 Operator-host drift audit | preserve | Service health with correct red/yellow exit semantics; JSON mode exists for scripting. |
| `mosaic gateway verify` | `packages/mosaic/src/commands/gateway/verify.ts` | ts | Post-install liveness: daemon meta via HTTP with retries, admin token on file, bootstrap endpoint reachable; aggregated pass/fail | `local`; consumed by `tools/e2e-install-test.sh` | QC-18 Operator-host drift audit | preserve | The first-run proof the installer E2E relies on; retry-aware so startup races don't false-red. |
| `mosaic fleet doctor` | `packages/mosaic/src/commands/fleet-reconciler-command.ts` | ts | Classifies local roster-owned drift (no mutation) from the parsed v2 roster | `local` (operator) | QC-18 Operator-host drift audit | preserve | Dry-run classification is the correct non-mutating audit shape. |
### 1.5 Git hooks (developer machine)
| check | location | kind | what it actually verifies | enforcement point | canonical check | disposition | rationale |
| ------------------------- | --------------------------------------------------------- | ----- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | --------------------------- | ----------- | ----------------------------------------------------------------------------------------------------------------- |
| Pre-commit staged hygiene | `.husky/pre-commit``npx lint-staged` (`.lintstagedrc`) | shell | On staged files only: `prettier --write` + `eslint --fix` for ts/tsx/js/jsx; `prettier --write` for json/md/yaml/yml. **Mutating** (fixes and re-stages); commit blocks only if a fixer itself fails | `pre-commit` (every local commit; hooks activated by `install-hooks.mjs` via `core.hooksPath .husky/_`) | QC-13 Staged-change hygiene | preserve | Correct scoped fast gate; note it auto-fixes rather than rejects (deliberate). Gap: no secret scan here — see §3. |
| Pre-push gate | `.husky/pre-push` | shell | `pnpm preflight && pnpm typecheck && pnpm lint && pnpm format:check` (no test run — documented in AGENTS.md) | `pre-push` | QC-14 Pre-push gate | preserve | Composes QC-1..4 exactly as specified in AGENTS.md; tests intentionally left to CI. |
| Hook installer | `scripts/install-hooks.mjs` (`pnpm prepare`) | ts | Stages husky hooks into a scratch repo first, asserts husky produced its `h` shim, quarantines incomplete previous sets, verifies idempotence via full directory snapshot comparison, then sets `core.hooksPath`; skips cleanly with `HUSKY=0` or no git | `installer` (runs on `pnpm install`) | QC-13 Staged-change hygiene | preserve | Self-verifying wiring for the hook gates — a corrupted half-install cannot silently disable them. |
### 1.6 CI pipeline steps (`.woodpecker/`)
Step-to-probe mapping for container steps: `ci.yml#sanitization` = QC-7+QC-8+QC-9 (rows §1.2, plus `apk add bash` env prep); `ci.yml#upgrade-guard` = QC-10 (rows §1.2, plus `apk add rsync`); `ci.yml#typecheck`/`#lint`/`#format`/`#test` = QC-2/3/4/5 (rows §1.1). Rows below are mechanisms that exist only in a pipeline.
| check | location | kind | what it actually verifies | enforcement point | canonical check | disposition | rationale |
| -------------------------------------- | -------------------------------------------------------------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------- | ----------------------------------- | --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------- |
| Frozen install | `ci.yml#install` | pipeline-step | `pnpm install --frozen-lockfile --prefer-offline` against the baked ci-base store — lockfile supply integrity; a drifted lockfile fails the build before any gate runs | `CI ci.yml#install` | QC-1 Checkout integrity | preserve | Lockfile-pinned dep resolution is the supply-chain floor under every later gate. |
| Test-step readiness prelude | `ci.yml#test` prologue | pipeline-step | Installs pinned `@earendil-works/[email protected]` (Invariant R suite requires the real binary) + openssl; waits up to 60×1s on `pg_isready` for the `ci-postgres` service and fails fast if it never comes up; runs `db:migrate` before tests | `CI ci.yml#test` | QC-5 Test suite execution | preserve | Fail-fast environment preconditions — a missing service produces a legible failure, not a wall of red tests. |
| Publish verify step (pending RI-1-001) | `publish.yml#verify` (branch `feat/ri-050-publish-gate` @ `46784c8d`, not yet on next) | pipeline-step | (a) Commit identity: fails closed if `CI_COMMIT_SHA` empty, `git rev-parse HEAD` empty, or the two differ; (b) runs the canonical `pnpm verify:release`. **Every publish effect depends on this step; it carries no path filter** | `publish.yml#verify` | QC-11 Terminal release verification | preserve | The RI-N1 exact-commit binding; until it merges, publish steps on next depend on `build` only (see §3 gap 1). |
| Publish error classification | `publish.yml#publish-npm` | pipeline-step | Publishes `@mosaicstack/*` (minus web) and classifies outcome: success, or the **only tolerated failure** = already-published (EPUBLISHCONFLICT / "cannot publish over" / "previously published"); explicit fatal on npm `E404/E401/ENEEDAUTH/ECONNREFUSED/ETIMEDOUT/ENOTFOUND` and on any unrecognized failure (replacing the old ` | | echo` that hid a registry 404) | `publish.yml#publish-npm` (main/tags, path-filtered on `packages/**`) | QC-12 Publish-effect integrity | preserve | Converts silent publish fall-on-floor into loud failure; allowlist-of-one error tolerance is the right shape. |
| Next-lane publish assertions | `publish.yml#publish-next-npm` | pipeline-step | Guards: branch must be `next`, `CI_PIPELINE_NUMBER` required; registry dist-tags JSON must be usable; walks all manifests, strictly parses stable semver, rewrites `X.Y.(Z+1)-next.<N>`; publishes with `--tag next` (never latest); post-publish asserts `npm view @mosaicstack/mosaic@next` resolves to the exact expected version | `publish.yml#publish-next-npm` (push/manual on next) | QC-12 Publish-effect integrity | preserve | Durable prerelease lane with end-to-end resolution proof — the published artifact is verified, not assumed. |
| Image destination policy | `publish.yml#build-gateway` / `#build-appservice` / `#build-web` | pipeline-step | Kaniko builds with destination policy: `next` ⇒ sha-tag only (fatal if a tag event sneaks in); `main` ⇒ sha + `latest`; tag events ⇒ sha + `<tag>`; anything else fatal. Path filters only skip **effects**, never the verify step | `publish.yml#build-*` | QC-12 Publish-effect integrity | preserve | Fail-closed tagging matrix; the exclude-list default-safe design keeps stale images impossible. |
Adjacent pipeline surface (not a probe): `.woodpecker/ci-image.yml` rebuilds the ci-base image on `pnpm-lock.yaml`/`Dockerfile.ci` change with an immutable `lock-<hash>` tag; pipelines consume `:latest`. Recorded for completeness — no code-quality property is checked.
### 1.7 Root installer tooling (`tools/`)
| check | location | kind | what it actually verifies | enforcement point | canonical check | disposition | rationale |
| --------------------------- | --------------------------------------------------------- | ----- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------- | ------------------------------- | ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Next-lane installer test | `tools/install-next-lane.test.sh` (`pnpm test:installer`) | shell | Drives `tools/install.sh --next` with faked `node`/`npm` binaries (no network): Node 20 must be rejected; installs must pin **exact** versions (mutable `@next` forbidden); fast path must not unexpectedly fall back to source; gateway-install failure takes the documented fallback | `turbo`-external tail of `pnpm test``CI ci.yml#test` | QC-5 Test suite execution | preserve | Hermetic (shimmed) regression net for the installer lane; runs as part of the standard test command. |
| Clean-container install E2E | `tools/e2e-install-test.sh` | shell | Full first-run flow in a node:22-alpine container: `install.sh --yes``mosaic wizard` (non-interactive) → `mosaic gateway install``mosaic gateway verify` exit check (with EXPECTED-SKIP if the installed CLI predates `gateway verify`); skips gracefully without Docker | `local` (manual; requires Docker); **not wired in CI** | QC-5 Test suite execution | strengthen (review) | The only end-to-end proof of the install→verify path; currently operator-initiated only — wire into a periodic/manual CI lane or sign its exclusion explicitly. |
| Host installer advisories | `tools/install.sh` (`--check`; `check_fleet_transport`) | shell | `--check` = version comparison only, no install; `check_fleet_transport` warns (non-blocking, by design — tmux is the fleet's dependency, not mosaic's) when the roster-declared transport binary is absent, naming exactly what it blocks; PATH-persistence warnings | `installer` (operator-run) | QC-18 Operator-host drift audit | preserve | Advisory-by-design warnings; the parallel doctor check is drift-tested by §1.3's harness. |
### 1.8 Pending workstream additions (branch `feat/ri-050-publish-gate` @ `46784c8d`)
| check | location | kind | what it actually verifies | enforcement point | canonical check | disposition | rationale |
| ------------------------------- | ---------------------------------------------------- | ---- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- | ----------------------------------- | ----------- | ---------------------------------------------------------------------------------------------- |
| Canonical terminal verification | `scripts/verify-release.mjs` (`pnpm verify:release`) | ts | One command replaying the full mandatory set as stages — sanitization, upgrade-guard, typecheck (incl. preflight), lint, format, test, build — mirroring `ci.yml` step-for-step; fail-fast on first failing command; requires `bash`+`rsync` on PATH; `--stage <name>` for wiring smoke-tests only | `publish.yml#verify` (pending); `local` (`pnpm verify:release`) | QC-11 Terminal release verification | preserve | The RI-N1 canonical command — CI and publication share one semantic checklist by construction. |
| Verify-parity contract test | `scripts/verify-release.test.mjs` | ts | Parses the real `ci.yml`/`publish.yml`: stage table must match ci.yml step-for-step; every publish-effect step (name `publish*` or image-pushing) must transitively depend on `verify`; commit-identity assertion must be present; `verify` must carry no path filter | `test:checkout``CI ci.yml#test` (once merged) | QC-11 Terminal release verification | preserve | Guard-of-the-guard at checkout time — the two surfaces cannot drift apart silently. |
## 2. Canonical check set
The deduplicated checks every row above maps onto. IDs are stable for RI-3-002 to consume.
- **QC-1 Checkout integrity.** Owns: the checkout can run its gates — frozen-lockfile dependency resolution, required gate binaries present, no stale build lock, and the `apps/web/.next` generated-state trust chain (real directory, uid ownership, certified source fingerprint, certified symlink manifest). Implemented by `scripts/preflight.mjs` + frozen install steps.
- **QC-2 Workspace typecheck.** Owns workspace-wide TypeScript soundness: per-package `tsc --noEmit` over built dependencies (`turbo typecheck`). The single definition invoked by CI, pre-push, and terminal verification.
- **QC-3 Workspace lint.** Owns static-analysis policy: per-package ESLint under the root config. One config, one task, every surface.
- **QC-4 Format check.** Owns formatting uniformity: Prettier check with the repo ignore list. (The pre-commit variant additionally fixes; the verdict form is this check.)
- **QC-5 Test suite execution.** Owns execution of all test surfaces: checkout script units (`node --test`), per-package Vitest suites (including the framework shell chain and its python unitests), the installer-lane shim test, and — once wired — `test-roster-schema.py` and container E2E. Also owns guards-of-the-gate that live inside the chain (terminal-green contract, RCE regression).
- **QC-6 Workspace build.** Owns artifact buildability: `turbo build` producing the artifacts publication consumes.
- **QC-7 Framework sanitization.** Owns the open-source guarantee for the shipped framework package: no operator-identity tokens anywhere (examples included), no private `$HOME` defaults in shipped scripts, with a self-test that keeps the regexes honest.
- **QC-8 Resident-context budget.** Owns the line-count ceilings on framework files injected into every agent's context (Constitution, dispatcher, RUNTIME.md slices) — the CI-enforceable half of the resident-prompt budget.
- **QC-9 Test-membership enumeration.** Owns the property that no test suite can silently fall out of CI: disk population vs parsed enumeration surfaces, both-directions staleness, and signed exclusions with reasons. Includes its needle/control harness.
- **QC-10 Upgrade/install safety.** Owns the #791 family: operator-path byte-identity across keep-mode upgrades (manifest guard), mid-failure rollback (errtrace-proven), durable pre-update snapshot + verify net + CWE-59 leaf guard, and the v2→v3 migration matrix with shell/TS parity.
- **QC-11 Terminal release verification.** Owns the RI-N1 exact-commit binding: commit-identity assertion plus one canonical command (`pnpm verify:release`) replaying the complete mandatory set, with every publish effect depending on it; plus the checkout-time parity/DAG contract test that keeps pipeline and command in sync.
- **QC-12 Publish-effect integrity.** Owns publication correctness: npm publish error classification (only already-published tolerated), next-lane versioning and post-publish resolution proof, and image destination/tag policy.
- **QC-13 Staged-change hygiene.** Owns commit-time hygiene on staged files (prettier/eslint fix-and-restage) and the self-verifying hook wiring that guarantees the gates are actually installed.
- **QC-14 Pre-push gate.** Owns the local push composition: preflight + typecheck + lint + format:check (tests deliberately deferred to CI).
- **QC-15 Lease-gate architecture invariant.** Owns "no ungated runtime launches in production code": the scan + allowlist over `packages/`, `apps/`, `plugins/`, `tools/`.
- **QC-16 Agent-runtime edit-time checks.** Owns edit-time feedback on agent hosts: the deps-preflight legibility sentinel and typecheck-on-edit, plus their regression harnesses.
- **QC-17 Lease-enforcement wiring safety.** Owns the #869 C1/C2/C5 trio: activation capability probe (versioned contract), enforcement-hook wiring gate (default-deny with explicit opt-out), and the doctor check that surfaces a bricked host — with their shell/TS harnesses.
- **QC-18 Operator-host drift audit.** Owns host-state health CI cannot see: `mosaic doctor` drift audit (+ fleet transport, both implementations), `fleet doctor` roster classification, `gateway doctor`/`gateway verify` service health, and installer advisories. Advisory exits are part of the contract.
- **QC-19 Downstream rails presence check.** Owns "does a scaffolded project still carry its rails files" — today the TS `quality-rails check/doctor` presence loop; per RI-N4 this is the seed that must become the typed evaluator (presence alone is explicitly not parity).
- **QC-20 Downstream enforcement verification.** Owns "do the rails actually block" on scaffolded projects: the behavioral planted-commit probe (type error, `any`, lint, gitleaks secret) currently in `verify.sh`/`verify.ps1` behind the `mosaic-quality-verify` adapter.
- **QC-21 Downstream rails scaffolding.** Owns putting rails files into a target project: the shell template installer (+ PowerShell twin) and the TS `quality-rails init` scaffolder — currently two paths that must converge.
## 3. Coverage gaps
Enforced nowhere but implied, or named in docs/tooling but not wired:
1. **Publication not yet bound to verification on `next`.** At this base (`8199261c`), `publish.yml` publish steps depend on `build` only; the `verify` step and `scripts/verify-release.mjs` exist on `feat/ri-050-publish-gate` (`46784c8d`) but are not merged. Until RI-1-001 lands, AC-RI-1's negative control cannot hold on the real pipeline.
2. **Playwright E2E unwired.** `apps/web` ships `test:e2e` (`playwright test`) with real suites (`admin/auth/chat/navigation.spec.ts`); neither `pnpm test` nor any CI step invokes it. The web UI's user flows are verified only when an operator runs them manually.
3. **No secret scanning on this repo.** The framework's own template pre-commit makes gitleaks **required**, and `verify.sh` proves detection with a planted key — but this repository's `.husky/pre-commit` (lint-staged only) and CI run no secret scan. The repo ships the control it does not use.
4. **No dependency audit.** The quality `.woodpecker.yml` templates and `docs/CI-SETUP.md` specify `npm audit --audit-level=high` as a pipeline stage; nothing equivalent runs for this repo.
5. **No coverage thresholds.** Templates enforce 80% Jest coverage thresholds; this repo's Vitest configs collect coverage with no thresholds — coverage is measured nowhere and enforced nowhere.
6. **`test-roster-schema.py` invisible.** A real jsonschema regression suite wired to no surface and invisible to the enumeration guard (its population is `*.sh`; the suite is `.py`). Either enumerate it or sign an exclusion — silence here is the #1017 defect shape.
7. **Presence-checker expectations ≠ this repo.** `quality-rails check` expects `.eslintrc`, `biome.json`, `.githooks/pre-commit`, `PR-CHECKLIST.md` for node projects — none describe this monorepo (husky, flat eslint config, no biome, no PR-CHECKLIST.md). The evaluator's check set must be per-subject (versioned, digested), not one global file list.
8. **Chain-ordering residual (documented).** `test:framework-shell` is one `&&` chain: a failing link skips every later suite while the step still fails (measured in #1270 — four suites after position 44 had not run since a prior merge). The enumeration guard proves naming, not reachability; both residuals are in-file documented but structurally unfixed.
9. **Signed-exclusion burndown open.** 16 signed exclusions remain in `test-enumeration-exclusions.txt`; several are "unmeasured in CI image" or blocked on missing CI tooling (tmux, setsid) — tracked under #1017/#1271. Each is an enforcement promise deferred, not delivered.
10. **Windows twins unexercised.** `verify.ps1`, `install.ps1`, `mosaic-doctor.ps1` have no runner anywhere (no Windows CI); behavioral drift from their bash twins is undetectable by construction.
11. **QA hook name vs behavior.** `qa-hook-handler.sh` files remediation report templates but performs no verification; the seam's actual gate value is only the deps-preflight sentinel. Anything relying on "QA automation hook" as a check is relying on report-filing.
12. **Two test paths, one gated.** CI runs tests against ci-postgres (`DATABASE_URL` set); the local PGlite path is the documented default (AGENTS.md) until KBN-101-02/101-05. Only the CI path is enforced by pipeline.
## 4. Disposition summary
| disposition | rows | checks |
| ------------------- | ---- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| preserve | 43 | Every canonical owner (QC-1..QC-18) plus correct guards-of-the-guard and thin adapters: all of §1.1, the CI-invoked framework probes and adapters in §1.2, all of §1.3, the C1/C2/C5 trio and doctors in §1.4, all of §1.5, all pipeline-only steps in §1.6, §1.7 rows 1 and 3, and §1.8. |
| strengthen | 2 | `quality-rails check` and `quality-rails doctor` (QC-19) — the RI-N4 evaluator seed: typed verdicts, versioned/digested check definitions, per-subject check sets. |
| strengthen (review) | 9 | `verify.sh` + `verify.ps1` (QC-20), quality `install.sh`/`install.ps1` + `quality-rails init` (QC-21 — scaffold-path convergence), `test-roster-schema.py` (QC-5 — wire or sign), `qa-hook-stdin.sh` seam + `typecheck-hook.sh` (QC-16), `tools/e2e-install-test.sh` (QC-5 — CI lane). |
| retire | 0 | None meet the bar: RI-N4 requires effective shell probes be **absorbed before** their paths retire, and no absorption exists yet. The `strengthen (review)` rows are the retirement candidates for RI-3-002 once the evaluator owns their behavior. |
Row total: 54. Canonical checks: 21 (QC-1..QC-21).
@@ -0,0 +1,83 @@
# #1099 pipefail + early-exit sweep
Baseline: `df4c591ab42aa1ae62c12935fdc0e772684864a0`
This is a site inventory, not a risk count. `FIXED` means the early-exiting consumer no longer has a piped upstream process whose SIGPIPE can become the result under `pipefail`. `NOT-LOAD-BEARING` means the pipeline status is explicitly discarded. `UNREACHABLE-AND-WHY` describes designed input, not a payload-size safety claim.
## Tranche 1 — runtime and general scripts
| Baseline site | Verdict | Construction / reason |
| --- | --- | --- |
| `tools/matrix-presence-harness/run.sh:38` | FIXED | nullglob array selects the first path; no pipeline |
| `tools/e2e-install-test.sh:139` | FIXED | capture help completely, then grep via redirection |
| `tools/install.sh:312` | FIXED | NUL `mapfile` reads all roots; count != 1 reaches the named malformed-archive diagnostic |
| `scripts/analysis/reflect-board-history.sh:76` | FIXED | capture Git history completely, then grep via redirection |
| `scripts/analysis/reflect-git-history.sh:67` | FIXED | grep reads from a here-string |
| `scripts/analysis/reflect-git-history.sh:69` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/authentik/user-create.sh:72` | FIXED | jq `first(...)` reads the response directly |
| `packages/mosaic/framework/tools/git/mutate-push-guard.sh:87` | FIXED | grep `-m1` reads the file directly; downstream `cut` consumes its complete scalar output |
| `packages/mosaic/framework/tools/orchestrator/session-resume.sh:94` | FIXED | `mapfile` plus bounded indexed loop replaces `head` pipeline |
| `packages/mosaic/framework/tools/prdy/prdy-status.sh:69` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/qa/reflect-stop-hook.sh:172` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/qa/reflect-stop-hook.sh:173` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/qa/reflect-stop-hook.sh:174` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/qa/reflect-stop-hook.sh:175` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/qa/reflect-stop-hook.sh:176` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/qa/reflect-stop-hook.sh:177` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/qa/reflect-stop-hook.sh:178` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/qa/typecheck-hook.sh:16` | FIXED | Bash regex extracts the first field without a pipeline |
| `packages/mosaic/framework/tools/qa/typecheck-hook.sh:56` | FIXED | grep and bounded sed each read from a here-string |
| `packages/mosaic/framework/tools/tmux/send-message.sh:113` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/tmux/send-message.sh:124` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/wake/detector.sh:126` | FIXED | one awk reads the manifest directly and exits after the first exact key |
| `packages/mosaic/framework/tools/wake/detector.sh:270` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/wake/detector.sh:278` | FIXED | grep reads from a here-string |
| `packages/mosaic/framework/tools/wake/digest.sh:647` | FIXED | capture complete locator output, then select first line by parameter expansion |
| `packages/mosaic/framework/tools/wake/reconcile.sh:149` | FIXED | one awk reads the manifest directly and exits after the first exact key |
## Explicit withdrawn / non-load-bearing sites
| Baseline site | Verdict | Reason |
| --- | --- | --- |
| `tools/install.sh:182` | NOT-LOAD-BEARING | `|| true` explicitly discards lookup status |
| `tools/install.sh:356` | UNREACHABLE-AND-WHY | `pnpm pack` writes one matching CLI tarball into a fresh directory immediately before lookup; citation withdrawn in #1099 |
| `tools/install.sh:357` | UNREACHABLE-AND-WHY | same fresh-directory invariant for gateway tarball; citation withdrawn in #1099 |
| `tools/install.sh:627` | NOT-LOAD-BEARING | `|| true` explicitly discards lookup status |
| `scripts/agent/session-start.sh:70` | NOT-LOAD-BEARING | optional scratchpad lookup has `|| true` |
| `packages/mosaic/framework/templates/repo/scripts/agent/session-start.sh:58` | NOT-LOAD-BEARING | optional scratchpad lookup has `|| true` |
| `packages/mosaic/framework/tools/qa/qa-hook-stdin.sh:25` | UNREACHABLE-AND-WHY | withdrawn in #1099 after designed-input reachability measurement; preserved without re-litigation |
| `packages/mosaic/framework/tools/qa/qa-hook-stdin.sh:27` | UNREACHABLE-AND-WHY | same withdrawn designed-input finding |
| `packages/mosaic/framework/tools/qa/qa-hook-stdin.sh:30` | UNREACHABLE-AND-WHY | same withdrawn designed-input finding |
| `packages/mosaic/framework/tools/qa/qa-hook-stdin.sh:32` | UNREACHABLE-AND-WHY | same withdrawn designed-input finding |
| `packages/mosaic/framework/tools/qa/qa-hook-stdin.sh:34` | UNREACHABLE-AND-WHY | same withdrawn designed-input finding |
## Tranche 2 — non-wake test harnesses
All 22 baseline sites below are `FIXED`; the checked-in tranche fixture is passed through the same scanner and asserts all 22 occurrences and 21 normalized identities (the same response-split line occurs twice).
| Baseline site(s) | Verdict | Construction |
| --- | --- | --- |
| `systemd/user/test-fleet-units.sh:148` | FIXED | capture tmux output, then grep via redirection |
| `git/test-issue-comment-readback.sh:283,302` | FIXED | parameter expansion splits status/body without `head` |
| `git/test-pr-review-gitea-comment.sh:228` | FIXED | parameter expansion splits status/body |
| `git/test-lane-brief-pr-linkage.sh:72` | FIXED | grep reads from a here-string |
| `git/test-pr-review-repo-host-override.sh:225-226` | FIXED | grep reads from a here-string |
| `orchestrator/smoke-test.sh:67,72` | FIXED | parameter expansion selects first line |
| `orchestrator/test-board-roll.sh:99-100` | FIXED | grep reads from a here-string |
| `quality/scripts/test-upgrade-durable-snapshot.sh:180` | FIXED | complete sorted output is read with `mapfile`, then indexed |
| `quality/scripts/test-upgrade-rollback.sh:339,356` | FIXED | direct `grep -m1` file reads; cleanup captures before testing |
| `tmux/test-send-message-socket.sh:37,38,44-46,68,72` | FIXED | capture commands complete before redirected grep assertions |
| `tmux/test-send-message-verdict.sh:34` | FIXED | grep reads from a here-string |
## Tranche 3 — wake validation harnesses
All 26 baseline occurrences (25 normalized identities; one preimage selector occurs twice) are `FIXED` and mechanically bound through the wake fixture and shared scanner.
| Baseline site(s) | Verdict | Construction |
| --- | --- | --- |
| `wake/test-wake-digest-quarantine.sh:567` | FIXED | complete match populations are captured, then first line selected by parameter expansion |
| `wake/test-wake-preimage.sh:182-183,346-347` | FIXED | jq `first(...)` reads each JSONL file directly |
| `wake/validate-973/microtest-wake-assert.sh:153,170-171,176,204-209,233-234,251-252,286-287` | FIXED | scalar assertions use here-strings; diagnostics use non-early sed ranges; source line captured before matching |
| `wake/validate-973/validate-973.sh:110,119,180,182,187` | FIXED | scalar assertions use here-strings; diagnostic truncation uses consuming sed ranges |
The scoped inventory is complete: 26 runtime/general + 22 non-wake tests + 26 wake tests fixed; 11 explicitly withdrawn or non-load-bearing sites retain their documented verdicts.
+229
View File
@@ -0,0 +1,229 @@
# #1043 — Fleet pane git-identity propagation
## Objective
Ensure a fleet seat's launched runtime process receives its roster-derived `MOSAIC_GIT_IDENTITY`, and lock the complete generated-environment propagation boundary with an enumerated set comparison.
## Tracking
- External issue: `mosaicstack/stack#1043`
- Branch: `fix/1043-pane-git-identity`
- Coordinator: `tl-mosaic`
- `docs/TASKS.md`: read-only by project worker contract; not modified.
## Constraints
- RED-first bug reproducer is mandatory.
- R7 delete-the-subject mutation must turn the behavioral test red.
- Assert launched-process environment, not source text.
- One push only; do not poll CI after push.
- Run the CI queue guard immediately before push and report its `state=` line as state, not evidence.
- Do not modify a live host launcher or obtain/copy another credential.
- Self-post the PR, verify provider attribution, then stop.
- Final status wording: `believed-fixed, pending jarvis validation`.
## Scope inventory
Re-derived against `origin/main` at `85d2108e`:
- Launch consumer: `packages/mosaic/framework/tools/fleet/start-agent-session.sh`
- Behavioral launch test: `packages/mosaic/framework/tools/fleet/test-start-agent-session.sh`
- Generated-environment contract/parser: `packages/mosaic/src/fleet/generated-env-boundary.ts`
- Roster projection producers:
- `packages/mosaic/src/commands/fleet.ts`
- `packages/mosaic/src/fleet/fleet-reconciler.ts`
- `packages/mosaic/src/fleet/fleet-agent-crud.ts`
- `packages/mosaic/src/fleet/v1-v2-migration.ts`
- Contract and producer tests discovered by repository search.
- Generated-environment operator/developer docs and their executable documentation contract test.
Discrepancy sent to `tl-mosaic`: current main no longer contains the charter's `PANE_SHELL_SNIPPET`; #772 replaced it with an `/usr/bin/env -i` argv launch boundary, and current generated projections do not declare git identity. Code-read inventory is **NOT MEASURED** behavior.
## Plan
1. Add the process-environment set-comparison regression first and record RED.
2. Add roster-derived `MOSAIC_GIT_IDENTITY=<agent name>` to the complete generated projection contract.
3. Validate identity syntax and equality with `MOSAIC_AGENT_NAME`; pass it through the clean pane environment.
4. Update affected projection tests and generated-environment docs.
5. Run focused and baseline gates.
6. Perform R7 by deleting the pane propagation entry, prove RED, restore, and prove GREEN.
7. Run independent review, remediate, commit, queue guard, one push, self-post PR, verify provider attribution, and stop without CI polling.
## Budget
No explicit token cap was provided. Working cap: one narrow logical unit, no dependency installation unless existing tooling requires it, no unrelated refactor.
## Evidence log
### TDD and mutation evidence
- RED-first, repository launcher: `bash packages/mosaic/framework/tools/fleet/test-start-agent-session.sh` exited 64 on pre-fix source with `code=unknown-key key=MOSAIC_GIT_IDENTITY`. The generated seat could not launch with the required declared identity.
- GREEN: the same repository launcher test emitted `ok - start-agent-session generated environment boundary`.
- R7 delete-the-subject: removed only `"MOSAIC_GIT_IDENTITY=$MOSAIC_GIT_IDENTITY"` from the repository launch array; the same test exited 1 with `FAIL: runtime pane omitted or changed generated environment keys: MOSAIC_GIT_IDENTITY`.
- R7 restoration: restored that launch entry; the same test returned green.
- Launcher under test is explicitly `packages/mosaic/framework/tools/fleet/start-agent-session.sh` through the test's `$START`, **not** the stale installed host copy.
### Situational and focused tests
- Repository launcher boundary: green, including set comparison of all nine generated projection entries and fail-before-tmux cases for missing, unsafe, mismatched, and local-shadow Git identity.
- Fleet systemd launcher integration: `bash packages/mosaic/framework/systemd/user/test-fleet-units.sh` — green.
- Focused Mosaic Vitest set: 6 files, 311 tests — green.
- `bash -n` on changed shell files — green.
- `git diff --check` — green.
### Baseline gates
- `pnpm typecheck` — 45/45 tasks green.
- `pnpm lint` — 25/25 tasks green.
- `pnpm format:check` — green.
- `pnpm test:checkout` — green.
- Repository-wide Vitest under a hermetic current-version npm prefix: Mosaic 81/81 files and 1510/1510 tests green; other workspace test tasks shown green before the framework-shell phase.
- Canonical `pnpm test` is not fully green on this host for unrelated environment-sensitive gates:
1. the first two runs exposed the globally installed Mosaic 0.0.48 update banner in three CLI smoke tests expecting empty stderr;
2. after isolating that global-version input, the framework wake assertion aborted at the known `#973` Bash `BASH_LINENO` convention check (exit 97; observed `[3 5]`, expected `[3 4]`).
No tests were weakened or bypassed; focused changed-surface tests are green. CI remains the canonical clean-environment result and is intentionally not polled after push per charter.
### Independent review
- Codex code review first pass: request changes for missing shell rejection-path coverage.
- Remediation: added table-driven missing/unsafe/mismatch/local-shadow launcher cases, each asserting no tmux call.
- Codex code re-review: **approve**, no findings, confidence 0.88.
- Codex security review: risk `none`, no findings, confidence 0.97.
### Acceptance criteria mapping
| Acceptance criterion | Evidence |
| --- | --- |
| AC-FGI-01: launched process receives every generated key/value | Repository launcher process-environment `comm -23` set comparison; GREEN and R7 RED evidence above |
| AC-FGI-02: missing, unsafe, or split identity fails before tmux | Table-driven shell cases plus TypeScript generated-boundary tests |
| AC-FGI-03: focused/baseline/review evidence recorded | Commands and review outcomes above; host-sensitive full-suite limitations stated explicitly |
### Documentation checklist
- PRD updated with #1043 requirements and acceptance criteria.
- Fleet launch runbook, generated-env concept, and generated-env reference updated.
- No API/OpenAPI, sitemap, user publishing target, deployment, or external docs publication change applies.
- `docs/TASKS.md` remains unmodified per its single-writer project contract.
## Round 2 — PR #1073 review 97 remediation
### Review blocker
The launched-process suite was signed-excluded from CI enumeration. Manual GREEN/R7 evidence therefore did not prove a PR workflow could detect regression.
### RED-first and canonical wiring
1. Removed the suite's signed exclusion before adding a CI execution path.
2. `check-test-enumeration.sh` went RED with exact `UNENUMERATED` output for `test-start-agent-session.sh`: population 49, enumerated 30, excluded 18.
3. Added both `framework/tools/fleet/test-start-agent-session.sh` and `framework/systemd/user/test-fleet-units.sh` to `@mosaicstack/mosaic`'s canonical `test:framework-shell` chain.
4. The guard returned GREEN: population 49, enumerated 32, excluded 18, surfaces 45. The systemd suite is outside the guard's tools-only population but now has the same explicit canonical execution disposition.
### Workflow-level R7
- Deleted only the pane launch entry `"MOSAIC_GIT_IDENTITY=$MOSAIC_GIT_IDENTITY"`.
- Ran the exact `.woodpecker/ci.yml` test-step command, `pnpm test`, with only a temporary PATH-scoped npm shim reporting the checkout's current 0.0.49 version so the unrelated global 0.0.48 banner could not preempt the shell chain.
- Result: exit 1 at `@mosaicstack/mosaic#test`, with the enumeration guard GREEN followed by `FAIL: runtime pane omitted or changed generated environment keys: MOSAIC_GIT_IDENTITY`.
- Restored the launch entry. The canonical `test:framework-shell` chain then reached both newly wired suites and printed both GREEN markers before the known unrelated #973 host-only `BASH_LINENO` abort.
- An actual provider PR workflow on the intentionally broken mutant is **NOT MEASURED**: the one-push constraint forbids pushing a red mutant and then a repaired head. Local execution proves the exact PR workflow command and dependency chain go RED on the subject deletion; CI on the repaired pushed head remains canonical.
### Workflow population
- **DEFINED:** 3 workflows (`ci.yml`, `ci-image.yml`, `publish.yml`).
- **ELIGIBLE for `pull_request`:** 1/3 (`ci.yml`), based on top-level `when:` clauses.
- **REPORTED:** Round-1 exact-head provider read reported 1/1 eligible context (`ci/woodpecker/pr/ci`). Post-remediation-head reported count is **NOT MEASURED** by this seat because CI polling is prohibited; workflow definitions and eligibility did not change.
### Independent remediation review
- First Round-2 review identified a CI-image blocker: the newly wired launcher suite used Perl, which the Alpine CI base does not install.
- Replaced the suite's three Perl-only fixture mutations with POSIX/BusyBox-compatible `sed -i` substitutions; production behavior and assertions are unchanged.
- Codex re-review: **APPROVE**, confidence 0.93, no findings.
### Vitest denominator reconciliation
The PR's `311/311` is correct for its explicitly named six-file command at both the original and remediation worktrees:
- generated environment boundary: 24
- fleet documentation: 23
- Tess service profile: 6
- fleet regen command: 27
- fleet agent CRUD command: 22
- fleet command: 209
- total: **311**
Review 97 reported 312/312 without naming its six files. That is a different or miscounted population and cannot replace the command-scoped 311 denominator; the PR follow-up will name the exact files and arithmetic.
## Round 3 — Alpine stale-marker portability
### Objective and plan
- Replace the GNU-only relative-date fixture with a deterministic POSIX/BusyBox timestamp while preserving the required stale-marker assertion.
- Re-run the launcher suite in the canonical `ci-base:latest` Alpine image, then run applicable repository gates and independent review.
- Update the PR body to name the repeated GNU-host/Alpine-CI portability pattern, run the mandatory queue guard, push once, verify provider attribution, and stop without CI polling.
- Working budget: 8K tokens; scope is one fixture line plus delivery evidence. No production behavior changes.
### RED-first evidence
Before the fix, the canonical CI image command
`docker run --rm -v "$PWD:/work" -w /work git.mosaicstack.dev/mosaicstack/stack/ci-base:latest bash packages/mosaic/framework/tools/fleet/test-start-agent-session.sh`
exited 1 at the stale-marker setup with exact BusyBox output
`touch: invalid date '10 seconds ago'`. The prior fresh-marker assertions had already executed, matching pipeline 2233's failure location.
### Root cause and fix
The test used GNU `touch -d` relative-date parsing although the PR workflow runs on Alpine/BusyBox. The fixture now uses POSIX `touch -t 200001010000.00`, a fixed timestamp that is unconditionally stale; the stale assertion remains mandatory and was not made tolerant of missing timestamp metadata.
### Structural pattern
This is the third GNU-host/Alpine-CI portability defect in the lane: GNU `grep` multi-match counting, Perl-only fixture mutation, and GNU `touch -d` date parsing. The repeated cause is shell suites authored on a GNU host but executed in an Alpine CI image; durable prevention belongs in CI-image execution or portability lint, not assertion weakening.
### GREEN and quality evidence
- Focused launcher suite in `ci-base:latest`: exit 0, `ok - start-agent-session generated environment boundary`.
- Canonical test step in `ci-base:latest` with the pipeline's `pgvector/pgvector:pg17` service, readiness check, migration, and `pnpm test`: exit 0; 46/46 Turbo tasks; Mosaic 81/81 files and 1510/1510 tests; Gateway 57 passed/5 skipped files and 629 passed/11 skipped tests; enumeration 49 population / 32 enumerated / 18 signed exclusions / 45 named surfaces.
- The first image-only `pnpm test` attempt lacked the pipeline PostgreSQL service and failed only on connection refusal after the launcher suite was GREEN. The rerun supplied the canonical service precondition and passed.
- Canonical-image baseline: typecheck 45/45 tasks, lint 25/25 tasks, format check GREEN; `git diff --check` GREEN.
- Independent Codex code review: APPROVE, confidence 0.96, 2/2 Round-3 files, no findings.
- Independent Codex security review: risk none, confidence 0.99, 2/2 Round-3 files, no findings.
### Re-derived inventory and denominators
- Round-3 git delta: **2/2 files** — launcher suite and task scratchpad; 25 insertions / 1 deletion before evidence finalization.
- Full PR path inventory against `origin/main` at `85d2108e`: **19/19 changed paths**; Round 3 adds no new PR path.
- Workflow definition population: **1/3 pull-request-eligible** (`ci.yml` of `ci.yml`, `ci-image.yml`, `publish.yml`).
- Do not re-litigate the settled 311/312 populations; both are valid for their separately named Tess6 and CRUD-core7 sets.
## Round 4 — bound stale-marker observation
### Objective and plan
- Make the heartbeat assertion discriminate an initially stale native marker from a fresh marker without changing the production staleness threshold or shortening the polling window.
- Freeze only the sidecar's numeric observation clock during the stale-fixture arm so elapsed assertion time cannot turn a fresh mutant stale.
- Prove two independent mutants RED: disable production stale-marker detection while retaining the stale fixture; replace the stale fixture with a fresh marker. Restore the tree and prove GREEN in the canonical Alpine image.
- Re-derive the changed-path inventory, run applicable quality and independent review gates, commit with environment-only author/committer identity, queue-guard, push once, verify provider attribution using curl stdin config, and stop without CI polling.
- Working budget: 8K tokens. Scope is the launcher test and its scratchpad evidence; production launcher behavior remains unchanged.
### Root cause and bounded observation
The 30 × 0.1-second assertion window overlaps the production `now - marker > interval * 2 + 1` threshold at interval 1. Depending on second boundaries and load, a fresh marker can age past the threshold before the assertion ends. A focused pre-fix fresh-mutant attempt returned RED while Review 101's full-suite run returned GREEN; the differing result is itself timing dependence, not a discriminating assertion.
The test now supplies a fixed numeric epoch only to the stale-fixture sidecar. Its real marker mtime is still read from the filesystem, but assertion runtime cannot advance `now`. Date formatting still delegates to the image's real `/bin/date`. Neither the production threshold nor the 30 × 0.1-second polling window changed.
### Two-mutant RED / restored GREEN
All three runs used `git.mosaicstack.dev/mosaicstack/stack/ci-base:latest`:
1. **Stale-detection mutant RED:** replaced only the production stale-age predicate with `false` while retaining the fixed stale marker; suite exit 1 with `FAIL: heartbeat sidecar did not resume after native marker became stale or absent`.
2. **Fresh-marker mutant RED:** replaced only `touch -t 200001010000.00` with fresh `touch`; suite exit 1 with the same failed stale-resumption assertion. The fixed observation epoch kept the mutant fresh throughout all 30 polls.
3. **Restored tree GREEN:** suite exit 0 with `ok - start-agent-session generated environment boundary`.
### Re-derived inventory
- Round-4 delta: **2/2 files** — launcher test plus task scratchpad; production launcher delta is empty.
- Full PR inventory against `origin/main`: **19/19 paths**; Round 4 adds no path.
- Production stale threshold remains `now - marker > iv * 2 + 1`; assertion polling remains 30 × 0.1 seconds.
- Review 101's confirmed enumeration/workflow/CI and attribution evidence is accepted without re-polling or re-derivation.
## Residual risk
- Landing on `main` does not update the currently installed host launcher. Host framework installation/reseed and Jarvis live-seat validation are separate downstream events.
- Canonical CI result is pending and will not be polled by this seat.
@@ -0,0 +1,97 @@
# #1098 — Framework shell portability / red main
## Objective
Restore terminal-green `main` by making the `test-start-agent-session.sh` clean-environment assertion semantic and portable without removing either newly enumerated framework-shell suite.
## Scope
- Tracking issue: `mosaicstack/stack#1098`
- Branch: `fix/framework-shell-portability`
- Base: `origin/main` at `4fa2768962702d53e16e8b67ee6ad52ebcb0910e`
- Primary file: `packages/mosaic/framework/tools/fleet/test-start-agent-session.sh`
- Requirements source: `docs/PRD.md` § Framework shell assertion portability (#1098)
- Out of scope: deployed files under `~/.config/mosaic`, pnpm-store cleanup, checkout deletion, and changes to the launchers `/usr/bin/env -i` behavior.
## Acceptance criteria
1. The test inspects the captured NUL-delimited tmux argv semantically and accepts an adjacent `/usr/bin/env`, `-i` pair regardless of trailing payload size or pipe scheduling.
2. Missing `/usr/bin/env`, missing `-i`, and non-adjacent `-i` remain failures.
3. Failure output includes the observed argv records with stable indexes and shell escaping; it exposes no credentials because this fixture supplies only generated non-secret launch data.
4. The focused suite passes on the dev host and in the repository CI image; the blocking PR/main pipeline returns terminal green.
5. Independent review passes; PR is squash-merged and #1098 is closed only after merged-main CI is terminal green.
## Budget
- ASSUMPTION: 30K-token working budget; rationale: one shell-test defect plus full PR/CI lifecycle.
- Auto-reduction: focused shell and package gates first; rely on canonical Woodpecker for the full monorepo suite rather than duplicating a dependency install under constrained `/home`.
- Disk baseline before clone/build: `/home` 7.1G free (99% used), `/tmp` 2.4G free (92% used).
## Investigation
### First-hand CI evidence
- Public log: `GET https://ci.mosaicstack.dev/api/repos/47/logs/2269/53041`
- Decoded 1,436 entries (11 null `data` entries treated as empty log rows), 190,756 bytes.
- Failure: `FAIL: pane command did not clear its environment` immediately after the expected pane-PID warning.
- BusyBox primitives, complete assertion pipeline, real CI image, stale/current image digests, Turbo cache masking, gateway failure, and heartbeat-sidecar concurrent writing were independently excluded.
### Root cause
The assertion ends in:
```bash
printf '%s\n' "$pane_args" | tail -n +"$after_pane_env" | grep -qxF -- '-i'
```
The script has `set -o pipefail`. `grep -q` exits as soon as it finds the valid `-i` record. Upstream `tail`/`printf` can then receive SIGPIPE, making the aggregate pipeline nonzero even though grep returned 0 and the semantic property is true. This depends on payload size, pipe capacity, and scheduling, explaining a local/image pass with a CI failure.
Discriminating stress control with `/usr/bin/env` followed immediately by `-i`:
- 8,192-byte trailing payload: `printf=0 tail=0 grep=0`, aggregate 0.
- 16,384-byte trailing payload: `printf=0 tail=141 grep=0`, aggregate 141.
- 32,768+ bytes: `printf=141 tail=141 grep=0`, aggregate 141.
- A full-reading `grep -xF` control remained 0 for every payload.
This is a third branch omitted by the earlier present-vs-corrupted split: the pair can be present and intact while `pipefail` reports an upstream SIGPIPE.
## TDD plan
1. RED: preserve the one-off stress reproducer above and add an automated large-argv semantic regression that fails under the current pipeline implementation.
2. GREEN: parse the authoritative NUL-delimited capture into a Bash array and search for an adjacent `/usr/bin/env`, `-i` pair without a short-circuit pipeline.
3. Add negative controls for missing, detached, and reversed tokens.
4. On failure, print indexed `%q` argv records before returning nonzero.
5. Run focused suite, mutation controls, shell syntax/format checks, then repository baseline gates feasible without dependency installation.
6. Independent review, queue guard, push, PR, CI, coordinator merge authorization, squash merge, merged-main CI, issue close.
## Progress
- [x] Checkout created and based on `origin/main` `4fa27689`.
- [x] CI log decoded directly.
- [x] Root-cause stress control reproduced semantic match + aggregate pipeline failure.
- [x] RED evidence: intact `/usr/bin/env`, `-i` fixture produced component statuses `0/141/0` and aggregate 141 under the former `grep -q` pipeline; full-reading semantic control stayed 0.
- [x] GREEN implementation: direct NUL-argv adjacency parser, indexed diagnostics, and full-reading scalar predicates replace all load-bearing early-exit pipelines in this test.
- [x] Baseline/situational tests:
- focused launcher suite: PASS on GNU host and cached Alpine CI image;
- paired `test-fleet-units.sh`: PASS;
- enumeration guard: PASS (`population=53`, `enumerated=36`, `excluded=18`), 14/14 mutation needles;
- `bash -n`, ShellCheck, `git diff --check`: PASS;
- static denominator after change: zero load-bearing `grep -q`/`head`/`-m1` pipeline candidates in `test-start-agent-session.sh`;
- delete-the-subject mutation removing production `-i`: RED with 78 indexed argv records, byte count, and explicit boundary failure.
- [x] Independent review:
- first Codex review: request changes — negative fixtures did not each assert diagnostics;
- remediation: centralized predicate + diagnostic wrapper and exercised all four negative fixtures;
- second Codex review: APPROVE, 0 blockers/should-fix/suggestions;
- Codex security review: risk none, 0 findings.
- [ ] PR CI, formal fleet review, merge, merged-main CI, issue closure.
## Documentation disposition
- Updated canonical `docs/PRD.md` with FSP requirements and acceptance criteria.
- This is an internal test/reliability change with no API, user workflow, deployment, navigation, or publishing-surface change; no user/admin/API/sitemap update is required.
- `docs/TASKS.md` remains unchanged because the project contract makes it orchestrator-only.
## Risks
- The CI failure did not print its captured argv, so the exact CI payload is unavailable. The stress control proves the assertion is non-portable and can emit the exact false verdict; branch CI is the canonical confirmation that replacing it resolves pipeline 2269s failure class.
- Printing fixture argv is safe only while this tests projection remains non-secret. The diagnostic must stay scoped to the test capture and shell-escaped.
+41
View File
@@ -0,0 +1,41 @@
# #1099 — pipefail + early-exit sweep
## Scope and decisions
- Baseline `df4c591ab42aa1ae62c12935fdc0e772684864a0`, after #1100 removed its 35 sites.
- Split into review-sized non-closing tranches: runtime/general; tmux/git/quality tests; wake validation/tests.
- Do not equate class membership with demonstrated risk. Do not use payload size or pipeline stage count as a safety proxy.
- Preserve the issue's withdrawn findings for `qa-hook-stdin.sh` and the two fresh-directory `pnpm pack` lookups. Fix `install.sh:312` because malformed multi-root input must reach its named handler.
## Tranche 1 TDD
RED-first control: `node --test scripts/pipefail-early-exit.test.mjs` reported exactly 26 non-accepted runtime/general sites, including `install.sh:312`, and exited 1. A checked-in fixture generated from immutable baseline `df4c591a` records all 26 normalized sites; the control passes every fixture entry through the same scanner, asserts exact identity/count/uniqueness, and separately requires zero findings in the current tree. It also inventories accepted sites rather than silently excluding whole files.
Construction choices:
- here-string/file redirection for scalar grep assertions;
- full capture then parameter expansion for first-line selection;
- arrays/`mapfile` for complete populations;
- direct jq/awk/grep selection where one tool can express the property;
- no `|| true` added to a load-bearing assertion.
Site-by-site verdicts: `docs/reports/quality/1099-pipefail-sweep.md`.
## Tranche 2 TDD
Expanded the unconditional scanner over 11 non-wake test harnesses. RED named exactly 22 source lines; a second immutable-baseline fixture now asserts those 22 entries through the same scanner. Rewrites preserve command status by capturing producers before redirected assertions, use parameter expansion for line selection, and use complete `mapfile` populations where ordering matters. Current-tree finding count is zero for tranches 1 and 2.
## Tranche 3 TDD
Expanded the shared scanner over four wake validation harnesses. RED named 26 occurrences. The wake fixture asserts 26 occurrences / 25 normalized identities through the same scanner; all scalar assertions now use redirection, direct jq selection, complete capture, or consuming diagnostic ranges. Current-tree finding count is zero across the full scoped population.
## Verification so far
- `bash -n` on every changed shell script: pass.
- structural Node control: pass.
- `test-mutate-push-guard.sh`: 8/8 pass.
- `test-send-message-verdict.sh`: 3/3 pass.
- `test-send-message-socket.sh`: pass.
- Independent review 143 found two semantic regressions: a help-probe `|| true` changed the failure truth table, and an unguarded Git capture changed non-Git data-dir behavior from rc 0 + JSON to silent rc 128. Both received RED-first regressions before correction; help status is now separate and required, and Git status remains condition-guarded.
- Wake static inventory remains aligned at 261/261 after line-neutral rewrites; no static-set mismatch. Wake detector/reconcile/digest/preimage suites terminate at their existing fail-closed #973 `BASH_LINENO` environment probe (exit 97, observed `[3 5]`, expected `[3 4]`) before subject tests. No bypass or skip was used; canonical CI remains required.
- ShellCheck reports only pre-existing source-following, unused-variable, and untouched `ls | head` findings; no new diagnostic was introduced.
+156
View File
@@ -0,0 +1,156 @@
# #1150 — Pi persistent goal extension
- **Task ID:** ISSUE-1150 (no `docs/TASKS.md` row; that file is orchestrator-only)
- **Issue:** #1150`pi: add persistent /goal controller extension to Mosaic framework`
- **Branch:** `feat/1150-pi-goal-extension`
- **Mode:** Delivery
- **Status:** in progress
## Objective
Build and locally validate a Mosaic-owned Pi `/goal` extension. Source must ship from
`packages/mosaic/framework/runtime/pi/`, framework sync must deploy it under
`~/.config/mosaic/runtime/pi/`, and no extension/configuration asset may be written into `~/.pi`.
Pi's native session manager remains the owner of session entries.
## Scope and acceptance source
- Canonical requirements: `docs/PRD.md`, section **Pi Persistent Goal Loop (#1150)**.
- User intent: continuous goal orientation and status checking after each Pi turn and compaction,
tested locally before framework delivery.
- Documentation target: canonical in-repo user/developer/runtime docs; no external publication.
## Assumptions
- `ASSUMPTION:` Initial semantic verification uses two consecutive structured, evidence-bearing
reports from the working agent rather than a second model request after every turn. This keeps the
loop testable and avoids doubling model cost while making the limitation explicit.
- `ASSUMPTION:` Default autonomous bounds are 40 turns and 6 repeated no-progress reports, with only
bounded numeric environment overrides.
- `ASSUMPTION:` A local smoke copy to `~/.config/mosaic/runtime/pi/goal-extension.ts` is authorized by
the user's explicit request. Full framework reseed into the live home is not required for the smoke
test and would touch unrelated framework-owned files.
## Budget
- Working estimate: 30K implementation/review tokens.
- Hard user cap: none stated.
- Cost control: deterministic fake-Pi tests; no nested evaluator calls; only bounded arithmetic/load
smoke workflows against the installed runtime.
## Plan
1. Update PRD and create tracking/scratchpad artifacts.
2. Read launcher, installer ownership, Pi extension, and documentation surfaces.
3. TDD: add fake-Pi behavior tests for commands, state restoration, turn checks, compaction, limits,
verification, and continuation deduplication.
4. Implement `runtime/pi/goal-extension.ts` and deterministic launcher discovery.
5. Add framework-sync/deployment acceptance coverage.
6. Update user, developer, runtime, framework README, and sitemap documentation.
7. Run focused tests, local Mosaic-path smoke test, then baseline repository gates.
8. Run independent review, remediate, commit, push/PR/CI/merge/issue closure per delivery gates.
## TDD decision
Applied. The continuation state machine and lifecycle scheduling are control-path logic where a race
or false terminal state can cause unbounded work or premature completion.
## Progress checkpoints
- [x] Issue #1150 created through Mosaic wrapper.
- [x] Isolated worktree created from `origin/main`.
- [x] PRD requirements and acceptance criteria added.
- [x] Task scratchpad created.
- [x] RED controller and security-regression tests written and observed failing before implementation.
- [x] Goal controller, launcher discovery, framework deployment coverage, and bounded state machine
implemented.
- [x] User, admin, developer, runtime, adapter, README, and sitemap documentation updated.
- [x] Final source copied additively to `~/.config/mosaic/runtime/pi/goal-extension.ts`; source and
deployed SHA-256 are identical.
- [x] Live Pi RPC smoke from the exact Mosaic path reached `achieved` with two verification passes and
no extension errors.
- [x] Baseline and situational checks completed, except the explicitly documented unavailable
PostgreSQL-only root integration case.
- [x] Independent code and OWASP/security reviews completed; all findings remediated and re-reviewed.
- [ ] Commit, push, PR, terminal-green CI, squash merge, and issue closure complete.
## Tests and evidence
### Situational
- `pnpm --filter @mosaicstack/mosaic exec vitest run src/runtime/pi-goal-extension.spec.ts`
- final: 25 passed.
- Covers commands, per-turn checks, context injection, two-pass verification, mixed-report
rejection, bounded limits, compaction, branch restore, stale timers, credential redaction,
typed-field false-positive protection, and append-only legacy-state fail-closed behavior.
- Final focused launcher/controller/file-adapter run: 3 files / 67 tests passed.
- Final V8 coverage for `framework/runtime/pi/goal-extension.ts`:
- 99.17% statements/lines, 93.78% branches, 100% functions.
- Installer migration fixture: 24 passed and byte-compared the deployed framework asset.
- Standalone extension TypeScript check against installed Pi 0.84.1 types passed:
`pnpm --filter @mosaicstack/mosaic exec tsc --noEmit --pretty false --module NodeNext
--moduleResolution NodeNext --target ES2022 --skipLibCheck framework/runtime/pi/goal-extension.ts`.
- Live deployment/load evidence:
- source/deployed SHA-256:
`1f0a3806e0948ad5f49684273a7e535e9880c148f7fd16d13ee487fcd601f637`.
- `get_commands` identified `/goal` as an extension command sourced from
`~/.config/mosaic/runtime/pi/goal-extension.ts`; `/goal help` succeeded; zero extension errors.
- live arithmetic goal ended `achieved`, verification `2/2`, with 3 goal reports / 3 agent starts
and zero extension errors.
- no goal extension exists under `~/.pi` extension paths.
### Baseline
- `pnpm build`: passed before the final framework-only redaction remediation; the extension is not a
package build input and its final source passed the standalone Pi type check.
- `pnpm typecheck`: 45/45 tasks passed.
- `pnpm lint`: 25/25 tasks passed.
- `pnpm format:check`: passed.
- Final Mosaic package components:
- Vitest: 82 files / 1,539 tests passed.
- full `test:framework-shell` harness passed.
- the discovered pre-existing tmux loader-marker race was reproduced with constructor PID
evidence, fixed with a pane readiness/FIFO barrier, passed 3 consecutive focused runs, and passed
in the full shell harness.
- one combined rerun encountered the separate existing real-lease probe TOCTOU in
`install-ordering-guard.spec.ts`; an earlier final Vitest run was fully green and the changed
focused suites remained green.
- Gateway safe baseline excluding the prohibited PostgreSQL-only fixture: 55 files / 600 tests passed
(6 files / 12 tests skipped by their existing environment gates).
- Root `pnpm test` reached 43 successful workspace tasks and all changed-package Vitest tests, but
the unchanged `apps/gateway/src/__tests__/cross-user-isolation.test.ts` afterAll hook retried a
PostgreSQL connection and failed authentication (`28P01`). This checkout explicitly forbids local
PostgreSQL startup/access; the failure is unrelated to #1150 and cannot be remediated by starting
the database. The gateway suite excluding that PostgreSQL-only file and required CI are used as
the safe verification paths.
### Independent review
- Codex code review: approved, 0 findings across 15 files.
- Initial Codex security review: one medium CWE-532/A09 finding for raw report persistence.
- Remediation added central credential-pattern redaction, prompt/docs guidance, canary tests, typed
field false-positive guards, and sticky fail-closed restore for credential-bearing append-only
history.
- Codex security re-review: risk `none`, 0 findings, confidence 0.87.
- Focused remediation review findings were fixed; final focused re-review verdict: `APPROVE`.
- Focused independent review of the tmux readiness barrier: `APPROVE`, no actionable findings.
## Risks and blockers
- Live `~/.config/mosaic` is shared by active Pi/fleet processes. Local deployment remained a single
additive framework file and did not reload or restart unrelated sessions.
- Completion verification is semantic, not mathematical: the active agent supplies structured
evidence twice. Operators must still inspect consequential outcomes.
- Credential redaction is pattern-based defense-in-depth, not a secret store. It covers
controller-owned state/status/tool details, not Pi's separate model-message/tool-call history.
Goals and reports must never contain real secrets or raw sensitive output. Because Pi session
entries are append-only, a detected credential-bearing legacy branch fails closed and the affected
session must be removed.
- Current installed Pi is newer than the repository's historical gateway Pi dependency. The
extension was checked and smoke-tested against installed Pi 0.84.1 using stable documented APIs.
- Local root testing cannot safely execute the unchanged PostgreSQL-only integration fixture under
the checkout's explicit database safety constraints. Terminal-green PR CI remains mandatory before
merge.
- The unchanged real-lease default-probe test can observe different broker availability across its two
sequential probes; one combined package rerun hit that existing TOCTOU. The same final Vitest suite
passed in a separate run, and CI remains the merge authority.
@@ -0,0 +1,89 @@
# #1174 — Wrapper guard rounds 1011
## Objective
Make checkout enforcement judge Git placement operands rather than every HOME-shaped word in the command, without reopening `--separate-git-dir` placement under HOME.
## Plan
1. Reproduce the four over-blocks and the placement-option control at head `20d86e39`.
2. Add RED fixtures before production changes.
3. Extract clone/worktree placement operands from the existing shell-aware normalized stream.
4. Run the full guard corpus, historical-head discrimination, syntax/static checks, probes, review, and CI.
## Progress and evidence
- Reproduced: `NOTE=$HOME`, `--reference=$HOME`, `GIT_DIR=$HOME/x`, and `--template=$HOME/t` all blocked despite explicit `/src/wt` destinations.
- RED at `20d86e39`: expanded suite had 8 failures, all HOME-valued non-placement cases.
- GREEN: expanded suite passes 242/242.
- Round-10 probes: 7/7 placement expectations and 4/4 placement-option controls pass.
- Earlier path probes remain green: 60/60, 24/24, and 17/17.
- Historical discrimination with the 242-fixture suite:
- `3d0a882a`: 216 pass / 26 fail.
- `4b8eba95`: 222 pass / 20 fail.
- `20d86e39`: 234 pass / 8 fail.
- `bash -n`, ShellCheck warning-or-higher, and `git diff --check`: pass.
## Residual / risk
- Relative destinations whose effective path depends on cwd are tracked separately by #1197 and remain out of scope.
- Unknown future Git options with a separate following value fail closed when that value is HOME-shaped. This may require classification when Git adds an unrelated path-taking option, but prevents a new placement option from silently bypassing the guard.
## Round 11 objective and intake
- **Issue / PR:** #1174.
- **Objective:** Remove the finite boolean-flag allowlists that turn accepted clone/worktree flags into fake placement operands, while preserving all real HOME placement blocks.
- **Scope:** `wrapper-guard.sh`, its hermetic fixtures, and task documentation. Relative cwd-dependent destinations remain in #1197.
- **Surfaces:** security-sensitive Bash hook behavior and shell/Git option grammar; no API, DB, UI, auth, deploy, or dependency changes.
- **Budget assumption:** 25K working tokens; reduce exploratory matrices before reducing acceptance coverage.
### Round 11 plan
1. Use Git itself to classify accepted/rejected clone and worktree options, and Bash itself to resolve path-word expectations.
2. Add RED fixtures for all six reported clone flags, generated negations, and equivalent worktree grammar.
3. Replace the open-ended unknown-option fail-closed fallback with a parser based on the closed value-taking option surface; keep explicit placement options special.
4. Run the full corpus, historical discrimination, shell/static checks, targeted probes, independent code/security review, one push, and exact-head CI.
### Root-cause evidence
- Git 2.39.5 accepts all six reported clone flags and the broader generated family measured in the brief: `--bare`, `--mirror`, `--ipv4`, `--ipv6`, `-4`, `-6`, `--no-local`, `--no-reject-shallow`, `--no-bare`, `--no-sparse`, `--no-dissociate`, `--no-shallow-submodules`, `--no-quiet`, `--no-progress`, and `--no-recurse-submodules`; it rejects `--relative-paths` as unknown.
- Git 2.39.5 accepts worktree negations including `--no-force`, `--no-detach`, `--no-lock`, `--no-guess-remote`, and `--no-track`; the current finite worktree flag list does not describe that generated family.
- `bash -c "printf '%s' <word>"` resolves `$HOME/source`, `${HOME}/source`, and `"$HOME"/source` under HOME while `/src/wt` remains outside it.
- **Hypothesis:** only separate-value options need positive classification. Treat every other option token as a no-value flag unless it is the explicit placement option; this matches Git's non-enumerable boolean family and confines the residual to genuinely new future value-taking options.
### TDD and verification checkpoints
- RED against the unmodified `91cc37bc` guard: 253 pass / 22 fail in the initial expanded 275-fixture suite. Failures include all 15 accepted clone flags, accepted long abbreviations, short value-taking bundles, abbreviated placement, worktree metadata abbreviation, and both directions of bundled worktree branch parsing.
- An exploratory fail-closed residual test drove emission of every worktree positional. Re-review correctly showed that this over-blocked HOME-shaped commit-ish metadata; a new commit-ish fixture failed RED against that intermediate implementation (278 pass / 2 fail, including one transient message assertion) and the parser was restored to emit only the actual path.
- GREEN after remediation: 280/280.
- Ultron's 13-shape option probe: 13/13 correct, including the six reported over-blocks, HOME destinations, end-of-options, worktree controls, and a later-command placement.
- Round-10 probes remain green: 7/7 subject-placement expectations and 4/4 `--separate-git-dir` controls.
- Earlier shell/path probes remain green: 60/60, 24/24, and 17/17.
- `bash -n`, ShellCheck warning-or-higher, and `git diff --check`: pass.
### Deliberate residual
A future Git release could add a new separate-value option absent from the closed value grammar. It defaults to no-value flag parsing, which leaves the following word positional. For clone, this can fail open if that future option itself creates repository state at its value. For worktree, it can shift which word is read as the path. This hypothetical future ambiguity is accepted deliberately because failing closed on every unclassified option is proven to over-block Git's open-ended present-day boolean/`--no-*` family. Every value-taking and placement option Git currently supports is classified, including accepted abbreviations of `--separate-git-dir`. Relative cwd-dependent targets remain in #1197.
### Independent review checkpoint
- Initial Codex code/security review raised `--orphan` as value-taking. Upstream Git `master` contradicts that premise: the synopsis is `[--orphan] [(-b | -B) <new-branch>] <path> [<commit-ish>]`, and the prose derives the branch from the path when `-b`/`-B` is absent. `--orphan` is therefore correctly handled as a boolean flag.
- The security review separately identified the generic future worktree shift residual. An attempted fail-closed remediation emitted every positional, but code re-review correctly rejected it because valid grammar has only one placement positional and an optional commit-ish. Final behavior checks only the path and documents the hypothetical future option shift deliberately; paired actual-grammar `--orphan` fixtures cover safe/HOME paths and `-b` metadata.
- Security re-review initially had no findings. Code re-review's commit-ish blocker was remediated with a RED fixture and path-only restoration; final code re-review approved with no findings.
- Final security review then found non-canonical absolute and symlink aliases. Eight lexical fixtures failed RED against the prior implementation, followed by three symlink fixtures failing RED. Remediation expands only shell-visible HOME tokens, resolves the longest existing directory prefix physically, and lexically normalizes the nonexistent suffix. The suite is now 292/292.
- Inherent residual: a symlink can be replaced between pre-tool inspection and Git execution. Existing aliases are resolved; eliminating the race requires enforcement inside the filesystem mutation path rather than a text pre-hook. Security review classified this medium, and architectural closure is tracked in #1199.
- Final independent code review: APPROVE, 0 findings. Final security review: no critical/high findings; the single medium TOCTOU residual is explicitly tracked in #1199.
### Final local evidence
- Final hermetic suite: 292/292; the same suite against `91cc37bc` discriminates at 256 pass / 36 fail.
- Ultron option probe: 13/13; round-10 probes: 7/7 plus 4/4 controls; earlier shell/path probes: 60/60, 24/24, and 17/17.
- `bash -n`, ShellCheck warning-or-higher, `git diff --check`, sanitization gate, and test-enumeration gate (population 55; 38 enumerated; 18 signed exclusions): pass.
- Independent code review: APPROVE, 0 findings. Security review's remaining medium TOCTOU architecture residual is tracked in #1199; no critical/high findings remain.
- Repository-wide TypeScript gates require dependencies absent from this worktree; the canonical Woodpecker pipeline will run them against the pushed exact head.
### Documentation checklist
- `docs/PRD.md` updated with WPG requirements, acceptance, canonicalization, and residual risk.
- Task scratchpad updated in the same logical change set; `docs/TASKS.md` remains orchestrator-only.
- No API, auth, UI, navigation, deployment, user-guide, or admin-guide surface changed; OpenAPI, endpoint index, sitemap, and publishing are not applicable.
@@ -0,0 +1,71 @@
# #1194 — Installed framework-tool drift detection and refresh analysis
## Decision
The reported queue-guard source defect was already fixed on `main` by `58b971ab`; the live failure came from a stale `~/.config/mosaic/tools/git/ci-queue-wait.sh`. The durable fix is therefore a detector, not a duplicate queue-guard patch.
`mosaic doctor` now compares the framework tools bundled with the executing Mosaic package against the deployed tools tree. Doctor is the selected visibility boundary because it is observational and operator-invoked: unlike session start, it does not add a repository/network scan to every seat launch, and it cannot silently replace identity or messaging tools while seats are active. It reports drift without changing files. `--fail-on-warn` converts detected drift into a non-zero doctor result.
## Classification
The existing `framework-manifest.txt` is authoritative. The detector invokes the canonical shared `tools/_lib/manifest.sh classify` implementation over the complete source census and refuses missing, unreadable, malformed, incomplete, or zero-framework ownership output. Policy is therefore read rather than duplicated:
- Current policy classifies source files under `tools/**` as framework-owned and required in the deployed tools tree.
- Current policy explicitly classifies `tools/_lib/credentials.json` operator-owned and excludes it from byte comparison; future policy changes take effect without a detector edit.
- A file present only in the deployed tools tree is operator-owned/unknown by the manifest's fail-safe default. The detector reports it as `INSTALLED_ONLY operator-or-unknown` under `--verbose` but does not fail or delete it.
- Empty/partial source traversal, unreadable directories/files, symlinked census entries, root aliases, and descendant source aliases all return `CANNOT_ASSERT` rather than manufacturing agreement.
This means `NOT_INSTALLED` is not suppressed by filename guesses such as “test” or “README”: if it ships below source `tools/**`, the installer contract says it should be installed. Source-only implementation files outside `tools/**` are outside this detector population by construction.
## Current host analysis (observation only; no refresh performed)
A direct source-vs-installed census showed broad drift, including identity and messaging behavior:
- Identity/provider operations: stale `git/detect-platform.sh`, `issue-comment.sh`, `issue-create.sh`, `issue-close.sh`, `issue-view.sh`, `pr-create.sh`, `pr-merge.sh`, `pr-review.sh`, `pr-metadata.sh`; missing `pr-edit.sh` and several identity/read-back regression tools.
- Messaging/session: stale `tmux/agent-send.sh`, `tmux/send-message.sh`, their regressions, and `fleet/start-agent-session.sh`.
- Gate enforcement: stale `git/ci-queue-wait.sh`; missing the queue tri-state/process-level suites and terminal-green verifier.
- Lease/QA behavior: stale lease-broker launch/mutation/receipt tools and QA hooks.
Counts vary with source head and installed local/operator files; the detector prints measured counts every run rather than baking this snapshot into policy.
## Reviewed refresh command — analyse only, do not run during active seats
Use the package/release updater's manifest-driven keep-mode sync during a quiet maintenance window:
```bash
MOSAIC_SYNC_ONLY=1 \
MOSAIC_INSTALL_MODE=keep \
MOSAIC_HOME="$HOME/.config/mosaic" \
bash /path/to/reviewed/@mosaicstack/mosaic/framework/install.sh
```
For the globally installed package, resolve the reviewed installer rather than guessing its path:
```bash
PACKAGE_ROOT="$(dirname "$(node -p "require.resolve('@mosaicstack/mosaic/package.json')")")"
MOSAIC_SYNC_ONLY=1 MOSAIC_INSTALL_MODE=keep MOSAIC_HOME="$HOME/.config/mosaic" \
bash "$PACKAGE_ROOT/framework/install.sh"
```
Do not run this while agent seats are active: the stale set includes identity selection, provider mutation, messaging, queue/merge guards, lease enforcement, and session launch. Syncing those files in place can change behavior between a seat's preflight and mutation.
## Post-refresh verification
1. Run `mosaic doctor --fail-on-warn`; require `stale=0 not-installed=0` from the framework drift summary (other unrelated doctor warnings must also be adjudicated).
2. Re-run the constructed process-level queue probes against the **installed path**, not the source checkout. Use the source suite while overriding its subject path in a reviewed scratch copy, or reproduce these exact observations:
- pending provider payload: guard must print `state=pending`, print the pending context, wait, and exit non-zero/timeout — never return immediately with rc 0;
- malformed payload: guard must print `state=malformed` and exit non-zero;
- unsupported but valid status vocabulary: guard must print `state=unknown` and exit non-zero.
3. Run provider author read-back for one deliberately low-risk wrapper operation before resuming fleet mutation work; wrapper self-report is not identity evidence.
4. Relaunch seats only after the quiet-window verification, because existing processes retain loaded environment/context.
## Probe evidence
The detector regression constructs a stale installed tool plus a missing shipped tool and observes rc 1 with distinct `STALE` and `NOT_INSTALLED` lines. That case would pass or be invisible before this change because no installed-vs-shipped comparison existed. Additional review-red controls prove:
- empty and unreadable source censuses return `CANNOT_ASSERT` (they returned clean rc 0 at the first PR head);
- deleting the manifest returns `CANNOT_ASSERT`, while changing manifest ownership changes the verdict through the canonical resolver (the first head never opened the manifest);
- root and descendant symlink/source aliases cannot return clean (the first head returned clean for a source-backed installed subtree);
- a checker hung during doctor is terminated by a bounded watchdog, emits `CANNOT_ASSERT`, and doctor reaches its final warnings line (the first head hung and suppressed the remaining audit).
Controls retain byte-identical success, exact credential carve-out behavior, and installed-only preservation.
+1
View File
@@ -34,6 +34,7 @@ export default tseslint.config(
'packages/storage/vitest.config.ts',
'packages/mosaic/vitest.config.ts',
'packages/mosaic/__tests__/*.ts',
'packages/forge/__tests__/*.ts',
'tools/federation-harness/*.ts',
],
},
+40
View File
@@ -539,3 +539,43 @@ Not every brief needs full Board of Directors review. The classification system
### Backward compatibility
Existing briefs without a `class` field are auto-classified. The default (no matching keywords) is `strategic`, so all existing runs get the full pipeline unless keywords trigger `technical`.
---
## Fail-Closed Execution & Explicit Simulation (SDLC-D-035)
**Added:** 2026-08-17
Forge fails closed when a required capability is missing. It never runs a
pipeline with a stub executor and reports success.
### Normal mode (default)
- No task executor wired → the CLI exits nonzero with the typed capability
error `FORGE_NO_EXECUTOR`. No run is created.
- A stage whose gate is approval-based (board approval, planning approvals,
remediation re-review, discovery/analysis attestations) records a typed
`waiting-for-authority` stage result and raises `FORGE_AUTHORITY_REQUIRED`.
It never passes vacuously.
- A stage whose gate requires an unwired provider (AI reviewer, CI pipeline)
records a typed `blocked` stage result and raises `FORGE_NO_REVIEWER` /
`FORGE_NO_CI_PIPELINE`. The synthetic echo-review approval in `06-review`
and all vacuous `true` gates were removed.
### Explicit simulation (`--simulate`)
Opts into stub/synthetic execution. Every stage result, every gate result, and
the run manifest carry the distinct typed status `simulated` (manifest also
records `mode: "simulated"`). `simulated` is a non-satisfying outcome:
`isSatisfyingOutcome()` and all completion/gate consumers treat only `passed`
as satisfying. The CLI exits 0 for a simulated run only because the caller
explicitly passed `--simulate`, and prints a loud SIMULATED banner.
### Typed outcome model
Every gate/task outcome is one of the closed set
`passed | failed | blocked | error | waiting-for-authority | simulated |
not-applicable`, with the reason recorded on the stage status and each gate
result in `manifest.json`. Missing implementations, missing gate evidence,
unknown stages, process errors, and timeouts map to fail-closed members —
never to `passed`.
@@ -0,0 +1,319 @@
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { generateBoardTasks } from '../src/board-tasks.js';
import { STAGE_SPECS } from '../src/constants.js';
import { ForgeCapabilityError } from '../src/errors.js';
import {
evaluateStageGates,
gateLabel,
isCommandGate,
isSatisfyingOutcome,
} from '../src/outcomes.js';
import { loadManifest, runPipeline } from '../src/pipeline-runner.js';
import type { ForgeTask, ForgeTaskResult, TaskExecutor } from '../src/types.js';
/**
* Mock real executor that returns typed results.
*
* Command gates are "verified" by the mock so normal-mode runs can pass
* mechanically gated stages; authority/provider gates are never reported
* because they have no mechanical implementation.
*/
function createTypedExecutor(options?: {
failStage?: string;
gateOutcomes?: Record<string, 'passed' | 'failed' | 'simulated' | 'error' | 'blocked'>;
}): TaskExecutor & { submittedTasks: ForgeTask[] } {
const submittedTasks: ForgeTask[] = [];
return {
submittedTasks,
async submitTask(task: ForgeTask) {
submittedTasks.push(task);
},
async waitForCompletion(taskId: string): Promise<ForgeTaskResult> {
const task = submittedTasks.find((t) => t.id === taskId);
const stageName = task?.metadata?.['stageName'] as string | undefined;
if (options?.failStage && stageName === options.failStage) {
return {
task_id: taskId,
outcome: 'failed',
reason: 'mock task failure',
completed_at: new Date().toISOString(),
exit_code: 1,
gate_results: [],
};
}
const gateResults = (task?.qualityGates ?? [])
.filter((gate) => isCommandGate(gate))
.map((gate) => {
const label = gateLabel(gate);
const outcome = options?.gateOutcomes?.[label] ?? 'passed';
return {
gate: label,
outcome,
reason: outcome === 'passed' ? 'mock verified' : `mock gate outcome: ${outcome}`,
};
});
return {
task_id: taskId,
outcome: 'passed',
reason: 'mock verified',
completed_at: new Date().toISOString(),
exit_code: 0,
gate_results: gateResults,
};
},
async getTaskStatus() {
return 'completed' as const;
},
};
}
describe('fail-closed: no executor wired', () => {
let tmpDir: string;
let briefPath: string;
beforeEach(() => {
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'forge-failclosed-'));
briefPath = path.join(tmpDir, 'brief.md');
fs.writeFileSync(briefPath, '# Fix bug\n\nA bugfix for lint cleanup.');
});
afterEach(() => {
fs.rmSync(tmpDir, { recursive: true, force: true });
});
it('throws a typed FORGE_NO_EXECUTOR capability error without --simulate', async () => {
await expect(
runPipeline(briefPath, tmpDir, {
// no executor, no simulate — must fail closed, never run with a stub
stages: ['00-intake'],
}),
).rejects.toMatchObject({
name: 'ForgeCapabilityError',
code: 'FORGE_NO_EXECUTOR',
capability: 'task-executor',
});
});
it('does not create a run directory when failing closed on a missing executor', async () => {
try {
await runPipeline(briefPath, tmpDir, { stages: ['00-intake'] });
} catch {
// expected
}
expect(fs.existsSync(path.join(tmpDir, '.forge', 'runs'))).toBe(false);
});
it('completes with every result typed simulated when simulate is set', async () => {
const result = await runPipeline(briefPath, tmpDir, {
simulate: true,
stages: ['00-intake', '00b-discovery', '02-planning-1', '06-review'],
});
expect(result.manifest.mode).toBe('simulated');
expect(result.manifest.status).toBe('simulated');
for (const stage of result.stages) {
const stageStatus = result.manifest.stages[stage];
expect(stageStatus?.status, `stage ${stage}`).toBe('simulated');
expect(stageStatus?.status, `stage ${stage}`).not.toBe('passed');
expect(stageStatus?.reason, `stage ${stage}`).toBeTruthy();
for (const gateResult of stageStatus?.gateResults ?? []) {
expect(gateResult.outcome, `gate ${gateResult.gate} of ${stage}`).toBe('simulated');
expect(gateResult.outcome, `gate ${gateResult.gate} of ${stage}`).not.toBe('passed');
}
}
// The persisted manifest agrees.
const persisted = loadManifest(result.runDir);
expect(persisted.mode).toBe('simulated');
expect(persisted.status).toBe('simulated');
expect(persisted.stages['02-planning-1']?.status).toBe('simulated');
});
});
describe('fail-closed: typed outcome model', () => {
it('only passed satisfies the gate/dependency predicate', () => {
expect(isSatisfyingOutcome('passed')).toBe(true);
expect(isSatisfyingOutcome('failed')).toBe(false);
expect(isSatisfyingOutcome('blocked')).toBe(false);
expect(isSatisfyingOutcome('error')).toBe(false);
expect(isSatisfyingOutcome('waiting-for-authority')).toBe(false);
expect(isSatisfyingOutcome('simulated')).toBe(false);
expect(isSatisfyingOutcome('not-applicable')).toBe(false);
});
it('a simulated gate result cannot satisfy the stage gate evaluation', () => {
const evaluation = evaluateStageGates('05-coding', STAGE_SPECS['05-coding']!.qualityGates, {
task_id: 'FORGE-x-05',
outcome: 'passed',
reason: 'executor claims success',
completed_at: new Date().toISOString(),
exit_code: 0,
gate_results: [{ gate: 'pnpm lint', outcome: 'simulated', reason: 'simulated gate' }],
});
expect(isSatisfyingOutcome(evaluation.outcome)).toBe(false);
expect(evaluation.outcome).toBe('error');
});
it('a simulated task outcome cannot satisfy evaluation in normal mode', () => {
const evaluation = evaluateStageGates('00-intake', [], {
task_id: 'FORGE-x-00',
outcome: 'simulated',
reason: 'executor reported simulated',
completed_at: new Date().toISOString(),
exit_code: 0,
gate_results: [],
});
expect(isSatisfyingOutcome(evaluation.outcome)).toBe(false);
});
it('a missing gate result blocks the stage instead of passing vacuously', () => {
const evaluation = evaluateStageGates('05-coding', STAGE_SPECS['05-coding']!.qualityGates, {
task_id: 'FORGE-x-05',
outcome: 'passed',
reason: 'executor claims success',
completed_at: new Date().toISOString(),
exit_code: 0,
gate_results: [],
});
expect(evaluation.outcome).toBe('blocked');
});
});
describe('fail-closed: authority and provider gates', () => {
let tmpDir: string;
let briefPath: string;
beforeEach(() => {
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'forge-authority-'));
briefPath = path.join(tmpDir, 'brief.md');
fs.writeFileSync(briefPath, '# Fix bug\n\nA bugfix for lint cleanup.');
});
afterEach(() => {
fs.rmSync(tmpDir, { recursive: true, force: true });
});
it.each(['02-planning-1', '03-planning-2', '04-planning-3', '07-remediate'])(
'planning/remediation stage %s yields waiting-for-authority (not passed) in normal mode',
async (stage) => {
const executor = createTypedExecutor();
let runDir: string | undefined;
try {
await runPipeline(briefPath, tmpDir, {
executor,
stages: [stage as string],
});
expect.unreachable('runPipeline should have failed closed');
} catch (err) {
expect(err).toBeInstanceOf(ForgeCapabilityError);
expect((err as ForgeCapabilityError).code).toBe('FORGE_AUTHORITY_REQUIRED');
runDir = path.join(tmpDir, '.forge', 'runs');
}
const runIds = fs.readdirSync(runDir!);
expect(runIds).toHaveLength(1);
const manifest = loadManifest(path.join(runDir!, runIds[0]!));
expect(manifest.stages[stage]?.status).toBe('waiting-for-authority');
expect(manifest.stages[stage]?.status).not.toBe('passed');
expect(manifest.status).toBe('waiting-for-authority');
},
);
it('review stage fails closed with a typed FORGE_NO_REVIEWER error in normal mode', async () => {
const executor = createTypedExecutor();
try {
await runPipeline(briefPath, tmpDir, {
executor,
stages: ['06-review'],
});
expect.unreachable('runPipeline should have failed closed');
} catch (err) {
expect(err).toBeInstanceOf(ForgeCapabilityError);
expect((err as ForgeCapabilityError).code).toBe('FORGE_NO_REVIEWER');
expect((err as ForgeCapabilityError).capability).toBe('reviewer');
}
const runsDir = path.join(tmpDir, '.forge', 'runs');
const runIds = fs.readdirSync(runsDir);
const manifest = loadManifest(path.join(runsDir, runIds[0]!));
expect(manifest.stages['06-review']?.status).toBe('blocked');
expect(manifest.stages['06-review']?.status).not.toBe('passed');
expect(manifest.status).toBe('failed');
});
it('review stage produces simulated results under --simulate', async () => {
const result = await runPipeline(briefPath, tmpDir, {
simulate: true,
stages: ['06-review'],
});
expect(result.manifest.mode).toBe('simulated');
expect(result.manifest.stages['06-review']?.status).toBe('simulated');
for (const gateResult of result.manifest.stages['06-review']?.gateResults ?? []) {
expect(gateResult.outcome).toBe('simulated');
}
});
it('deploy stage fails closed without a wired ci-pipeline provider in normal mode', async () => {
const executor = createTypedExecutor();
await expect(
runPipeline(briefPath, tmpDir, {
executor,
stages: ['09-deploy'],
}),
).rejects.toMatchObject({
name: 'ForgeCapabilityError',
code: 'FORGE_NO_CI_PIPELINE',
});
});
});
describe('fail-closed: no vacuous gate commands remain', () => {
it('stage constants contain no echo/synthetic-approval, vacuous true, or empty gate commands', () => {
for (const [stageName, spec] of Object.entries(STAGE_SPECS)) {
for (const gate of spec.qualityGates) {
const serialized = JSON.stringify(gate);
// The echo-review synthetic approval must be gone.
expect(serialized, `stage ${stageName} gate ${serialized}`).not.toContain('echo');
expect(serialized, `stage ${stageName} gate ${serialized}`).not.toMatch(/"verdict"\s*:/);
expect(serialized, `stage ${stageName} gate ${serialized}`).not.toMatch(
/"summary"\s*:\s*"review-pass"/,
);
// No vacuous literal `true` gate.
expect(gate, `stage ${stageName}`).not.toBe('true');
// Command gates must carry a real, non-empty command.
if (isCommandGate(gate)) {
const command = typeof gate === 'string' ? gate : gate.command;
expect(command.trim().length, `stage ${stageName} gate ${serialized}`).toBeGreaterThan(0);
}
}
}
});
it('board tasks contain no vacuous true gates', () => {
const tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'forge-board-gates-'));
try {
const tasks = generateBoardTasks('# Brief', [], tmpDir, 'BOARD-TEST');
for (const task of tasks) {
for (const gate of task.qualityGates) {
expect(gate, `task ${task.id}`).not.toBe('true');
const serialized = JSON.stringify(gate);
expect(serialized, `task ${task.id} gate ${serialized}`).not.toContain('echo');
}
}
} finally {
fs.rmSync(tmpDir, { recursive: true, force: true });
}
});
});
+161 -34
View File
@@ -12,10 +12,10 @@ import {
resumePipeline,
getPipelineStatus,
} from '../src/pipeline-runner.js';
import type { ForgeTask, RunManifest, TaskExecutor } from '../src/types.js';
import type { TaskResult } from '@mosaicstack/macp';
import type { ForgeTask, ForgeTaskResult, RunManifest, TaskExecutor } from '../src/types.js';
import { gateLabel, isCommandGate } from '../src/outcomes.js';
/** Mock TaskExecutor that records submitted tasks and returns success. */
/** Mock TaskExecutor that records submitted tasks and returns typed results. */
function createMockExecutor(options?: {
failStage?: string;
}): TaskExecutor & { submittedTasks: ForgeTask[] } {
@@ -25,7 +25,7 @@ function createMockExecutor(options?: {
async submitTask(task: ForgeTask) {
submittedTasks.push(task);
},
async waitForCompletion(taskId: string): Promise<TaskResult> {
async waitForCompletion(taskId: string): Promise<ForgeTaskResult> {
const failStage = options?.failStage;
const task = submittedTasks.find((t) => t.id === taskId);
const stageName = task?.metadata?.['stageName'] as string | undefined;
@@ -33,7 +33,8 @@ function createMockExecutor(options?: {
if (failStage && stageName === failStage) {
return {
task_id: taskId,
status: 'failed',
outcome: 'failed',
reason: 'mock task failure',
completed_at: new Date().toISOString(),
exit_code: 1,
gate_results: [],
@@ -41,10 +42,17 @@ function createMockExecutor(options?: {
}
return {
task_id: taskId,
status: 'completed',
outcome: 'passed',
reason: 'mock verified',
completed_at: new Date().toISOString(),
exit_code: 0,
gate_results: [],
gate_results: (task?.qualityGates ?? [])
.filter((gate) => isCommandGate(gate))
.map((gate) => ({
gate: gateLabel(gate),
outcome: 'passed' as const,
reason: 'mock verified',
})),
};
},
async getTaskStatus() {
@@ -156,12 +164,13 @@ describe('runPipeline', () => {
const executor = createMockExecutor();
const result = await runPipeline(briefPath, tmpDir, {
executor,
stages: ['00-intake', '00b-discovery'],
stages: ['00-intake', '05-coding'],
});
expect(result.runId).toMatch(/^\d{8}-\d{6}$/);
expect(result.stages).toEqual(['00-intake', '00b-discovery']);
expect(result.stages).toEqual(['00-intake', '05-coding']);
expect(result.manifest.status).toBe('completed');
expect(result.manifest.mode).toBe('normal');
expect(executor.submittedTasks).toHaveLength(2);
});
@@ -180,12 +189,17 @@ describe('runPipeline', () => {
const executor = createMockExecutor();
const result = await runPipeline(briefPath, tmpDir, {
executor,
stages: ['00-intake', '00b-discovery'],
stages: ['00-intake', '05-coding'],
});
const manifest = loadManifest(result.runDir);
expect(manifest.stages['00-intake']?.status).toBe('passed');
expect(manifest.stages['00b-discovery']?.status).toBe('passed');
expect(manifest.stages['05-coding']?.status).toBe('passed');
expect(manifest.stages['05-coding']?.gateResults?.map((g) => g.outcome)).toEqual([
'passed',
'passed',
'passed',
]);
});
it('respects CLI class override', async () => {
@@ -215,7 +229,7 @@ describe('runPipeline', () => {
const executor = createMockExecutor();
await runPipeline(briefPath, tmpDir, {
executor,
stages: ['00-intake', '00b-discovery', '02-planning-1'],
stages: ['00-intake', '05-coding', '08-test'],
});
expect(executor.submittedTasks[0]!.dependsOn).toBeUndefined();
@@ -224,14 +238,14 @@ describe('runPipeline', () => {
});
it('handles stage failure', async () => {
const executor = createMockExecutor({ failStage: '00b-discovery' });
const executor = createMockExecutor({ failStage: '05-coding' });
await expect(
runPipeline(briefPath, tmpDir, {
executor,
stages: ['00-intake', '00b-discovery'],
stages: ['00-intake', '05-coding'],
}),
).rejects.toThrow('Stage 00b-discovery failed');
).rejects.toThrow('Stage 05-coding failed');
});
it('marks manifest as failed on stage failure', async () => {
@@ -270,30 +284,143 @@ describe('resumePipeline', () => {
fs.rmSync(tmpDir, { recursive: true, force: true });
});
it('resumes from first incomplete stage', async () => {
// First run fails on discovery
const executor1 = createMockExecutor({ failStage: '00b-discovery' });
let runDir: string;
it('resumes from first incomplete stage and fails closed at the next provider gate', async () => {
// Simulate a run whose authority stages were approved out-of-band
// (recorded as passed) and whose coding stage failed mechanically.
const runId = '20260101-000000';
const runDir = path.join(tmpDir, '.forge', 'runs', runId);
fs.mkdirSync(runDir, { recursive: true });
const passed = { status: 'passed' as const, startedAt: '2026-01-01T00:00:00Z' };
saveManifest(runDir, {
runId,
brief: briefPath,
codebase: tmpDir,
briefClass: 'hotfix',
classSource: 'frontmatter',
forceBoard: false,
mode: 'normal',
createdAt: '2026-01-01T00:00:00Z',
updatedAt: '2026-01-01T00:00:00Z',
currentStage: '05-coding',
status: 'failed',
stages: {
'00-intake': passed,
'00b-discovery': passed,
'02-planning-1': passed,
'03-planning-2': passed,
'04-planning-3': passed,
'05-coding': { status: 'failed', reason: 'gate failed' },
},
});
try {
await runPipeline(briefPath, tmpDir, {
executor: executor1,
stages: ['00-intake', '00b-discovery', '02-planning-1'],
});
} catch {
// expected
// Resume re-runs 05-coding (the first non-passed stage), then fails
// closed at 06-review because no reviewer provider is wired.
const executor = createMockExecutor();
await expect(resumePipeline(runDir, executor)).rejects.toMatchObject({
name: 'ForgeCapabilityError',
code: 'FORGE_NO_REVIEWER',
});
const manifest = loadManifest(runDir);
expect(manifest.stages['05-coding']?.status).toBe('passed');
expect(manifest.stages['06-review']?.status).toBe('blocked');
expect(manifest.status).toBe('failed');
});
it('resumes to completion as simulated under explicit simulate', async () => {
const runId = '20260101-000003';
const runDir = path.join(tmpDir, '.forge', 'runs', runId);
fs.mkdirSync(runDir, { recursive: true });
const passed = { status: 'passed' as const, startedAt: '2026-01-01T00:00:00Z' };
saveManifest(runDir, {
runId,
brief: briefPath,
codebase: tmpDir,
briefClass: 'hotfix',
classSource: 'frontmatter',
forceBoard: false,
mode: 'normal',
createdAt: '2026-01-01T00:00:00Z',
updatedAt: '2026-01-01T00:00:00Z',
currentStage: '05-coding',
status: 'failed',
stages: {
'00-intake': passed,
'00b-discovery': passed,
'02-planning-1': passed,
'03-planning-2': passed,
'04-planning-3': passed,
'05-coding': { status: 'failed', reason: 'gate failed' },
},
});
const result = await resumePipeline(runDir, undefined, { simulate: true });
expect(result.manifest.status).toBe('simulated');
expect(result.manifest.mode).toBe('simulated');
expect(result.stages[0]).toBe('05-coding');
for (const stage of result.stages) {
expect(result.manifest.stages[stage]?.status).toBe('simulated');
}
});
const runsDir = path.join(tmpDir, '.forge', 'runs');
runDir = path.join(runsDir, fs.readdirSync(runsDir)[0]!);
it('fails closed on resume when the next stage needs authority sign-off', async () => {
const runId = '20260101-000001';
const runDir = path.join(tmpDir, '.forge', 'runs', runId);
fs.mkdirSync(runDir, { recursive: true });
saveManifest(runDir, {
runId,
brief: briefPath,
codebase: tmpDir,
briefClass: 'hotfix',
classSource: 'frontmatter',
forceBoard: false,
mode: 'normal',
createdAt: '2026-01-01T00:00:00Z',
updatedAt: '2026-01-01T00:00:00Z',
currentStage: '00-intake',
status: 'in_progress',
stages: {
'00-intake': { status: 'passed' },
},
});
// Resume should pick up from 00b-discovery
const executor2 = createMockExecutor();
const result = await resumePipeline(runDir, executor2);
const executor = createMockExecutor();
await expect(resumePipeline(runDir, executor)).rejects.toMatchObject({
name: 'ForgeCapabilityError',
code: 'FORGE_AUTHORITY_REQUIRED',
});
expect(result.manifest.status).toBe('completed');
// Should have re-run from 00b-discovery onward
expect(result.stages[0]).toBe('00b-discovery');
const manifest = loadManifest(runDir);
expect(manifest.stages['00b-discovery']?.status).toBe('waiting-for-authority');
expect(manifest.status).toBe('waiting-for-authority');
});
it('fails closed on resume without an executor or --simulate', async () => {
const runId = '20260101-000002';
const runDir = path.join(tmpDir, '.forge', 'runs', runId);
fs.mkdirSync(runDir, { recursive: true });
saveManifest(runDir, {
runId,
brief: briefPath,
codebase: tmpDir,
briefClass: 'hotfix',
classSource: 'frontmatter',
forceBoard: false,
mode: 'normal',
createdAt: '2026-01-01T00:00:00Z',
updatedAt: '2026-01-01T00:00:00Z',
currentStage: '00-intake',
status: 'in_progress',
stages: {
'00-intake': { status: 'passed' },
},
});
await expect(resumePipeline(runDir)).rejects.toMatchObject({
name: 'ForgeCapabilityError',
code: 'FORGE_NO_EXECUTOR',
});
});
});
+15 -2
View File
@@ -95,7 +95,14 @@ export function generateBoardTasks(
briefPath,
resultPath: resultRelPath,
timeoutSeconds: 120,
qualityGates: ['true'],
qualityGates: [
{
kind: 'authority',
capability: 'board-approval',
reason:
'persona evaluation is judged by board synthesis (authority review); no mechanical gate exists',
},
],
metadata: {
personaName: persona.name,
personaSlug: persona.slug,
@@ -121,7 +128,13 @@ export function generateBoardTasks(
timeoutSeconds: 120,
dependsOn: personaTaskIds,
dependsOnPolicy: 'all_terminal',
qualityGates: ['true'],
qualityGates: [
{
kind: 'authority',
capability: 'board-approval',
reason: 'board synthesis is an authority decision; no mechanical gate exists',
},
],
metadata: {
resultOutputPath: synthesisResult,
inputResultPaths: personaResultPaths,
+96 -1
View File
@@ -1,7 +1,11 @@
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { Command } from 'commander';
import { describe, expect, it } from 'vitest';
import { describe, expect, it, vi, beforeEach, afterEach } from 'vitest';
import { registerForgeCommand } from './cli.js';
import { loadManifest } from './pipeline-runner.js';
describe('registerForgeCommand', () => {
it('registers a "forge" command on the parent program', () => {
@@ -55,3 +59,94 @@ describe('registerForgeCommand', () => {
}).not.toThrow();
});
});
describe('forge run fail-closed behavior (SDLC-D-035)', () => {
let tmpDir: string;
let briefPath: string;
let errSpy: ReturnType<typeof vi.spyOn>;
let logSpy: ReturnType<typeof vi.spyOn>;
let prevExitCode: string | number | null | undefined;
const parse = (args: string[]) => {
const program = new Command();
registerForgeCommand(program);
return program.parseAsync(['forge', ...args], { from: 'user' });
};
beforeEach(() => {
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'forge-cli-failclosed-'));
briefPath = path.join(tmpDir, 'brief.md');
fs.writeFileSync(briefPath, '# Fix bug\n\nA bugfix for lint cleanup.');
errSpy = vi.spyOn(console, 'error').mockImplementation(() => {});
logSpy = vi.spyOn(console, 'log').mockImplementation(() => {});
prevExitCode = process.exitCode;
});
afterEach(() => {
errSpy.mockRestore();
logSpy.mockRestore();
process.exitCode = prevExitCode;
fs.rmSync(tmpDir, { recursive: true, force: true });
});
it('exits nonzero with a typed FORGE_NO_EXECUTOR error when no executor is wired and --simulate is absent', async () => {
await parse(['run', '--brief', briefPath, '--codebase', tmpDir]);
expect(process.exitCode).toBe(1);
const errText = errSpy.mock.calls.map((c) => c.join(' ')).join('\n');
expect(errText).toContain('FORGE_NO_EXECUTOR');
// It must never run the pipeline with a stub and report success.
expect(fs.existsSync(path.join(tmpDir, '.forge', 'runs'))).toBe(false);
});
it('completes with typed simulated results and exit 0 under explicit --simulate', async () => {
await parse(['run', '--brief', briefPath, '--codebase', tmpDir, '--simulate']);
expect(process.exitCode).toBeUndefined();
// Loud simulated-mode summary.
const logText = logSpy.mock.calls.map((c) => c.join(' ')).join('\n');
expect(logText).toContain('SIMULATED');
// Manifest records the mode and simulated per-result statuses.
const runsDir = path.join(tmpDir, '.forge', 'runs');
const runIds = fs.readdirSync(runsDir);
expect(runIds).toHaveLength(1);
const manifest = loadManifest(path.join(runsDir, runIds[0]!));
expect(manifest.mode).toBe('simulated');
expect(manifest.status).toBe('simulated');
for (const stageStatus of Object.values(manifest.stages)) {
expect(stageStatus?.status).toBe('simulated');
for (const gateResult of stageStatus?.gateResults ?? []) {
expect(gateResult.outcome).toBe('simulated');
}
}
});
it('resume exits nonzero with a typed FORGE_NO_EXECUTOR error without --simulate', async () => {
const runDir = path.join(tmpDir, '.forge', 'runs', '20260101-000000');
fs.mkdirSync(runDir, { recursive: true });
fs.writeFileSync(
path.join(runDir, 'manifest.json'),
JSON.stringify({
runId: '20260101-000000',
brief: briefPath,
codebase: tmpDir,
briefClass: 'hotfix',
classSource: 'frontmatter',
forceBoard: false,
createdAt: '2026-01-01T00:00:00Z',
updatedAt: '2026-01-01T00:00:00Z',
currentStage: '00-intake',
status: 'in_progress',
stages: { '00-intake': { status: 'passed' } },
}),
);
await parse(['resume', '20260101-000000', '--project', tmpDir]);
expect(process.exitCode).toBe(1);
const errText = errSpy.mock.calls.map((c) => c.join(' ')).join('\n');
expect(errText).toContain('FORGE_NO_EXECUTOR');
});
});
+122 -48
View File
@@ -5,37 +5,47 @@ import type { Command } from 'commander';
import { classifyBrief } from './brief-classifier.js';
import { STAGE_LABELS, STAGE_SEQUENCE } from './constants.js';
import { ForgeCapabilityError } from './errors.js';
import { getEffectivePersonas, loadBoardPersonas } from './persona-loader.js';
import { generateRunId, getPipelineStatus, loadManifest, runPipeline } from './pipeline-runner.js';
import type { PipelineOptions, RunManifest, TaskExecutor } from './types.js';
// ---------------------------------------------------------------------------
// Stub executor — used when no real executor is wired at CLI invocation time.
// ---------------------------------------------------------------------------
const stubExecutor: TaskExecutor = {
async submitTask(task) {
console.log(` [forge] stage submitted: ${task.id} (${task.title})`);
},
async waitForCompletion(taskId, _timeoutMs) {
console.log(` [forge] stage complete: ${taskId}`);
return {
task_id: taskId,
status: 'completed' as const,
completed_at: new Date().toISOString(),
exit_code: 0,
gate_results: [],
};
},
async getTaskStatus(_taskId) {
return 'completed' as const;
},
};
import { createSimulatedExecutor } from './simulated-executor.js';
import type { PipelineOptions, RunManifest, RunMode } from './types.js';
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
/** Resolve a run's effective mode, defaulting legacy manifests to normal. */
function runModeOf(manifest: RunManifest): RunMode {
return manifest.mode ?? 'normal';
}
/** Print a loud banner so a simulated run can never be misread as verified. */
function printSimulatedBanner(): void {
console.log('');
console.log('[forge] ===============================================================');
console.log('[forge] MODE: SIMULATED — no stage or gate was really executed.');
console.log('[forge] All results are synthetic and MUST NOT be read as verified');
console.log('[forge] success. Wire a real executor/providers and re-run to verify.');
console.log('[forge] ===============================================================');
}
/** Print a typed error line for fail-closed capability errors. */
function printCapabilityError(err: ForgeCapabilityError): void {
console.error(`[forge] error ${err.code}: ${err.message}`);
console.error(`[forge] missing capability: ${err.capability}`);
}
/** Handle a pipeline error uniformly: typed capability errors get their code. */
function handlePipelineError(err: unknown): void {
if (err instanceof ForgeCapabilityError) {
printCapabilityError(err);
} else {
console.error(`[forge] pipeline failed: ${err instanceof Error ? err.message : String(err)}`);
}
process.exitCode = 1;
}
function formatDuration(startedAt?: string, completedAt?: string): string {
if (!startedAt || !completedAt) return '-';
const ms = new Date(completedAt).getTime() - new Date(startedAt).getTime();
@@ -44,19 +54,24 @@ function formatDuration(startedAt?: string, completedAt?: string): string {
}
function printManifestTable(manifest: RunManifest): void {
const mode = runModeOf(manifest);
console.log(`\nRun ID : ${manifest.runId}`);
console.log(`Status : ${manifest.status}`);
console.log(`Mode : ${mode}`);
if (mode === 'simulated') {
console.log('WARNING: SIMULATED RUN — results are synthetic, not verified success.');
}
console.log(`Brief : ${manifest.brief}`);
console.log(`Class : ${manifest.briefClass} (${manifest.classSource})`);
console.log(`Updated: ${manifest.updatedAt}`);
console.log('');
console.log('Stage'.padEnd(22) + 'Status'.padEnd(14) + 'Duration');
console.log('-'.repeat(50));
console.log('Stage'.padEnd(22) + 'Status'.padEnd(24) + 'Duration');
console.log('-'.repeat(60));
for (const stage of STAGE_SEQUENCE) {
const s = manifest.stages[stage];
if (!s) continue;
const label = (STAGE_LABELS[stage] ?? stage).padEnd(22);
const status = s.status.padEnd(14);
const status = s.status.padEnd(24);
const dur = formatDuration(s.startedAt, s.completedAt);
console.log(`${label}${status}${dur}`);
}
@@ -90,23 +105,58 @@ function listRecentRuns(projectRoot?: string): void {
}
console.log('\nRecent runs:');
console.log('Run ID'.padEnd(22) + 'Status'.padEnd(14) + 'Brief');
console.log('-'.repeat(70));
console.log('Run ID'.padEnd(22) + 'Status'.padEnd(24) + 'Mode'.padEnd(12) + 'Brief');
console.log('-'.repeat(80));
for (const runId of entries) {
const runDir = path.join(runsDir, runId);
try {
const manifest = loadManifest(runDir);
const status = manifest.status.padEnd(14);
const status = manifest.status.padEnd(24);
const mode = runModeOf(manifest).padEnd(12);
const brief = path.basename(manifest.brief);
console.log(`${runId.padEnd(22)}${status}${brief}`);
console.log(`${runId.padEnd(22)}${status}${mode}${brief}`);
} catch {
console.log(`${runId.padEnd(22)}${'(unreadable)'.padEnd(14)}`);
console.log(`${runId.padEnd(22)}${'(unreadable)'.padEnd(24)}`);
}
}
console.log('');
}
/**
* Apply the exit-code policy for a finished pipeline run (SDLC-D-035):
*
* - exit 0 only for a verified `completed` normal run, or for an overall
* `simulated` run when the caller explicitly passed --simulate;
* - anything else exits nonzero so it can never be read as success.
*/
function applyRunExitPolicy(result: { manifest: RunManifest; runDir: string }, simulate: boolean) {
const { manifest } = result;
if (runModeOf(manifest) === 'simulated') {
if (!simulate || manifest.status !== 'simulated') {
console.error(
'[forge] error FORGE_MODE_MISMATCH: run reports simulated results without an explicit, ' +
'consistent --simulate request; refusing to report success.',
);
process.exitCode = 1;
return;
}
printSimulatedBanner();
console.log(`[forge] run directory: ${result.runDir}`);
return; // exit 0 — the caller explicitly opted into simulation
}
if (manifest.status !== 'completed') {
console.error(`[forge] run did not complete: terminal status '${manifest.status}'`);
process.exitCode = 1;
return;
}
console.log(`[forge] pipeline complete (mode: normal): ${manifest.runId}`);
console.log(`[forge] run directory: ${result.runDir}`);
}
// ---------------------------------------------------------------------------
// Register function
// ---------------------------------------------------------------------------
@@ -129,6 +179,11 @@ export function registerForgeCommand(parent: Command): void {
.option('--config <path>', 'Path to forge config file (.forge/config.yaml)')
.option('--codebase <path>', 'Codebase root to pass to the pipeline', process.cwd())
.option('--dry-run', 'Print planned stages without executing', false)
.option(
'--simulate',
'Simulate execution without real providers (every result is typed simulated, never verified)',
false,
)
.action(
async (opts: {
brief: string;
@@ -137,6 +192,7 @@ export function registerForgeCommand(parent: Command): void {
config?: string;
codebase: string;
dryRun: boolean;
simulate: boolean;
}) => {
const briefPath = path.resolve(opts.brief);
@@ -149,14 +205,22 @@ export function registerForgeCommand(parent: Command): void {
const briefContent = fs.readFileSync(briefPath, 'utf-8');
const briefClass = classifyBrief(briefContent);
const projectRoot = opts.codebase;
// A real executor is never wired at CLI invocation time today, so the
// only executor we may construct is the explicitly-requested simulated
// one. Normal mode fails closed with FORGE_NO_EXECUTOR.
const executor = opts.simulate ? createSimulatedExecutor() : undefined;
if (opts.resume) {
const runId = opts.runId ?? generateRunId();
const runDir = resolveRunDir(runId, projectRoot);
console.log(`[forge] resuming run: ${runId}`);
const { resumePipeline } = await import('./pipeline-runner.js');
const result = await resumePipeline(runDir, stubExecutor);
console.log(`[forge] pipeline complete: ${result.runId}`);
try {
const { resumePipeline } = await import('./pipeline-runner.js');
const result = await resumePipeline(runDir, executor, { simulate: opts.simulate });
applyRunExitPolicy(result, opts.simulate);
} catch (err) {
handlePipelineError(err);
}
return;
}
@@ -164,7 +228,8 @@ export function registerForgeCommand(parent: Command): void {
briefClass,
codebase: projectRoot,
dryRun: opts.dryRun,
executor: stubExecutor,
executor,
simulate: opts.simulate,
};
if (opts.dryRun) {
@@ -180,16 +245,15 @@ export function registerForgeCommand(parent: Command): void {
console.log(`[forge] starting pipeline for brief: ${briefPath}`);
console.log(`[forge] classified as: ${briefClass}`);
if (opts.simulate) {
console.log('[forge] mode: SIMULATED (explicit --simulate)');
}
try {
const result = await runPipeline(briefPath, projectRoot, pipelineOptions);
console.log(`[forge] pipeline complete: ${result.runId}`);
console.log(`[forge] run directory: ${result.runDir}`);
applyRunExitPolicy(result, opts.simulate);
} catch (err) {
console.error(
`[forge] pipeline failed: ${err instanceof Error ? err.message : String(err)}`,
);
process.exitCode = 1;
handlePipelineError(err);
}
},
);
@@ -224,7 +288,12 @@ export function registerForgeCommand(parent: Command): void {
.command('resume <runId>')
.description('Resume a stopped or failed pipeline run')
.option('--project <path>', 'Project root (defaults to cwd)', process.cwd())
.action(async (runId: string, opts: { project: string }) => {
.option(
'--simulate',
'Simulate execution without real providers (every result is typed simulated, never verified)',
false,
)
.action(async (runId: string, opts: { project: string; simulate: boolean }) => {
const runDir = resolveRunDir(runId, opts.project);
if (!fs.existsSync(runDir)) {
@@ -234,15 +303,20 @@ export function registerForgeCommand(parent: Command): void {
}
console.log(`[forge] resuming run: ${runId}`);
if (opts.simulate) {
console.log('[forge] mode: SIMULATED (explicit --simulate)');
}
// No real executor is wired at CLI invocation time; only the explicitly
// requested simulated executor may be constructed (fail closed otherwise).
const executor = opts.simulate ? createSimulatedExecutor() : undefined;
try {
const { resumePipeline } = await import('./pipeline-runner.js');
const result = await resumePipeline(runDir, stubExecutor);
console.log(`[forge] pipeline complete: ${result.runId}`);
console.log(`[forge] run directory: ${result.runDir}`);
const result = await resumePipeline(runDir, executor, { simulate: opts.simulate });
applyRunExitPolicy(result, opts.simulate);
} catch (err) {
console.error(`[forge] resume failed: ${err instanceof Error ? err.message : String(err)}`);
process.exitCode = 1;
handlePipelineError(err);
}
});
+72 -12
View File
@@ -9,7 +9,16 @@ export const PACKAGE_ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.
/** Pipeline asset directory (stages, agents, rails, gates, templates). */
export const PIPELINE_DIR = path.join(PACKAGE_ROOT, 'pipeline');
/** Stage specifications — defines every pipeline stage. */
/** Stage specifications defines every pipeline stage.
*\n * Gate semantics (SDLC-D-035): every gate is one of
* - a real command string / GateEntry a mechanical runner can execute,
* - an `authority` gate (human/board sign-off; produces waiting-for-authority),
* - a `provider` gate (requires a wired provider such as a reviewer or CI pipeline).
*
* Vacuous gates (`true`, echo'd synthetic approvals, placeholder ci-pipeline
* commands) are forbidden: a stage whose gate has no real implementation
* fails closed instead of passing.
*/
export const STAGE_SPECS: Record<string, StageSpec> = {
'00-intake': {
number: '00',
@@ -27,7 +36,13 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
type: 'research',
gate: 'discovery-complete',
promptFile: '00b-discovery.md',
qualityGates: ['true'],
qualityGates: [
{
kind: 'authority',
capability: 'discovery-complete',
reason: 'discovery completion is attested by an authority; no mechanical check exists',
},
],
},
'01-board': {
number: '01',
@@ -36,7 +51,13 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
type: 'review',
gate: 'board-approval',
promptFile: '01-board.md',
qualityGates: [{ type: 'ci-pipeline', command: 'board-approval (via board-tasks)' }],
qualityGates: [
{
kind: 'authority',
capability: 'board-approval',
reason: 'board approval is a board/human decision; no mechanical gate exists',
},
],
},
'01b-brief-analyzer': {
number: '01b',
@@ -45,7 +66,13 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
type: 'research',
gate: 'brief-analysis-complete',
promptFile: '01-board.md',
qualityGates: ['true'],
qualityGates: [
{
kind: 'authority',
capability: 'brief-analysis-complete',
reason: 'brief analysis completion is attested by an authority; no mechanical check exists',
},
],
},
'02-planning-1': {
number: '02',
@@ -54,7 +81,13 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
type: 'research',
gate: 'architecture-approval',
promptFile: '02-planning-1-architecture.md',
qualityGates: ['true'],
qualityGates: [
{
kind: 'authority',
capability: 'architecture-approval',
reason: 'ADR approval requires authority sign-off; no mechanical check exists',
},
],
},
'03-planning-2': {
number: '03',
@@ -63,7 +96,14 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
type: 'research',
gate: 'implementation-approval',
promptFile: '03-planning-2-implementation.md',
qualityGates: ['true'],
qualityGates: [
{
kind: 'authority',
capability: 'implementation-approval',
reason:
'implementation spec approval requires authority sign-off; no mechanical check exists',
},
],
},
'04-planning-3': {
number: '04',
@@ -72,7 +112,14 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
type: 'research',
gate: 'decomposition-approval',
promptFile: '04-planning-3-decomposition.md',
qualityGates: ['true'],
qualityGates: [
{
kind: 'authority',
capability: 'decomposition-approval',
reason:
'task decomposition approval requires authority sign-off; no mechanical check exists',
},
],
},
'05-coding': {
number: '05',
@@ -92,9 +139,10 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
promptFile: '06-review.md',
qualityGates: [
{
type: 'ai-review',
command:
'echo \'{"summary":"review-pass","verdict":"approve","findings":[],"stats":{"blockers":0,"should_fix":0,"suggestions":0}}\'',
kind: 'provider',
capability: 'reviewer',
reason:
'review verdicts require a wired reviewer provider; synthetic approvals are not permitted',
},
],
},
@@ -105,7 +153,13 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
type: 'coding',
gate: 're-review',
promptFile: '07-remediate.md',
qualityGates: ['true'],
qualityGates: [
{
kind: 'authority',
capability: 're-review',
reason: 'remediation re-review is an approval-based gate; no mechanical check exists',
},
],
},
'08-test': {
number: '08',
@@ -123,7 +177,13 @@ export const STAGE_SPECS: Record<string, StageSpec> = {
type: 'deploy',
gate: 'deploy-verification',
promptFile: '09-deploy.md',
qualityGates: [{ type: 'ci-pipeline', command: 'deploy-verification' }],
qualityGates: [
{
kind: 'provider',
capability: 'ci-pipeline',
reason: 'deploy verification requires a wired CI pipeline provider',
},
],
},
};
+46
View File
@@ -0,0 +1,46 @@
/**
* Typed fail-closed capability errors (SDLC-D-035).
*
* A Forge run must fail closed when a required capability (executor, reviewer
* provider, CI pipeline, authority sign-off) is missing. These typed errors
* name the missing capability so callers can distinguish "not wired" from
* ordinary execution failures.
*/
/** Closed set of typed Forge capability error codes. */
export const FORGE_ERROR_CODES = [
'FORGE_NO_EXECUTOR',
'FORGE_NO_REVIEWER',
'FORGE_NO_CI_PIPELINE',
'FORGE_NO_PROVIDER',
'FORGE_AUTHORITY_REQUIRED',
] as const;
export type ForgeErrorCode = (typeof FORGE_ERROR_CODES)[number];
/** Raised when a required capability is missing and the pipeline must fail closed. */
export class ForgeCapabilityError extends Error {
/** Typed error code from the closed FORGE_ERROR_CODES set. */
readonly code: ForgeErrorCode;
/** The missing capability, e.g. `task-executor`, `reviewer`, `board-approval`. */
readonly capability: string;
constructor(code: ForgeErrorCode, capability: string, message: string) {
super(message);
this.name = 'ForgeCapabilityError';
this.code = code;
this.capability = capability;
}
}
/** Map a provider gate capability to its typed error code. */
export function providerErrorCode(capability: string): ForgeErrorCode {
switch (capability) {
case 'reviewer':
return 'FORGE_NO_REVIEWER';
case 'ci-pipeline':
return 'FORGE_NO_CI_PIPELINE';
default:
return 'FORGE_NO_PROVIDER';
}
}
+26
View File
@@ -5,6 +5,13 @@ export type {
StageSpec,
BriefClass,
ClassSource,
ForgeOutcome,
AuthorityGate,
ProviderGate,
ForgeGate,
ForgeGateResult,
ForgeTaskResult,
RunMode,
StageStatus,
RunManifest,
ForgeTaskStatus,
@@ -81,5 +88,24 @@ export {
getPipelineStatus,
} from './pipeline-runner.js';
// Fail-closed errors and typed outcome model (SDLC-D-035)
export { FORGE_ERROR_CODES, ForgeCapabilityError, providerErrorCode } from './errors.js';
export type { ForgeErrorCode } from './errors.js';
export {
isSatisfyingOutcome,
isCapabilityGate,
isCommandGate,
gateLabel,
uniformGateResults,
simulatedGateResults,
waitingGateResults,
blockedGateResults,
evaluateStageGates,
} from './outcomes.js';
export type { StageEvaluation } from './outcomes.js';
// Simulated executor (explicit --simulate only)
export { createSimulatedExecutor } from './simulated-executor.js';
// CLI
export { registerForgeCommand } from './cli.js';
+147
View File
@@ -0,0 +1,147 @@
import type { GateEntry } from '@mosaicstack/macp';
import type {
AuthorityGate,
ForgeGate,
ForgeGateResult,
ForgeOutcome,
ForgeTaskResult,
ProviderGate,
} from './types.js';
/**
* Gate and dependency satisfaction predicate (SDLC-D-035).
*
* ONLY a verified `passed` outcome satisfies. Every other member of the closed
* outcome set including `simulated` is non-satisfying, so a simulated or
* authority-blocked result can never be read as success-by-verification.
*/
export function isSatisfyingOutcome(outcome: ForgeOutcome): boolean {
return outcome === 'passed';
}
/** Whether a gate is an authority or provider gate (capability-based, command-less). */
export function isCapabilityGate(gate: ForgeGate): gate is AuthorityGate | ProviderGate {
if (typeof gate !== 'object' || gate === null) return false;
const kind = (gate as Record<string, unknown>)['kind'];
return kind === 'authority' || kind === 'provider';
}
/** Whether a gate definition carries a real command a mechanical runner can execute. */
export function isCommandGate(gate: ForgeGate): gate is string | GateEntry {
if (typeof gate === 'string') {
return gate.trim().length > 0;
}
if (isCapabilityGate(gate)) {
// Authority and provider gates are satisfied by a capability, not a command.
return false;
}
return typeof gate.command === 'string' && gate.command.trim().length > 0;
}
/** Typed label identifying a gate in results and logs. */
export function gateLabel(gate: ForgeGate): string {
if (typeof gate === 'string') return gate;
if (isCapabilityGate(gate)) return `${gate.kind}:${gate.capability}`;
return gate.command || gate.type || 'unnamed-gate';
}
/** Reason string stamped on every simulated gate result. */
export const SIMULATED_GATE_REASON =
'simulated execution (--simulate): gate was not evaluated by a real implementation';
/** Build typed gate results with a uniform outcome for a stage's declared gates. */
export function uniformGateResults(
gates: ForgeGate[],
outcome: ForgeOutcome,
reason: string,
): ForgeGateResult[] {
return gates.map((gate) => ({ gate: gateLabel(gate), outcome, reason }));
}
/** Typed simulated gate results — used exclusively in `--simulate` runs. */
export function simulatedGateResults(gates: ForgeGate[]): ForgeGateResult[] {
return uniformGateResults(gates, 'simulated', SIMULATED_GATE_REASON);
}
/** Typed waiting-for-authority gate results for approval-based stages. */
export function waitingGateResults(gates: ForgeGate[], reason: string): ForgeGateResult[] {
return uniformGateResults(gates, 'waiting-for-authority', reason);
}
/** Typed blocked gate results for stages whose provider capability is not wired. */
export function blockedGateResults(gates: ForgeGate[], reason: string): ForgeGateResult[] {
return uniformGateResults(gates, 'blocked', reason);
}
/** Outcome of evaluating a completed stage in normal mode. */
export interface StageEvaluation {
outcome: ForgeOutcome;
reason: string;
gateResults: ForgeGateResult[];
}
/**
* Evaluate a stage's declared gates against the executor's typed result.
*
* Fail-closed mapping:
* - a `simulated` task or gate outcome in normal mode maps to `error`
* - a missing gate result for a required command gate maps to `blocked`
* - a non-passing task outcome propagates as the stage outcome
* - only verified `passed` task and gate outcomes yield a `passed` stage
*/
export function evaluateStageGates(
stageName: string,
gates: ForgeGate[],
result: ForgeTaskResult,
): StageEvaluation {
const gateResults = result.gate_results ?? [];
if (result.outcome === 'simulated') {
return {
outcome: 'error',
reason: `executor reported a simulated outcome for stage '${stageName}' in normal mode — refusing to treat simulated results as verified`,
gateResults,
};
}
if (!isSatisfyingOutcome(result.outcome)) {
return {
outcome: result.outcome,
reason: `task outcome is '${result.outcome}': ${result.reason}`,
gateResults,
};
}
for (const gate of gates) {
// Authority and provider gates are pre-flighted before execution; they have
// no mechanical result to verify here.
if (!isCommandGate(gate)) continue;
const label = gateLabel(gate);
const gateResult = gateResults.find((r) => r.gate === label);
if (!gateResult) {
return {
outcome: 'blocked',
reason: `no gate result was reported for required gate '${label}' (stage '${stageName}')`,
gateResults,
};
}
if (!isSatisfyingOutcome(gateResult.outcome)) {
return {
outcome: gateResult.outcome === 'simulated' ? 'error' : gateResult.outcome,
reason: `gate '${label}' outcome is '${gateResult.outcome}': ${gateResult.reason}`,
gateResults,
};
}
}
return {
outcome: 'passed',
reason:
gates.length === 0
? "stage declares no gates; task outcome 'passed' accepted"
: 'all declared gates verified passed',
gateResults,
};
}
+227 -99
View File
@@ -1,18 +1,33 @@
import fs from 'node:fs';
import path from 'node:path';
import { STAGE_SEQUENCE } from './constants.js';
import { STAGE_SEQUENCE, STAGE_SPECS } from './constants.js';
import { determineBriefClass, stagesForClass } from './brief-classifier.js';
import { ForgeCapabilityError, providerErrorCode } from './errors.js';
import {
blockedGateResults,
evaluateStageGates,
isCapabilityGate,
simulatedGateResults,
waitingGateResults,
} from './outcomes.js';
import { mapStageToTask } from './stage-adapter.js';
import { createSimulatedExecutor } from './simulated-executor.js';
import type {
ForgeTask,
ForgeTaskResult,
PipelineOptions,
PipelineResult,
RunManifest,
RunMode,
StageStatus,
TaskExecutor,
} from './types.js';
/** Reason stamped on stages that complete under explicit simulation. */
const SIMULATED_STAGE_REASON =
'simulated execution (--simulate): stage was not executed by a real executor';
/**
* Generate a timestamp-based run ID.
*/
@@ -47,6 +62,7 @@ function createManifest(opts: {
briefClass: RunManifest['briefClass'];
classSource: RunManifest['classSource'];
forceBoard: boolean;
mode: RunMode;
runDir: string;
}): RunManifest {
const ts = nowISO();
@@ -57,6 +73,7 @@ function createManifest(opts: {
briefClass: opts.briefClass,
classSource: opts.classSource,
forceBoard: opts.forceBoard,
mode: opts.mode,
createdAt: ts,
updatedAt: ts,
currentStage: '',
@@ -108,20 +125,199 @@ export function selectStages(stages?: string[], skipTo?: string): string[] {
return selected.slice(skipIndex);
}
/**
* Fail closed when the required executor capability is missing (SDLC-D-035).
*/
function requireExecutor(executor: TaskExecutor | undefined, simulate: boolean): TaskExecutor {
if (executor) return executor;
if (simulate) return createSimulatedExecutor({ log: false });
throw new ForgeCapabilityError(
'FORGE_NO_EXECUTOR',
'task-executor',
'no task executor is wired; refusing to run the pipeline with a stub executor (fail closed). ' +
'Pass --simulate to opt into explicitly simulated execution.',
);
}
/**
* Pre-flight a stage's gates in normal mode (fail closed, SDLC-D-035).
*
* - authority gates: record a typed `waiting-for-authority` stage result and
* raise FORGE_AUTHORITY_REQUIRED approval-based gates never pass vacuously.
* - provider gates: record a typed `blocked` stage result and raise the typed
* capability error for the missing provider.
*
* Returns the stage status to record when the pre-flight blocks, or undefined
* when the stage may proceed.
*/
function preflightStageGates(
stageName: string,
manifest: RunManifest,
): { status: StageStatus; error: ForgeCapabilityError } | undefined {
const spec = STAGE_SPECS[stageName];
if (!spec) throw new Error(`Unknown Forge stage: ${stageName}`);
for (const gate of spec.qualityGates) {
if (!isCapabilityGate(gate)) continue;
const startedAt = manifest.stages[stageName]?.startedAt;
const completedAt = nowISO();
if (gate.kind === 'authority') {
const reason = `gate '${gate.capability}' requires authority sign-off; no mechanical implementation exists (${gate.reason})`;
return {
status: {
status: 'waiting-for-authority',
reason,
startedAt,
completedAt,
gateResults: waitingGateResults(spec.qualityGates, reason),
},
error: new ForgeCapabilityError(
'FORGE_AUTHORITY_REQUIRED',
gate.capability,
`stage '${stageName}' is blocked on authority gate '${gate.capability}': ${gate.reason}. ` +
'The pipeline fails closed instead of passing vacuously. Record the approval out-of-band ' +
'or run with --simulate for explicitly simulated execution.',
),
};
}
const reason = `gate '${gate.capability}' requires provider '${gate.capability}' and none is wired (${gate.reason})`;
return {
status: {
status: 'blocked',
reason,
startedAt,
completedAt,
gateResults: blockedGateResults(spec.qualityGates, reason),
},
error: new ForgeCapabilityError(
providerErrorCode(gate.capability),
gate.capability,
`stage '${stageName}' requires provider '${gate.capability}' which is not wired: ${gate.reason}. ` +
'The pipeline fails closed instead of passing vacuously.',
),
};
}
return undefined;
}
/**
* Execute the given stage tasks sequentially, updating the manifest.
*
* Normal mode requires a real executor and evaluates every declared command
* gate through the typed outcome model; any non-verified result fails closed.
* Simulate mode types every stage and gate result as `simulated`.
*/
async function executeStages(opts: {
manifest: RunManifest;
runDir: string;
tasks: ForgeTask[];
stageNames: string[];
executor: TaskExecutor;
simulate: boolean;
}): Promise<void> {
const { manifest, runDir, tasks, stageNames, executor, simulate } = opts;
for (let i = 0; i < tasks.length; i++) {
const task = tasks[i]!;
const stageName = stageNames[i]!;
const spec = STAGE_SPECS[stageName];
if (!spec) throw new Error(`Unknown Forge stage: ${stageName}`);
// Update manifest: stage in progress
manifest.currentStage = stageName;
manifest.stages[stageName] = {
status: 'in_progress',
startedAt: nowISO(),
};
saveManifest(runDir, manifest);
// Fail-closed pre-flight (normal mode only): authority/provider gates have
// no mechanical implementation and must never pass vacuously.
if (!simulate) {
const blocked = preflightStageGates(stageName, manifest);
if (blocked) {
manifest.stages[stageName] = blocked.status;
manifest.status =
blocked.status.status === 'waiting-for-authority' ? 'waiting-for-authority' : 'failed';
saveManifest(runDir, manifest);
throw blocked.error;
}
}
let result: ForgeTaskResult;
try {
await executor.submitTask(task);
result = await executor.waitForCompletion(task.id, task.timeoutSeconds * 1000);
} catch (error) {
// Process errors (including timeouts) map to the fail-closed `error` outcome.
const reason = error instanceof Error ? error.message : String(error);
manifest.stages[stageName] = {
status: 'error',
reason: `executor error: ${reason}`,
startedAt: manifest.stages[stageName]?.startedAt,
completedAt: nowISO(),
gateResults: [],
};
manifest.status = 'failed';
saveManifest(runDir, manifest);
throw error instanceof Error ? error : new Error(reason);
}
if (simulate) {
manifest.stages[stageName] = {
status: 'simulated',
reason: SIMULATED_STAGE_REASON,
startedAt: manifest.stages[stageName]?.startedAt,
completedAt: nowISO(),
gateResults: simulatedGateResults(spec.qualityGates),
};
saveManifest(runDir, manifest);
continue;
}
const evaluation = evaluateStageGates(stageName, spec.qualityGates, result);
manifest.stages[stageName] = {
status: evaluation.outcome,
reason: evaluation.reason,
startedAt: manifest.stages[stageName]?.startedAt,
completedAt: nowISO(),
gateResults: evaluation.gateResults,
};
if (evaluation.outcome !== 'passed') {
manifest.status =
evaluation.outcome === 'waiting-for-authority' ? 'waiting-for-authority' : 'failed';
saveManifest(runDir, manifest);
throw new Error(`Stage ${stageName} ${evaluation.outcome}: ${evaluation.reason}`);
}
saveManifest(runDir, manifest);
}
}
/**
* Run the Forge pipeline.
*
* 1. Classify the brief
* 2. Generate a run ID and create run directory
* 3. Map stages to tasks and submit to TaskExecutor
* 4. Track manifest with stage statuses
* 5. Return pipeline result
* 1. Fail closed unless a real executor is wired or simulation is explicit
* 2. Classify the brief
* 3. Generate a run ID and create run directory
* 4. Map stages to tasks and submit to TaskExecutor
* 5. Track manifest with typed stage outcomes
* 6. Return pipeline result
*/
export async function runPipeline(
briefPath: string,
projectRoot: string,
options: PipelineOptions,
): Promise<PipelineResult> {
const simulate = options.simulate ?? false;
const executor = requireExecutor(options.executor, simulate);
const mode: RunMode = simulate ? 'simulated' : 'normal';
const resolvedRoot = path.resolve(projectRoot);
const resolvedBrief = path.resolve(briefPath);
const briefContent = fs.readFileSync(resolvedBrief, 'utf-8');
@@ -146,6 +342,7 @@ export async function runPipeline(
briefClass,
classSource,
forceBoard: options.forceBoard ?? false,
mode,
runDir,
});
@@ -172,54 +369,10 @@ export async function runPipeline(
}
// Execute stages
const { executor } = options;
for (let i = 0; i < tasks.length; i++) {
const task = tasks[i]!;
const stageName = selectedStages[i]!;
await executeStages({ manifest, runDir, tasks, stageNames: selectedStages, executor, simulate });
// Update manifest: stage in progress
manifest.currentStage = stageName;
manifest.stages[stageName] = {
status: 'in_progress',
startedAt: nowISO(),
};
saveManifest(runDir, manifest);
try {
await executor.submitTask(task);
const result = await executor.waitForCompletion(task.id, task.timeoutSeconds * 1000);
// Update manifest: stage completed or failed
const stageStatus: StageStatus = {
status: result.status === 'completed' ? 'passed' : 'failed',
startedAt: manifest.stages[stageName]!.startedAt,
completedAt: nowISO(),
};
manifest.stages[stageName] = stageStatus;
if (result.status !== 'completed') {
manifest.status = 'failed';
saveManifest(runDir, manifest);
throw new Error(`Stage ${stageName} failed with status: ${result.status}`);
}
saveManifest(runDir, manifest);
} catch (error) {
if (!manifest.stages[stageName]?.completedAt) {
manifest.stages[stageName] = {
status: 'failed',
startedAt: manifest.stages[stageName]?.startedAt,
completedAt: nowISO(),
};
}
manifest.status = 'failed';
saveManifest(runDir, manifest);
throw error;
}
}
// All stages passed
manifest.status = 'completed';
// All stages reached a terminal state for this mode
manifest.status = simulate ? 'simulated' : 'completed';
saveManifest(runDir, manifest);
return {
@@ -234,22 +387,30 @@ export async function runPipeline(
}
/**
* Resume a pipeline from the last incomplete stage.
* Resume a pipeline from the last non-passed stage.
*/
export async function resumePipeline(
runDir: string,
executor: TaskExecutor,
executor?: TaskExecutor,
options?: { simulate?: boolean },
): Promise<PipelineResult> {
const simulate = options?.simulate ?? false;
const wiredExecutor = requireExecutor(executor, simulate);
const mode: RunMode = simulate ? 'simulated' : 'normal';
const manifest = loadManifest(runDir);
const resolvedRoot = path.dirname(path.dirname(path.dirname(runDir))); // .forge/runs/{id} → project root
const briefContent = fs.readFileSync(manifest.brief, 'utf-8');
const allStages = stagesForClass(manifest.briefClass, manifest.forceBoard);
// Find first non-passed stage
manifest.mode = mode;
// Find first non-satisfying stage (only a verified `passed` counts as done;
// simulated and waiting-for-authority stages are re-run).
const resumeFrom = allStages.find((s) => manifest.stages[s]?.status !== 'passed');
if (!resumeFrom) {
manifest.status = 'completed';
manifest.status = mode === 'simulated' ? 'simulated' : 'completed';
saveManifest(runDir, manifest);
return {
runId: manifest.runId,
@@ -284,49 +445,16 @@ export async function resumePipeline(
tasks.push(task);
}
for (let i = 0; i < tasks.length; i++) {
const task = tasks[i]!;
const stageName = remainingStages[i]!;
await executeStages({
manifest,
runDir,
tasks,
stageNames: remainingStages,
executor: wiredExecutor,
simulate,
});
manifest.currentStage = stageName;
manifest.stages[stageName] = {
status: 'in_progress',
startedAt: nowISO(),
};
saveManifest(runDir, manifest);
try {
await executor.submitTask(task);
const result = await executor.waitForCompletion(task.id, task.timeoutSeconds * 1000);
manifest.stages[stageName] = {
status: result.status === 'completed' ? 'passed' : 'failed',
startedAt: manifest.stages[stageName]!.startedAt,
completedAt: nowISO(),
};
if (result.status !== 'completed') {
manifest.status = 'failed';
saveManifest(runDir, manifest);
throw new Error(`Stage ${stageName} failed with status: ${result.status}`);
}
saveManifest(runDir, manifest);
} catch (error) {
if (!manifest.stages[stageName]?.completedAt) {
manifest.stages[stageName] = {
status: 'failed',
startedAt: manifest.stages[stageName]?.startedAt,
completedAt: nowISO(),
};
}
manifest.status = 'failed';
saveManifest(runDir, manifest);
throw error;
}
}
manifest.status = 'completed';
manifest.status = simulate ? 'simulated' : 'completed';
saveManifest(runDir, manifest);
return {
+32
View File
@@ -0,0 +1,32 @@
import type { ForgeTask, ForgeTaskResult, TaskExecutor } from './types.js';
/**
* Simulated executor used ONLY when the caller explicitly passes --simulate.
*
* It submits no real work and returns typed `simulated` results so a simulated
* run can never be confused with a verified one. In normal mode (no --simulate)
* the CLI refuses to run at all with FORGE_NO_EXECUTOR instead of wiring this
* stub (fail closed, SDLC-D-035).
*/
export function createSimulatedExecutor(options?: { log?: boolean }): TaskExecutor {
const log = options?.log ?? true;
return {
async submitTask(task: ForgeTask) {
if (log) console.log(` [forge:simulated] stage submitted: ${task.id} (${task.title})`);
},
async waitForCompletion(taskId: string): Promise<ForgeTaskResult> {
if (log) console.log(` [forge:simulated] stage complete: ${taskId}`);
return {
task_id: taskId,
outcome: 'simulated',
reason: 'no executor wired; simulated execution requested via --simulate',
completed_at: new Date().toISOString(),
exit_code: 0,
gate_results: [],
};
},
async getTaskStatus() {
return 'completed' as const;
},
};
}
+88 -7
View File
@@ -1,4 +1,4 @@
import type { GateEntry, TaskResult } from '@mosaicstack/macp';
import type { GateEntry } from '@mosaicstack/macp';
/** Stage dispatch mode. */
export type StageDispatch = 'exec' | 'yolo' | 'pi';
@@ -6,6 +6,58 @@ export type StageDispatch = 'exec' | 'yolo' | 'pi';
/** Stage type — determines agent selection and gate requirements. */
export type StageType = 'research' | 'review' | 'coding' | 'deploy';
/**
* Typed outcome for every gate and stage evaluation closed set (SDLC-D-035).
*
* Only `passed` means "verified by a real implementation". `simulated` is
* produced exclusively in explicit `--simulate` runs and is never satisfying.
*/
export type ForgeOutcome =
| 'passed'
| 'failed'
| 'blocked'
| 'error'
| 'waiting-for-authority'
| 'simulated'
| 'not-applicable';
/** A gate that requires authority (human/board) sign-off; no mechanical command can satisfy it. */
export interface AuthorityGate {
kind: 'authority';
capability: string;
reason: string;
}
/** A gate that requires a wired provider (e.g. an AI reviewer, CI pipeline) to evaluate. */
export interface ProviderGate {
kind: 'provider';
capability: string;
reason: string;
}
/** Forge quality gate: a real command, an authority sign-off, or a provider-backed check. */
export type ForgeGate = string | GateEntry | AuthorityGate | ProviderGate;
/** Typed result of evaluating a single quality gate. */
export interface ForgeGateResult {
gate: string;
outcome: ForgeOutcome;
reason: string;
exitCode?: number;
output?: string;
timedOut?: boolean;
}
/** Typed result of a task/stage execution returned by a TaskExecutor. */
export interface ForgeTaskResult {
task_id: string;
outcome: ForgeOutcome;
reason: string;
completed_at: string;
exit_code: number;
gate_results: ForgeGateResult[];
}
/** Stage specification — defines a single pipeline stage. */
export interface StageSpec {
number: string;
@@ -14,7 +66,7 @@ export interface StageSpec {
type: StageType;
gate: string;
promptFile: string;
qualityGates: (string | GateEntry)[];
qualityGates: ForgeGate[];
}
/** Brief classification. */
@@ -25,11 +77,18 @@ export type ClassSource = 'cli' | 'frontmatter' | 'auto';
/** Per-stage status within a run manifest. */
export interface StageStatus {
status: 'pending' | 'in_progress' | 'passed' | 'failed';
status: 'pending' | 'in_progress' | ForgeOutcome;
/** Why the stage reached its current (terminal) outcome, when applicable. */
reason?: string;
startedAt?: string;
completedAt?: string;
/** Typed per-gate results recorded alongside the stage outcome. */
gateResults?: ForgeGateResult[];
}
/** Execution mode of a run. */
export type RunMode = 'normal' | 'simulated';
/** Run manifest — persisted to disk as manifest.json. */
export interface RunManifest {
runId: string;
@@ -38,10 +97,23 @@ export interface RunManifest {
briefClass: BriefClass;
classSource: ClassSource;
forceBoard: boolean;
/**
* Execution mode. `simulated` runs stub execution; their results are typed
* `simulated` and must never be read as verified success. Optional because
* manifests written before this field existed default to `normal`.
*/
mode?: RunMode;
createdAt: string;
updatedAt: string;
currentStage: string;
status: 'in_progress' | 'completed' | 'failed' | 'interrupted' | 'rejected';
status:
| 'in_progress'
| 'completed'
| 'failed'
| 'interrupted'
| 'rejected'
| 'simulated'
| 'waiting-for-authority';
stages: Record<string, StageStatus>;
}
@@ -65,7 +137,7 @@ export interface ForgeTask {
briefPath: string;
resultPath: string;
timeoutSeconds: number;
qualityGates: (string | GateEntry)[];
qualityGates: ForgeGate[];
worktree?: string;
command?: string;
dependsOn?: string[];
@@ -76,7 +148,7 @@ export interface ForgeTask {
/** Abstract task executor — decouples from packages/coord. */
export interface TaskExecutor {
submitTask(task: ForgeTask): Promise<void>;
waitForCompletion(taskId: string, timeoutMs: number): Promise<TaskResult>;
waitForCompletion(taskId: string, timeoutMs: number): Promise<ForgeTaskResult>;
getTaskStatus(taskId: string): Promise<ForgeTaskStatus>;
}
@@ -122,7 +194,16 @@ export interface PipelineOptions {
stages?: string[];
skipTo?: string;
dryRun?: boolean;
executor: TaskExecutor;
/**
* Real task executor. Required in normal mode: the pipeline fails closed
* with FORGE_NO_EXECUTOR when it is absent.
*/
executor?: TaskExecutor;
/**
* Explicit opt-in to simulated execution. Every stage and gate result is
* typed `simulated` and is never satisfying.
*/
simulate?: boolean;
}
/** Pipeline run result. */
-253
View File
@@ -1,253 +0,0 @@
import { mkdirSync, readFileSync, rmSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { randomUUID } from 'node:crypto';
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { normalizeGate, countAIFindings, runGate, runGates } from '../src/gate-runner.js';
function makeTmpDir(): string {
const dir = join(tmpdir(), `macp-gate-${randomUUID()}`);
mkdirSync(dir, { recursive: true });
return dir;
}
describe('normalizeGate', () => {
it('normalizes a string to mechanical gate', () => {
expect(normalizeGate('echo test')).toEqual({
command: 'echo test',
type: 'mechanical',
fail_on: 'blocker',
});
});
it('normalizes an object gate with defaults', () => {
expect(normalizeGate({ command: 'lint' })).toEqual({
command: 'lint',
type: 'mechanical',
fail_on: 'blocker',
});
});
it('preserves explicit type and fail_on', () => {
expect(normalizeGate({ command: 'review', type: 'ai-review', fail_on: 'any' })).toEqual({
command: 'review',
type: 'ai-review',
fail_on: 'any',
});
});
it('handles non-string/non-object input', () => {
expect(normalizeGate(42)).toEqual({ command: '', type: 'mechanical', fail_on: 'blocker' });
expect(normalizeGate(null)).toEqual({ command: '', type: 'mechanical', fail_on: 'blocker' });
});
});
describe('countAIFindings', () => {
it('returns zeros for non-object', () => {
expect(countAIFindings(null)).toEqual({ blockers: 0, total: 0 });
expect(countAIFindings('string')).toEqual({ blockers: 0, total: 0 });
expect(countAIFindings([])).toEqual({ blockers: 0, total: 0 });
});
it('counts from stats block', () => {
const output = { stats: { blockers: 2, should_fix: 3, suggestions: 1 } };
expect(countAIFindings(output)).toEqual({ blockers: 2, total: 6 });
});
it('counts from findings array when stats has no blockers', () => {
const output = {
stats: { blockers: 0 },
findings: [{ severity: 'blocker' }, { severity: 'warning' }, { severity: 'blocker' }],
};
expect(countAIFindings(output)).toEqual({ blockers: 2, total: 3 });
});
it('uses stats blockers over findings array when stats has blockers', () => {
const output = {
stats: { blockers: 5 },
findings: [{ severity: 'blocker' }, { severity: 'warning' }],
};
// stats.blockers = 5, total from stats = 5+0+0 = 5, findings not used for total since stats total is non-zero
expect(countAIFindings(output)).toEqual({ blockers: 5, total: 5 });
});
it('counts findings length as total when stats has zero total', () => {
const output = {
findings: [{ severity: 'warning' }, { severity: 'info' }],
};
expect(countAIFindings(output)).toEqual({ blockers: 0, total: 2 });
});
});
describe('runGate', () => {
let tmp: string;
let logPath: string;
beforeEach(() => {
tmp = makeTmpDir();
logPath = join(tmp, 'gate.log');
});
afterEach(() => {
rmSync(tmp, { recursive: true, force: true });
});
it('passes mechanical gate on exit 0', () => {
const result = runGate('echo hello', tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.exit_code).toBe(0);
expect(result.type).toBe('mechanical');
expect(result.output).toContain('hello');
});
it('fails mechanical gate on non-zero exit', () => {
const result = runGate('exit 1', tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.exit_code).toBe(1);
});
it('ci-pipeline always passes', () => {
const result = runGate({ command: 'anything', type: 'ci-pipeline' }, tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.type).toBe('ci-pipeline');
expect(result.output).toBe('CI pipeline gate placeholder');
});
it('empty command passes', () => {
const result = runGate({ command: '' }, tmp, logPath, 30);
expect(result.passed).toBe(true);
});
it('ai-review gate parses JSON output', () => {
const json = JSON.stringify({ stats: { blockers: 0, should_fix: 1 } });
const result = runGate({ command: `echo '${json}'`, type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.blockers).toBe(0);
expect(result.findings).toBe(1);
});
it('ai-review gate fails on blockers', () => {
const json = JSON.stringify({ stats: { blockers: 2 } });
const result = runGate({ command: `echo '${json}'`, type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.blockers).toBe(2);
});
it('ai-review gate with fail_on=any fails on any findings', () => {
const json = JSON.stringify({ stats: { blockers: 0, should_fix: 1 } });
const result = runGate(
{ command: `echo '${json}'`, type: 'ai-review', fail_on: 'any' },
tmp,
logPath,
30,
);
expect(result.passed).toBe(false);
expect(result.fail_on).toBe('any');
});
it('ai-review gate fails on invalid JSON output', () => {
const result = runGate({ command: 'echo "not json"', type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.parse_error).toBeDefined();
});
it('writes to log file', () => {
runGate('echo logged', tmp, logPath, 30);
const log = readFileSync(logPath, 'utf-8');
expect(log).toContain('COMMAND: echo logged');
expect(log).toContain('logged');
expect(log).toContain('EXIT:');
});
});
describe('runGates', () => {
let tmp: string;
let logPath: string;
let eventsPath: string;
beforeEach(() => {
tmp = makeTmpDir();
logPath = join(tmp, 'gates.log');
eventsPath = join(tmp, 'events.ndjson');
});
afterEach(() => {
rmSync(tmp, { recursive: true, force: true });
});
it('runs multiple gates and returns results', () => {
const { allPassed, gateResults } = runGates(
['echo one', 'echo two'],
tmp,
logPath,
30,
eventsPath,
'task-1',
);
expect(allPassed).toBe(true);
expect(gateResults).toHaveLength(2);
});
it('reports failure when any gate fails', () => {
const { allPassed, gateResults } = runGates(
['echo ok', 'exit 1'],
tmp,
logPath,
30,
eventsPath,
'task-2',
);
expect(allPassed).toBe(false);
expect(gateResults[0]!.passed).toBe(true);
expect(gateResults[1]!.passed).toBe(false);
});
it('emits events for each gate', () => {
runGates(['echo test'], tmp, logPath, 30, eventsPath, 'task-3');
const events = readFileSync(eventsPath, 'utf-8')
.trim()
.split('\n')
.map((l) => JSON.parse(l));
expect(events).toHaveLength(2); // started + passed
expect(events[0].event_type).toBe('rail.check.started');
expect(events[1].event_type).toBe('rail.check.passed');
});
it('skips gates with empty command (non ci-pipeline)', () => {
const { gateResults } = runGates(
[{ command: '', type: 'mechanical' }, 'echo real'],
tmp,
logPath,
30,
eventsPath,
'task-4',
);
expect(gateResults).toHaveLength(1);
});
it('does not skip ci-pipeline even with empty command', () => {
const { gateResults } = runGates(
[{ command: '', type: 'ci-pipeline' }],
tmp,
logPath,
30,
eventsPath,
'task-5',
);
expect(gateResults).toHaveLength(1);
expect(gateResults[0]!.passed).toBe(true);
});
it('emits failed event with correct message', () => {
runGates(['exit 42'], tmp, logPath, 30, eventsPath, 'task-6');
const events = readFileSync(eventsPath, 'utf-8')
.trim()
.split('\n')
.map((l) => JSON.parse(l));
const failEvent = events.find(
(e: Record<string, unknown>) => e.event_type === 'rail.check.failed',
);
expect(failEvent).toBeDefined();
expect(failEvent.message).toContain('Gate failed (');
});
});
+163 -1
View File
@@ -1,5 +1,8 @@
import { describe, it, expect } from 'vitest';
import { describe, it, expect, afterEach, beforeEach, vi } from 'vitest';
import { Command } from 'commander';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { registerMacpCommand } from './cli.js';
describe('registerMacpCommand', () => {
@@ -75,3 +78,162 @@ describe('registerMacpCommand', () => {
expect(topLevel).toContain('events');
});
});
/**
* RI-N2 fail-closed CLI behavior: an unimplemented capability is a failure,
* never a success. Every stub exits nonzero with a typed message, and the
* implemented `macp gate` mirrors the typed gate-runner states.
*/
describe('registerMacpCommand fail-closed (RI-N2)', () => {
let tmpDir: string;
function buildProgram(): Command {
const program = new Command();
program.exitOverride();
program.configureOutput({ writeErr: () => {} });
registerMacpCommand(program);
return program;
}
beforeEach(() => {
tmpDir = fs.mkdtempSync(path.join(os.tmpdir(), 'macp-cli-failclosed-'));
process.exitCode = 0;
});
afterEach(() => {
process.exitCode = 0;
fs.rmSync(tmpDir, { recursive: true, force: true });
});
it('macp tasks list exits nonzero (unimplemented capability)', async () => {
const program = buildProgram();
await program.parseAsync(['macp', 'tasks', 'list'], { from: 'user' });
expect(process.exitCode).not.toBe(0);
});
it('macp submit exits nonzero with a typed MACP_NOT_IMPLEMENTED message', async () => {
const program = buildProgram();
const errSpy = vi.spyOn(console, 'error').mockImplementation(() => {});
try {
await program.parseAsync(['macp', 'submit', 'spec.json'], { from: 'user' });
expect(process.exitCode).not.toBe(0);
const errText = errSpy.mock.calls.map((c) => String(c[0])).join('\n');
expect(errText).toContain('MACP_NOT_IMPLEMENTED');
} finally {
errSpy.mockRestore();
}
});
it('macp events tail exits nonzero (unimplemented capability)', async () => {
const program = buildProgram();
await program.parseAsync(['macp', 'events', 'tail'], { from: 'user' });
expect(process.exitCode).not.toBe(0);
});
it('macp gate runs a green inline command and exits 0', async () => {
const program = buildProgram();
await program.parseAsync(
[
'macp',
'gate',
'exit 0',
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
expect(process.exitCode).toBe(0);
});
it('macp gate exits nonzero on a failing command', async () => {
const program = buildProgram();
await program.parseAsync(
[
'macp',
'gate',
'exit 9',
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
expect(process.exitCode).not.toBe(0);
});
it('macp gate with an unimplemented ci-pipeline capability exits nonzero', async () => {
const program = buildProgram();
const specPath = path.join(tmpDir, 'gates.json');
fs.writeFileSync(specPath, JSON.stringify([{ type: 'ci-pipeline' }]));
await program.parseAsync(
[
'macp',
'gate',
specPath,
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
expect(process.exitCode).not.toBe(0);
});
it('macp gate --simulate completes (exit 0) but reports simulated results', async () => {
const program = buildProgram();
const logSpy = vi.spyOn(console, 'log').mockImplementation(() => {});
try {
await program.parseAsync(
[
'macp',
'gate',
'exit 0',
'--simulate',
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
// completes only because the caller explicitly asked to simulate
expect(process.exitCode).toBe(0);
const outText = logSpy.mock.calls.map((c) => String(c[0])).join('\n');
expect(outText).toContain('simulated');
expect(outText).toContain('SIMULATED');
} finally {
logSpy.mockRestore();
}
});
it('macp gate with an empty spec exits nonzero with a typed error', async () => {
const program = buildProgram();
await program.parseAsync(
[
'macp',
'gate',
' ',
'--cwd',
tmpDir,
'--log',
path.join(tmpDir, 'g.log'),
'--timeout',
'10',
],
{ from: 'user' },
);
expect(process.exitCode).not.toBe(0);
});
});
+129 -19
View File
@@ -1,5 +1,73 @@
import { existsSync, readFileSync } from 'node:fs';
import type { Command } from 'commander';
import { runGates } from './gate-runner.js';
import { MACPCapabilityError, type MacpErrorCode } from './errors.js';
/**
* Load gates from a spec: an existing file (JSON gates array, a JSON object
* with `quality_gates`, a JSON gate object, or one command per line) or an
* inline command string. Fails closed with a typed capability error when the
* spec contains no executable gate definition.
*/
function loadGateSpec(spec: string): unknown[] {
if (existsSync(spec)) {
const raw = readFileSync(spec, 'utf-8');
try {
const parsed = JSON.parse(raw) as unknown;
if (Array.isArray(parsed)) {
if (parsed.length === 0) {
throw new MACPCapabilityError(
'MACP_NO_COMMAND',
'gate-spec',
`gate spec file '${spec}' contains an empty gates array`,
);
}
return parsed;
}
if (typeof parsed === 'object' && parsed !== null) {
const obj = parsed as Record<string, unknown>;
if (Array.isArray(obj['quality_gates'])) {
return obj['quality_gates'];
}
return [parsed];
}
throw new MACPCapabilityError(
'MACP_NO_COMMAND',
'gate-spec',
`gate spec file '${spec}' parsed to ${typeof parsed} — expected a gates array, a task with quality_gates, or a gate object`,
);
} catch (exc) {
if (exc instanceof MACPCapabilityError) throw exc;
// Not JSON — treat each non-empty line as a command gate.
const lines = raw
.split('\n')
.map((l) => l.trim())
.filter((l) => l.length > 0);
if (lines.length > 0) return lines;
throw new MACPCapabilityError(
'MACP_NO_COMMAND',
'gate-spec',
`gate spec file '${spec}' contains no gates`,
);
}
}
if (spec.trim().length > 0) return [spec];
throw new MACPCapabilityError('MACP_NO_COMMAND', 'gate-spec', 'gate spec is empty');
}
/** Print a typed not-implemented failure and exit nonzero (RI-N2 fail-closed). */
function notImplemented(subcommand: string, capability: string, hint: string): void {
const err = new MACPCapabilityError(
'MACP_NOT_IMPLEMENTED',
capability,
`${subcommand} is not implemented in @mosaicstack/macp yet (${capability} capability absent) — ${hint}`,
);
console.error(`[macp] ${subcommand}: ${err.message} [${err.code}]`);
process.exitCode = 1;
}
/**
* Register macp subcommands on an existing Commander program.
* This avoids cross-package Commander version mismatches by using the
@@ -24,15 +92,14 @@ export function registerMacpCommand(parent: Command): void {
'Filter by task type (coding|deploy|research|review|documentation|infrastructure)',
)
.action((opts: { status?: string; type?: string }) => {
// not yet wired — task persistence layer is not present in @mosaicstack/macp
console.log('[macp] tasks list: not yet wired — use macp package programmatically');
// unimplemented capability — a failure, never a success (RI-N2)
if (opts.status) {
console.log(` status filter: ${opts.status}`);
}
if (opts.type) {
console.log(` type filter: ${opts.type}`);
}
process.exitCode = 0;
notImplemented('tasks list', 'task-persistence', 'use the macp package programmatically');
});
// ─── submit ──────────────────────────────────────────────────────────────
@@ -41,12 +108,11 @@ export function registerMacpCommand(parent: Command): void {
.command('submit <path>')
.description('Submit a task from a JSON/YAML spec file')
.action((specPath: string) => {
// not yet wired — task submission requires a running MACP server
console.log('[macp] submit: not yet wired — use macp package programmatically');
// unimplemented capability — a failure, never a success (RI-N2)
console.log(` spec path: ${specPath}`);
console.log(' task id: (unavailable — no MACP server connected)');
console.log(' status: (unavailable — no MACP server connected)');
process.exitCode = 0;
notImplemented('submit', 'macp-server', 'use the macp package programmatically');
});
// ─── gate ────────────────────────────────────────────────────────────────
@@ -58,16 +124,58 @@ export function registerMacpCommand(parent: Command): void {
.option('--cwd <path>', 'Working directory for gate execution', process.cwd())
.option('--log <path>', 'Path to write gate log output', '/tmp/macp-gate.log')
.option('--timeout <seconds>', 'Gate timeout in seconds', '60')
.action((spec: string, opts: { failOn: string; cwd: string; log: string; timeout: string }) => {
// not yet wired — gate execution requires a task context and event sink
console.log('[macp] gate: not yet wired — use macp package programmatically');
console.log(` spec: ${spec}`);
console.log(` fail-on: ${opts.failOn}`);
console.log(` cwd: ${opts.cwd}`);
console.log(` log: ${opts.log}`);
console.log(` timeout: ${opts.timeout}s`);
process.exitCode = 0;
});
.option(
'--simulate',
'Simulate gates instead of executing them; results are typed simulated and never satisfy a check',
)
.action(
(
spec: string,
opts: { failOn: string; cwd: string; log: string; timeout: string; simulate?: boolean },
) => {
let gates: unknown[];
try {
gates = loadGateSpec(spec);
} catch (exc) {
if (exc instanceof MACPCapabilityError) {
console.error(`[macp] gate: ${exc.message} [${exc.code}]`);
} else {
console.error(`[macp] gate: ${String(exc)}`);
}
process.exitCode = 1;
return;
}
const timeoutSec = Number.parseInt(opts.timeout, 10) || 60;
const eventsPath = `${opts.log}.events.ndjson`;
const { state, gateResults } = runGates(
gates,
opts.cwd,
opts.log,
timeoutSec,
eventsPath,
'macp-cli-gate',
{
simulate: opts.simulate,
},
);
for (const r of gateResults) {
const label = r.command || r.type;
const reason = r.reason ? `${r.reason}` : '';
console.log(`[macp] gate ${r.status}: ${label}${reason}`);
}
if (opts.simulate) {
console.log(
'[macp] SIMULATED run — every result is typed simulated and can never satisfy a gate, dependency, or release check',
);
}
// Simulated runs may complete (exit 0) only because the caller
// explicitly passed --simulate; the typed state stays 'simulated'.
process.exitCode = state === 'passed' || state === 'simulated' ? 0 : 1;
},
);
// ─── events ──────────────────────────────────────────────────────────────
@@ -79,14 +187,16 @@ export function registerMacpCommand(parent: Command): void {
.option('--file <path>', 'Path to the MACP events NDJSON file')
.option('--follow', 'Follow the file for new events (like tail -f)')
.action((opts: { file?: string; follow?: boolean }) => {
// not yet wired — event streaming requires a live event source
console.log('[macp] events tail: not yet wired — use macp package programmatically');
// unimplemented capability — a failure, never a success (RI-N2)
if (opts.file) {
console.log(` file: ${opts.file}`);
}
if (opts.follow) {
console.log(' mode: follow');
}
process.exitCode = 0;
notImplemented('events tail', 'event-source', 'use the macp package programmatically');
});
}
// Re-export so CLI consumers can surface typed capability codes.
export type { MacpErrorCode };
+35
View File
@@ -0,0 +1,35 @@
/** Typed error code from the closed MACP_ERROR_CODES set. */
export type MacpErrorCode = (typeof MACP_ERROR_CODES)[number];
/**
* Typed fail-closed capability errors (RI-N2, SDLC-D-035).
*
* MACP must fail closed when a required capability (executor, reviewer,
* command, CI provider, human authority) is absent. These typed codes mirror
* the Forge failure vocabulary (FORGE_NO_*) so both packages speak the same
* language: an unimplemented capability is a failure, never a stub success.
*/
/** Closed set of typed MACP capability error codes. */
export const MACP_ERROR_CODES = [
'MACP_NOT_IMPLEMENTED',
'MACP_NO_COMMAND',
'MACP_NO_REVIEWER',
'MACP_NO_CI_PIPELINE',
'MACP_NO_PROVIDER',
'MACP_AUTHORITY_REQUIRED',
] as const;
/** Raised when a required capability is missing and execution must fail closed. */
export class MACPCapabilityError extends Error {
/** Typed error code from the closed MACP_ERROR_CODES set. */
readonly code: MacpErrorCode;
/** The missing capability, e.g. `ci-provider`, `task-persistence`, `command`. */
readonly capability: string;
constructor(code: MacpErrorCode, capability: string, message: string) {
super(message);
this.name = 'MACPCapabilityError';
this.code = code;
this.capability = capability;
}
}
+429
View File
@@ -0,0 +1,429 @@
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { afterEach, beforeEach, describe, expect, it } from 'vitest';
import { countAIFindings, normalizeGate, runGate, runGates } from './gate-runner.js';
function makeTmpDir(): string {
return fs.mkdtempSync(path.join(os.tmpdir(), 'macp-gate-'));
}
describe('normalizeGate', () => {
it('normalizes a string to mechanical gate', () => {
expect(normalizeGate('echo test')).toEqual({
command: 'echo test',
type: 'mechanical',
fail_on: 'blocker',
});
});
it('normalizes an object gate with defaults', () => {
expect(normalizeGate({ command: 'lint' })).toEqual({
command: 'lint',
type: 'mechanical',
fail_on: 'blocker',
});
});
it('preserves explicit type and fail_on', () => {
expect(normalizeGate({ command: 'review', type: 'ai-review', fail_on: 'any' })).toEqual({
command: 'review',
type: 'ai-review',
fail_on: 'any',
});
});
it('handles non-string/non-object input', () => {
expect(normalizeGate(42)).toEqual({ command: '', type: 'mechanical', fail_on: 'blocker' });
expect(normalizeGate(null)).toEqual({ command: '', type: 'mechanical', fail_on: 'blocker' });
});
});
describe('countAIFindings', () => {
it('returns zeros for non-object', () => {
expect(countAIFindings(null)).toEqual({ blockers: 0, total: 0 });
expect(countAIFindings('string')).toEqual({ blockers: 0, total: 0 });
expect(countAIFindings([])).toEqual({ blockers: 0, total: 0 });
});
it('counts from stats block', () => {
const output = { stats: { blockers: 2, should_fix: 3, suggestions: 1 } };
expect(countAIFindings(output)).toEqual({ blockers: 2, total: 6 });
});
it('counts from findings array when stats has no blockers', () => {
const output = {
stats: { blockers: 0 },
findings: [{ severity: 'blocker' }, { severity: 'warning' }, { severity: 'blocker' }],
};
expect(countAIFindings(output)).toEqual({ blockers: 2, total: 3 });
});
it('uses stats blockers over findings array when stats has blockers', () => {
const output = {
stats: { blockers: 5 },
findings: [{ severity: 'blocker' }, { severity: 'warning' }],
};
// stats.blockers = 5, total from stats = 5+0+0 = 5, findings not used for total since stats total is non-zero
expect(countAIFindings(output)).toEqual({ blockers: 5, total: 5 });
});
it('counts findings length as total when stats has zero total', () => {
const output = {
findings: [{ severity: 'warning' }, { severity: 'info' }],
};
expect(countAIFindings(output)).toEqual({ blockers: 0, total: 2 });
});
});
describe('runGate', () => {
let tmp: string;
let logPath: string;
beforeEach(() => {
tmp = makeTmpDir();
logPath = path.join(tmp, 'gate.log');
});
afterEach(() => {
fs.rmSync(tmp, { recursive: true, force: true });
});
it('passes mechanical gate on exit 0', () => {
const result = runGate('echo hello', tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.exit_code).toBe(0);
expect(result.type).toBe('mechanical');
expect(result.output).toContain('hello');
});
it('fails mechanical gate on non-zero exit', () => {
const result = runGate('exit 1', tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.exit_code).toBe(1);
});
it('ci-pipeline fails closed without a CI provider (no placeholder pass)', () => {
const result = runGate({ command: 'anything', type: 'ci-pipeline' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.status).toBe('capability_failure');
expect(result.capability_code).toBe('MACP_NO_CI_PIPELINE');
expect(result.type).toBe('ci-pipeline');
expect(result.output).not.toBe('CI pipeline gate placeholder');
});
it('empty command is a typed capability failure, never a pass', () => {
const result = runGate({ command: '' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.status).toBe('capability_failure');
expect(result.capability_code).toBe('MACP_NO_COMMAND');
});
it('ai-review gate parses JSON output', () => {
const json = JSON.stringify({ stats: { blockers: 0, should_fix: 1 } });
const result = runGate({ command: `echo '${json}'`, type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(true);
expect(result.blockers).toBe(0);
expect(result.findings).toBe(1);
});
it('ai-review gate fails on blockers', () => {
const json = JSON.stringify({ stats: { blockers: 2 } });
const result = runGate({ command: `echo '${json}'`, type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.blockers).toBe(2);
});
it('ai-review gate with fail_on=any fails on any findings', () => {
const json = JSON.stringify({ stats: { blockers: 0, should_fix: 1 } });
const result = runGate(
{ command: `echo '${json}'`, type: 'ai-review', fail_on: 'any' },
tmp,
logPath,
30,
);
expect(result.passed).toBe(false);
expect(result.fail_on).toBe('any');
});
it('ai-review gate fails on invalid JSON output', () => {
const result = runGate({ command: 'echo "not json"', type: 'ai-review' }, tmp, logPath, 30);
expect(result.passed).toBe(false);
expect(result.parse_error).toBeDefined();
});
it('writes to log file', () => {
runGate('echo logged', tmp, logPath, 30);
const log = fs.readFileSync(logPath, 'utf-8');
expect(log).toContain('COMMAND: echo logged');
expect(log).toContain('logged');
expect(log).toContain('EXIT:');
});
});
describe('runGates', () => {
let tmp: string;
let logPath: string;
let eventsPath: string;
beforeEach(() => {
tmp = makeTmpDir();
logPath = path.join(tmp, 'gates.log');
eventsPath = path.join(tmp, 'events.ndjson');
});
afterEach(() => {
fs.rmSync(tmp, { recursive: true, force: true });
});
it('runs multiple gates and returns results', () => {
const { allPassed, gateResults } = runGates(
['echo one', 'echo two'],
tmp,
logPath,
30,
eventsPath,
'task-1',
);
expect(allPassed).toBe(true);
expect(gateResults).toHaveLength(2);
});
it('reports failure when any gate fails', () => {
const { allPassed, gateResults } = runGates(
['echo ok', 'exit 1'],
tmp,
logPath,
30,
eventsPath,
'task-2',
);
expect(allPassed).toBe(false);
expect(gateResults[0]!.passed).toBe(true);
expect(gateResults[1]!.passed).toBe(false);
});
it('emits events for each gate', () => {
runGates(['echo test'], tmp, logPath, 30, eventsPath, 'task-3');
const events = fs
.readFileSync(eventsPath, 'utf-8')
.trim()
.split('\n')
.map((l) => JSON.parse(l));
expect(events).toHaveLength(2); // started + passed
expect(events[0].event_type).toBe('rail.check.started');
expect(events[1].event_type).toBe('rail.check.passed');
});
it('does not silently skip gates with empty command — they become capability failures', () => {
const { gateResults, allPassed, state } = runGates(
[{ command: '', type: 'mechanical' }, 'echo real'],
tmp,
logPath,
30,
eventsPath,
'task-4',
);
expect(gateResults).toHaveLength(2);
expect(gateResults[0]!.status).toBe('capability_failure');
expect(gateResults[1]!.status).toBe('passed');
expect(allPassed).toBe(false);
expect(state).toBe('capability_failure');
});
it('does not skip ci-pipeline even with empty command — typed capability failure', () => {
const { gateResults, allPassed, state } = runGates(
[{ command: '', type: 'ci-pipeline' }],
tmp,
logPath,
30,
eventsPath,
'task-5',
);
expect(gateResults).toHaveLength(1);
expect(gateResults[0]!.passed).toBe(false);
expect(gateResults[0]!.status).toBe('capability_failure');
expect(allPassed).toBe(false);
expect(state).toBe('capability_failure');
});
it('emits failed event with correct message', () => {
runGates(['exit 42'], tmp, logPath, 30, eventsPath, 'task-6');
const events = fs
.readFileSync(eventsPath, 'utf-8')
.trim()
.split('\n')
.map((l) => JSON.parse(l));
const failEvent = events.find(
(e: Record<string, unknown>) => e.event_type === 'rail.check.failed',
);
expect(failEvent).toBeDefined();
expect(failEvent.message).toContain('Gate failed (');
});
});
/**
* RI-N2 / SDLC-D-035 fail-closed controls for the MACP gate runner.
*
* Invariant under test: `passed: true` occurs ONLY when a gate really executed
* and really exited green (`status === 'passed'`). Absent capabilities,
* manual sign-offs, and simulated runs are typed distinctly and can never
* make the aggregate `passed`.
*/
describe('gate-runner fail-closed (RI-N2)', () => {
let tmpDir: string;
let logPath: string;
let eventsPath: string;
beforeEach(() => {
tmpDir = makeTmpDir();
logPath = path.join(tmpDir, 'gate.log');
eventsPath = path.join(tmpDir, 'events.ndjson');
});
afterEach(() => {
fs.rmSync(tmpDir, { recursive: true, force: true });
});
function run(gates: unknown[], options?: { simulate?: boolean }) {
return runGates(gates, tmpDir, logPath, 10, eventsPath, 'spec-task', options);
}
// ─── positive controls ───────────────────────────────────────────────────
it('a really-executed green command gate still passes', () => {
const result = run([{ command: 'exit 0', type: 'mechanical' }]);
expect(result.gateResults[0]!.status).toBe('passed');
expect(result.gateResults[0]!.passed).toBe(true);
expect(result.allPassed).toBe(true);
expect(result.state).toBe('passed');
});
it('explicit simulate completes and types every result simulated', () => {
const result = run([{ command: 'exit 0', type: 'mechanical' }, 'echo hello'], {
simulate: true,
});
expect(result.gateResults).toHaveLength(2);
for (const gate of result.gateResults) {
expect(gate.status).toBe('simulated');
expect(gate.passed).toBe(false);
}
expect(result.state).toBe('simulated');
});
it('a really-executed red command gate fails with typed status failed', () => {
const result = run([{ command: 'exit 3', type: 'mechanical' }]);
expect(result.gateResults[0]!.status).toBe('failed');
expect(result.gateResults[0]!.passed).toBe(false);
expect(result.allPassed).toBe(false);
expect(result.state).toBe('failed');
});
// ─── negative controls — each asserts typed status AND aggregate not passed ──
it('an empty-command gate is a capability_failure, not skipped and not passed', () => {
const result = run([{ command: '', type: 'mechanical' }]);
// runGates must not silently skip it — it produces a typed result
expect(result.gateResults).toHaveLength(1);
const gate = result.gateResults[0]!;
expect(gate.status).toBe('capability_failure');
expect(gate.capability_code).toBe('MACP_NO_COMMAND');
expect(gate.passed).toBe(false);
// aggregate is not passed
expect(result.allPassed).toBe(false);
expect(result.state).toBe('capability_failure');
expect(result.state).not.toBe('passed');
});
it('a commandless ai-review gate is a typed MACP_NO_REVIEWER capability_failure', () => {
const result = run([{ command: '', type: 'ai-review' }]);
expect(result.gateResults[0]!.status).toBe('capability_failure');
expect(result.gateResults[0]!.capability_code).toBe('MACP_NO_REVIEWER');
expect(result.allPassed).toBe(false);
expect(result.state).not.toBe('passed');
});
it('a ci-pipeline gate without a provider implementation is a capability_failure, never a placeholder pass', () => {
const result = run([{ command: '', type: 'ci-pipeline' }]);
const gate = result.gateResults[0]!;
expect(gate.status).toBe('capability_failure');
expect(gate.capability_code).toBe('MACP_NO_CI_PIPELINE');
expect(gate.passed).toBe(false);
// the old false-success placeholder must be gone
expect(gate.output).not.toBe('CI pipeline gate placeholder');
expect(result.allPassed).toBe(false);
expect(result.state).not.toBe('passed');
});
it('a ci-pipeline gate fails closed even alongside an otherwise green run', () => {
const result = run(['exit 0', { type: 'ci-pipeline', command: 'fake-ci' }]);
expect(result.gateResults[1]!.status).toBe('capability_failure');
expect(result.gateResults[0]!.status).toBe('passed');
expect(result.allPassed).toBe(false);
expect(result.state).toBe('capability_failure');
});
it('a manual gate with no automation enters typed waiting — neither pass nor fail', () => {
const result = run([{ type: 'manual' }]);
const gate = result.gateResults[0]!;
expect(gate.status).toBe('waiting');
expect(gate.passed).toBe(false);
expect(gate.exit_code).toBe(0);
// aggregate is not passed while any gate is waiting
expect(result.allPassed).toBe(false);
expect(result.state).toBe('waiting');
expect(result.state).not.toBe('passed');
});
it('a simulated result can never make the aggregate passed', () => {
const result = run(['exit 0', 'exit 0'], { simulate: true });
expect(result.gateResults.every((g) => g.status === 'simulated')).toBe(true);
expect(result.allPassed).toBe(false);
expect(result.state).toBe('simulated');
expect(result.state).not.toBe('passed');
});
it('waiting dominates an otherwise green aggregate', () => {
const result = run(['exit 0', { type: 'manual' }]);
expect(result.allPassed).toBe(false);
expect(result.state).toBe('waiting');
});
});
describe('runGate fail-closed (RI-N2)', () => {
let tmpDir: string;
let logPath: string;
beforeEach(() => {
tmpDir = makeTmpDir();
logPath = path.join(tmpDir, 'gate.log');
});
afterEach(() => {
fs.rmSync(tmpDir, { recursive: true, force: true });
});
it('simulate: true returns a typed simulated result without executing', () => {
const result = runGate('this-command-does-not-exist-xyz', tmpDir, logPath, 10, {
simulate: true,
});
expect(result.status).toBe('simulated');
expect(result.passed).toBe(false);
expect(result.exit_code).toBe(0);
});
it('normal mode executes for real and types a green gate passed', () => {
const result = runGate('echo ok', tmpDir, logPath, 10);
expect(result.status).toBe('passed');
expect(result.passed).toBe(true);
expect(result.output).toContain('ok');
});
it('a bare string gate normalizes to mechanical and executes', () => {
const result = runGate('exit 7', tmpDir, logPath, 10);
expect(result.type).toBe('mechanical');
expect(result.status).toBe('failed');
expect(result.passed).toBe(false);
});
});
+148 -25
View File
@@ -4,7 +4,20 @@ import { dirname } from 'node:path';
import { emitEvent } from './event-emitter.js';
import { nowISO } from './event-emitter.js';
import type { GateResult } from './types.js';
import type { GateResult, GateStatus, RunGatesResult } from './types.js';
/** Typed reason stamped on every simulated gate result. */
export const SIMULATED_GATE_REASON =
'simulated execution (explicit simulate opt-in): gate was not evaluated by a real implementation';
/** Options for gate execution (RI-N2 fail-closed / explicit simulation). */
export interface RunGateOptions {
/**
* Explicit caller opt-in to simulation. Simulated gates are NOT executed;
* every result is typed `simulated` and never satisfies anything.
*/
simulate?: boolean;
}
export interface NormalizedGate {
command: string;
@@ -103,36 +116,91 @@ export function countAIFindings(parsedOutput: unknown): { blockers: number; tota
return { blockers, total };
}
function simulatedResult(gateEntry: NormalizedGate): GateResult {
return {
command: gateEntry.command,
exit_code: 0,
type: gateEntry.type,
output: SIMULATED_GATE_REASON,
timed_out: false,
passed: false,
status: 'simulated',
reason: SIMULATED_GATE_REASON,
};
}
function capabilityFailureResult(
gateEntry: NormalizedGate,
code: GateResult['capability_code'],
reason: string,
): GateResult {
return {
command: gateEntry.command,
exit_code: 1,
type: gateEntry.type,
output: '',
timed_out: false,
passed: false,
status: 'capability_failure',
capability_code: code,
reason,
};
}
function waitingResult(gateEntry: NormalizedGate, reason: string): GateResult {
return {
command: gateEntry.command,
exit_code: 0,
type: gateEntry.type,
output: '',
timed_out: false,
passed: false,
status: 'waiting',
capability_code: 'MACP_AUTHORITY_REQUIRED',
reason,
};
}
export function runGate(
gate: unknown,
cwd: string,
logPath: string,
timeoutSec: number,
options: RunGateOptions = {},
): GateResult {
const gateEntry = normalizeGate(gate);
const gateType = gateEntry.type;
const command = gateEntry.command;
// Explicit simulation only: never executes, typed simulated, never satisfying.
if (options.simulate) {
return simulatedResult(gateEntry);
}
// Fail closed: no CI provider implementation exists in @mosaicstack/macp,
// so a ci-pipeline gate is an absent capability — never a placeholder pass.
if (gateType === 'ci-pipeline') {
return {
command,
exit_code: 0,
type: gateType,
output: 'CI pipeline gate placeholder',
timed_out: false,
passed: true,
};
return capabilityFailureResult(
gateEntry,
'MACP_NO_CI_PIPELINE',
`ci-pipeline gate '${gateEntry.command || gateType}' has no CI provider implementation wired — refusing placeholder pass`,
);
}
if (!command) {
return {
command: '',
exit_code: 0,
type: gateType,
output: '',
timed_out: false,
passed: true,
};
// A manual gate with no automation waits for human sign-off: not pass, not fail.
if (gateType === 'manual') {
return waitingResult(
gateEntry,
`manual gate has no automation — waiting for human sign-off (type: ${gateType})`,
);
}
// Any other commandless gate is an absent capability — never a vacuous pass.
return capabilityFailureResult(
gateEntry,
gateType === 'ai-review' ? 'MACP_NO_REVIEWER' : 'MACP_NO_COMMAND',
`gate of type '${gateType}' has no command to execute — refusing empty-command pass`,
);
}
const { exitCode, output, timedOut } = runShell(command, cwd, logPath, timeoutSec);
@@ -143,10 +211,12 @@ export function runGate(
output,
timed_out: timedOut,
passed: false,
status: 'failed',
};
if (gateType !== 'ai-review') {
result.passed = exitCode === 0;
result.status = result.passed ? 'passed' : 'failed';
return result;
}
@@ -170,6 +240,7 @@ export function runGate(
} else {
result.passed = exitCode === 0 && blockers === 0 && !timedOut && parseError === undefined;
}
result.status = result.passed ? 'passed' : 'failed';
result.fail_on = failOn;
result.blockers = blockers;
@@ -191,16 +262,19 @@ export function runGates(
timeoutSec: number,
eventsPath: string,
taskId: string,
): { allPassed: boolean; gateResults: GateResult[] } {
let allPassed = true;
options: RunGateOptions = {},
): RunGatesResult {
const gateResults: GateResult[] = [];
let hasCapabilityFailure = false;
let hasSimulated = false;
let hasFailed = false;
let hasWaiting = false;
for (const gate of gates) {
const gateEntry = normalizeGate(gate);
const gateCmd = gateEntry.command;
if (!gateCmd && gateEntry.type !== 'ci-pipeline') continue;
const label = gateCmd || gateEntry.type;
// NOTE: no silent skip — every gate produces a typed result (RI-N2).
emitEvent(
eventsPath,
'rail.check.started',
@@ -209,10 +283,10 @@ export function runGates(
'quality-gate',
`Running gate: ${label}`,
);
const result = runGate(gate, cwd, logPath, timeoutSec);
const result = runGate(gate, cwd, logPath, timeoutSec, options);
gateResults.push(result);
if (result.passed) {
if (result.status === 'passed') {
emitEvent(
eventsPath,
'rail.check.passed',
@@ -224,7 +298,46 @@ export function runGates(
continue;
}
allPassed = false;
if (result.status === 'waiting') {
hasWaiting = true;
emitEvent(
eventsPath,
'rail.check.waiting',
taskId,
'gated',
'quality-gate',
`Gate waiting: ${label}${result.reason ?? 'manual gate awaits sign-off'}`,
);
continue;
}
if (result.status === 'simulated') {
hasSimulated = true;
emitEvent(
eventsPath,
'rail.check.simulated',
taskId,
'gated',
'quality-gate',
`Gate simulated (non-satisfying): ${label}`,
);
continue;
}
if (result.status === 'capability_failure') {
hasCapabilityFailure = true;
emitEvent(
eventsPath,
'rail.check.failed',
taskId,
'gated',
'quality-gate',
`Gate capability failure (${result.capability_code ?? 'MACP_NO_PROVIDER'}): ${label}${result.reason ?? 'required capability is absent'}`,
);
continue;
}
hasFailed = true;
let message: string;
if (result.timed_out) {
message = `Gate timed out after ${timeoutSec}s: ${label}`;
@@ -236,5 +349,15 @@ export function runGates(
emitEvent(eventsPath, 'rail.check.failed', taskId, 'gated', 'quality-gate', message);
}
return { allPassed, gateResults };
const state: GateStatus = hasCapabilityFailure
? 'capability_failure'
: hasSimulated
? 'simulated'
: hasFailed
? 'failed'
: hasWaiting
? 'waiting'
: 'passed';
return { allPassed: state === 'passed', gateResults, state };
}
+16 -2
View File
@@ -6,11 +6,13 @@ export type {
DependsOnPolicy,
GateType,
GateFailOn,
GateStatus,
GateEntry,
Task,
EventType,
MACPEvent,
GateResult,
RunGatesResult,
TaskResult,
ProviderMeta,
ProviderRegistry,
@@ -18,6 +20,11 @@ export type {
export { CredentialError } from './types.js';
// Typed fail-closed capability errors (RI-N2, SDLC-D-035)
export { MACP_ERROR_CODES, MACPCapabilityError } from './errors.js';
export type { MacpErrorCode } from './errors.js';
// Credential resolver
export {
DEFAULT_CREDENTIALS_DIR,
@@ -35,9 +42,16 @@ export {
export type { ResolveCredentialsOptions } from './credential-resolver.js';
// Gate runner
export { normalizeGate, runShell, countAIFindings, runGate, runGates } from './gate-runner.js';
export {
normalizeGate,
runShell,
countAIFindings,
runGate,
runGates,
SIMULATED_GATE_REASON,
} from './gate-runner.js';
export type { NormalizedGate } from './gate-runner.js';
export type { NormalizedGate, RunGateOptions } from './gate-runner.js';
// Risk-floor (agent reflection loop — diff review classifier)
export { evaluateRiskFloor, DEFAULT_RISK_THRESHOLD } from './risk-floor.js';
+39 -2
View File
@@ -1,3 +1,5 @@
import type { MacpErrorCode } from './errors.js';
/** Task status values. */
export type TaskStatus = 'pending' | 'running' | 'gated' | 'completed' | 'failed' | 'escalated';
@@ -17,7 +19,17 @@ export type DispatchMode = 'yolo' | 'acp' | 'exec';
export type DependsOnPolicy = 'all' | 'any' | 'all_terminal';
/** Quality gate type. */
export type GateType = 'mechanical' | 'ai-review' | 'ci-pipeline';
export type GateType = 'mechanical' | 'ai-review' | 'ci-pipeline' | 'manual';
/**
* Typed execution state of a gate closed set (RI-N2, SDLC-D-035).
*
* Only `passed` means "really executed and green". `simulated` is produced
* exclusively under an explicit simulate opt-in and never satisfies anything.
* `capability_failure` means a required executor/provider/command was absent.
* `waiting` means a manual gate awaits human sign-off (neither pass nor fail).
*/
export type GateStatus = 'passed' | 'failed' | 'simulated' | 'waiting' | 'capability_failure';
/** Gate fail_on mode. */
export type GateFailOn = 'blocker' | 'any';
@@ -67,7 +79,9 @@ export type EventType =
| 'task.retry.scheduled'
| 'rail.check.started'
| 'rail.check.passed'
| 'rail.check.failed';
| 'rail.check.failed'
| 'rail.check.waiting'
| 'rail.check.simulated';
/** Structured event record. */
export interface MACPEvent {
@@ -88,7 +102,14 @@ export interface GateResult {
type: string;
output: string;
timed_out: boolean;
/** Back-compat boolean view — true ONLY when `status === 'passed'`. */
passed: boolean;
/** Typed discriminator — the authoritative gate outcome (RI-N2). */
status: GateStatus;
/** Typed capability error code, set when `status === 'capability_failure'`. */
capability_code?: MacpErrorCode;
/** Why a non-executed state (simulated/waiting/capability_failure) was reached. */
reason?: string;
fail_on?: string;
blockers?: number;
findings?: number;
@@ -96,6 +117,22 @@ export interface GateResult {
parse_error?: string;
}
/**
* Aggregate outcome of `runGates` (RI-N2).
*
* `state` is the typed aggregate: it is `passed` only when every gate really
* executed green. A `simulated` result makes the aggregate `simulated` (never
* `passed`); a `waiting` manual gate keeps the aggregate `waiting`; a missing
* capability makes it `capability_failure`. `allPassed` is exactly
* `state === 'passed'`, so a simulated or waiting result can never satisfy a
* dependency, acceptance criterion, gate, merge, or release check.
*/
export interface RunGatesResult {
allPassed: boolean;
gateResults: GateResult[];
state: GateStatus;
}
/** Result from a completed task. */
export interface TaskResult {
task_id: string;
+3 -1
View File
@@ -13,7 +13,8 @@ Pi is the native Mosaic agent runtime. The `mosaic pi` launcher:
1. Injects the full runtime contract via `--append-system-prompt`
2. Loads Mosaic skills via `--skill` flags
3. Loads the Mosaic extension via `--extension` for lifecycle hooks
3. Loads framework-owned `mosaic-extension.ts` and `goal-extension.ts` from
`~/.config/mosaic/runtime/pi/` via ordered `--extension` flags
4. Detects active missions and injects initial prompts
## Capabilities vs Other Runtimes
@@ -22,6 +23,7 @@ Pi is the native Mosaic agent runtime. The `mosaic pi` launcher:
- Native thinking levels replace sequential-thinking MCP
- Native skill discovery compatible with Mosaic SKILL.md format
- Native extension system for lifecycle hooks (TypeScript, not bash shims)
- Bounded persistent `/goal` loop with per-turn, post-compaction, and two-pass evidence checks
- Native session persistence and resume
- Model-agnostic (Anthropic, OpenAI, Google, Ollama, custom providers)
@@ -43,6 +43,8 @@ overwritten on upgrade. (Layer model: `constitution/LAYER-MODEL.md`.)
| Secrets / vault usage | `guides/VAULT-SECRETS.md` |
| Tool/credential reference (service CLIs, wrappers) | `guides/TOOLS-REFERENCE.md` |
| Memory protocol (OpenBrain capture/recall) | `guides/MEMORY.md` |
| Seat identity, git credentials, token slots | `guides/SEAT-IDENTITY.md` |
| Reaching another agent (fleet comms) | `guides/FLEET-COMMS.md` |
## Subagent Model Selection (Cost — Hard Rule)
+14 -7
View File
@@ -104,7 +104,14 @@ The launcher:
1. Verifies `~/.config/mosaic` exists
2. Verifies `SOUL.md` exists (auto-runs `mosaic init` if missing)
3. Injects `AGENTS.md` into the runtime
4. Forwards all arguments to the runtime CLI
4. For Pi, loads the framework-owned core and persistent-goal extensions from
`~/.config/mosaic/runtime/pi/`
5. Forwards all arguments to the runtime CLI
Inside `mosaic pi`, `/goal set <statement>` starts a bounded persistent goal loop. Use `/goal status`,
`/goal pause`, `/goal resume`, or `/goal cancel` to control it. The extension remains part of Mosaic
under `~/.config/mosaic/runtime/pi/goal-extension.ts`; it is not installed in Pi's main extension
directory.
You can still launch runtimes directly (`claude`, `codex`, etc.) — thin runtime adapters will tell the agent to read `~/.config/mosaic/AGENTS.md`.
@@ -124,9 +131,9 @@ You can still launch runtimes directly (`claude`, `codex`, etc.) — thin runtim
│ ├── claude/ ← CLAUDE.md, RUNTIME.md, settings.json, hooks
│ ├── codex/ ← instructions.md, RUNTIME.md
│ ├── opencode/ ← AGENTS.md, RUNTIME.md
│ ├── pi/ ← RUNTIME.md, mosaic-extension.ts
│ ├── pi/ ← RUNTIME.md, mosaic-extension.ts, goal-extension.ts
│ └── mcp/ ← MCP server configs
├── skills/ ← Universal skills (synced from mosaic/agent-skills)
├── skills/ ← Universal skills (shipped with the framework package)
├── skills-local/ ← Local cross-runtime skills
├── memory/ ← Persistent agent memory (preserved across upgrades)
└── templates/ ← SOUL.md template, project templates
@@ -136,7 +143,7 @@ You can still launch runtimes directly (`claude`, `codex`, etc.) — thin runtim
| Launch method | Injection mechanism |
| ------------------- | ----------------------------------------------------------------------------------------- |
| `mosaic pi` | `--append-system-prompt` with composed runtime contract + skills + extension |
| `mosaic pi` | `--append-system-prompt` with composed runtime contract + skills + Mosaic extensions |
| `mosaic claude` | `--append-system-prompt` with composed runtime contract (`AGENTS.md` + runtime reference) |
| `mosaic codex` | Writes composed runtime contract to `~/.codex/instructions.md` before launch |
| `mosaic opencode` | Writes composed runtime contract to `~/.config/opencode/AGENTS.md` before launch |
@@ -193,11 +200,11 @@ The installer rejects unrecognized flags or positional arguments before making c
## Universal Skills
The installer syncs skills from `mosaic/agent-skills` into `~/.config/mosaic/skills/`. Install, wizard finalization, and `mosaic update` automatically reconcile every canonical skill into Claude Code's `~/.claude/skills/` directory.
Canonical skills ship inside the framework package itself; the installer installs them into `~/.config/mosaic/skills/` together with the rest of the framework (there is no separate skills repository). Install, wizard finalization, and `mosaic update` automatically link every canonical skill into Claude Code's `~/.claude/skills/` directory.
```bash
mosaic sync # Full canonical catalog sync
~/.config/mosaic/tools/_scripts/mosaic-sync-skills --link-only # Re-link only
mosaic sync # Relink the full canonical catalog
~/.config/mosaic/tools/_scripts/mosaic-sync-skills --link-only # Re-link only (same as default)
mosaic skill list # Show registered, missing, dangling, and foreign entries
mosaic skill register <name> # Register or repair one canonical Claude link
mosaic skill unregister <name> # Remove one Mosaic-owned Claude link
@@ -60,6 +60,52 @@ If a repo does not expose these scripts, run equivalent local workflow commands
- Do not auto-resolve data conflicts in shared state files.
- Keep commits scoped to a single logical change set.
## Model Tiering
Model choice is a standard, not a preference. Delegating a mechanical grep to a
frontier reasoning model wastes budget; sending a security review to a cheap tier
produces a review that passes and proves nothing. Both are defects.
Tiers are named by **capability class**, so the standard survives a model
generation. An operator binds each class to a concrete model id.
| Class | Use for |
| ------------- | ----------------------------------------------------------------------------------------- |
| `search` | grep/glob, file location, status and health checks, one-line mechanical edits |
| `build` | feature implementation, test writing, bugfixes, routine refactors |
| `judge` | code review, planning, API/compat-sensitive changes |
| `adversarial` | security review, ambiguous architecture, anything where a wrong "looks fine" is expensive |
Rules:
1. **Start at the cheapest class that can do the task; escalate on evidence, not
on nerves.** Omitting a tier is not neutral — it inherits the caller's model,
which is usually the most expensive one.
2. **Compat-sensitive work escalates one class.** A change that must interoperate
with an existing contract is judged, not just built.
3. **A tier assignment is benchmarked, not asserted.** Move a task class to a
cheaper tier only against a blind A/B on real work from this codebase, ranked
by someone other than the author. "It seemed fine" is not evidence.
4. **Reviewer independence beats reviewer size.** An `adversarial` verdict from
the model that wrote the code is not a second opinion (see Constitution gate 16).
### Where the binding lives
The class→model map is operator configuration, never framework source: model
availability, cost, and quotas differ per operator and per host.
Resolution order, first hit wins:
1. the config service (DB-backed, surfaced and editable in the Mosaic webUI)
2. a local operator file (`STANDARDS.local.md`, or `policy/` where the runtime
injects it)
3. the framework default — the class names above, with no binding
Only layer 1 is auditable across a fleet, so it is the target end state; layers 2
and 3 exist so a host with no config service still runs. A local override that
silently disagrees with the config service is drift — the same failure class the
tool-index gate exists to catch, and it belongs in `mosaic doctor`.
## Prompting Contract
All runtime adapters should inject:
@@ -15,8 +15,78 @@ Merge strategy enforcement (HARD RULE):
- Merge to `main` MUST be squash-only.
- Use `~/.config/mosaic/tools/git/pr-merge.sh -n {PR_NUMBER} -m squash --expect-head {approved_full_sha}` (or PowerShell equivalent).
An estate MAY carry a documented exception for a repository whose gates are commit hooks rather
than review. Such an exception belongs in that estate's own working copy of this guide, is
scoped to the named repository, and is never precedent for a second one.
**Do not use `pr-review.sh` or `issue-comment.sh` to post a verdict** (mosaicstack#1280). Post
through a direct authenticated API call as your own seat, or hand the verdict to the requesting
seat. Handing it over is a legitimate delivery path, not a fallback.
## Evidence Discipline (applies to every finding)
The checklist below says what to look at. This section says when you are allowed to believe what
you saw. Every rule here was earned by a wrong conclusion that reached a report.
1. **A finding is a claim about behavior.** State the failing input, the path taken, and the
wrong result. "This looks fragile" is not a finding.
2. **A green check is not a result until you have shown it could go red.** Run the control. A
`0`, an empty result, or a column of identical values with no failing counterpart is a
non-result.
3. **Measurement and explanation are separate sentences.** Report the command and its output,
then, as its own sentence, what you think it means.
4. **Never widen the case you measured.** If you checked one path, the finding covers one path.
5. **Reproduce a reported failure before recording it, and say which tree you measured.** Two
correct measurements of two different trees disagree without either being wrong.
6. **Verify by content on the ref that ships**, never by ancestry of a local sha. A rebase mints
new shas; a commit being an ancestor of something local proves nothing about the remote.
Compare by digest against `origin/<branch>`.
7. **Confidence is part of the finding.** "I could not reproduce this" is a usable review
comment. A confident guess is not.
8. **Author is not reviewer** (Gate-16). Do not review your own work, or work you shaped closely
enough to be a co-author of. Say so and hand it back.
### Measuring a shell suite
Each of these produced a wrong conclusion before it was written down.
9. **`cmd | tail; echo rc=$?` reports `tail`'s exit code, not `cmd`'s.** It reads as a pass when
the command failed. Redirect to a file and check `rc` directly, or use `${PIPESTATUS[0]}`.
10. **Under `set -o pipefail`, a missed glob makes `ls` exit 2**, the pipeline inherits it, and
`set -e` kills the run. Iterate a glob with a `for` loop and an `-e` test instead of piping
`ls`.
11. **A suite that exits nonzero with ZERO output is an environment question, not a defect in
the code under review.** The usual cause is a sourced dependency that is absent, so `set -e`
kills the first case before anything prints. Extract whole tool trees — `tools/git` alone is
missing `tools/_lib/credentials.sh`. Isolate the variable and prove it by adding only that
back.
12. **`git -C <dir>` in a directory that is not itself a repo answers from the enclosing repo.**
A scratch tree under `~/.mosaic` reports `~/.mosaic`'s HEAD, not the PR's, and every
conclusion drawn from it describes the wrong tree. Confirm `git rev-parse --show-toplevel`
is the tree you think it is before trusting any git output.
### Feedback Categories
- **Blocker**: must fix before merge (security, bugs, test failures)
- **Should Fix**: important but not blocking (code quality, minor issues)
- **Suggestion**: optional improvement (style preference, nice-to-have)
- **Question**: seeking clarification
## Review Checklist
Reviewer seats split this checklist by class rather than duplicating it. A seat reviews its own
sections in full and may raise anything it notices outside them as a Suggestion, never as a
Blocker on someone else's ground.
| Reviewer class | Owns |
| ---------------- | ------------------------------------------------------------------------------------------------------- |
| `rev-code-*` | 1 Correctness, 3 Testing, 4 Code Quality, 4a TypeScript, 5 Documentation, 6 Performance, 7 Dependencies |
| `rev-security-*` | 2 Security, 2a OWASP |
Where two seats of the same class review the same change, they review independently and compare
after. A second seat that reads the first seat's findings before measuring is a proofreader, not
a second opinion.
### 1. Correctness
- [ ] Code does what the issue/PR description says
@@ -53,7 +123,7 @@ Merge strategy enforcement (HARD RULE):
- [ ] Tests cover happy path AND error cases
- [ ] Situational tests cover all impacted change surfaces (primary gate)
- [ ] Tests validate required behavior/outcomes, not only internal implementation details
- [ ] TDD was applied when required by `~/.config/mosaic/guides/QA-TESTING.md`
- [ ] TDD was applied when required by `guides/QA-TESTING.md`
- [ ] Coverage meets 85% minimum
- [ ] Tests are readable and maintainable
- [ ] No flaky tests introduced
@@ -82,7 +152,7 @@ Merge strategy enforcement (HARD RULE):
### 5. Documentation
- [ ] Complex logic has explanatory comments
- [ ] Required docs updated per `~/.config/mosaic/guides/DOCUMENTATION.md`
- [ ] Required docs updated per `guides/DOCUMENTATION.md`
- [ ] Public APIs are documented
- [ ] Private/internal APIs are documented
- [ ] API input/output schemas are documented
@@ -126,13 +196,6 @@ git diff main...HEAD
- Distinguish between blocking issues and suggestions
- Be constructive, not critical of the person
### Feedback Categories
- **Blocker**: Must fix before merge (security, bugs, test failures)
- **Should Fix**: Important but not blocking (code quality, minor issues)
- **Suggestion**: Optional improvements (style preferences, nice-to-haves)
- **Question**: Seeking clarification
### Review Comment Format
```
@@ -0,0 +1,86 @@
# Fleet Comms Guide
How one seat reaches another on a host. The mechanism is the framework's; the sessions and
sockets are per-host, so measure yours rather than trusting an example.
`mosaic <runtime>` would normally inject the addressing block from the roster. Where the composer
is unavailable, or where the roster is stale, this guide is the substitute.
## Measure the fleet; do not trust the roster
`fleet/roster.yaml` is a declaration of intent, not an observation. It routinely names a socket
that was never created, lists seats that are not running, and omits seats that are — this was
all three have been observed true at once on a live host. Find out what is actually
up before addressing anyone:
```bash
tmux list-sessions
tmux list-panes -a -F '#{session_name} #{pane_current_command} #{pane_current_path}'
```
The pane command tells you the runtime. A pane showing `bash` is an idle shell with no agent
attached — a send there lands in a shell prompt and is not read by anyone.
Use the **default socket**. Do not pass `-L mosaic-fleet` on the strength of the roster.
## Sending
```bash
~/.config/mosaic/tools/tmux/agent-send.sh -s <dst_session> -C <class> -m "<message>"
```
`-s` also accepts `session:window.pane`. `-f <file>` sends a file body; stdin works too.
### Classes
`-C` takes exactly one of these. Anything else exits 3.
| Class | Use for |
| -------------- | -------------------------------------------------------- |
| `terminal-log` | log only; never needs the agent's attention |
| `actionable` | a decision, blocker, gate, or question needing an answer |
| `human` | relayed from a human operator |
| `reaction` | an ack or acknowledgement token |
| `digest` | machine wake, coalescible |
An absent class is treated as `actionable` by consumers, which is the fail-safe direction. Prefer
naming it anyway.
### Addressing preamble
The wire format is `[<src> -> <dst> class=<class>] <body>`. Flip it when you reply — the tool
sends, it does not auto-reply.
### Exit codes
| rc | Meaning |
| --- | ---------------------------------------------- |
| 0 | delivered or queued |
| 1 | target session not found |
| 2 | text reached the pane but is **still a draft** |
| 3 | usage error (bad class, missing `-s`) |
**Never retry on rc=2.** The message is in the target pane; retrying double-sends it. Confirm
instead:
```bash
tmux capture-pane -p -t <session>:0.0 | tail -20
```
rc=2 is the normal result when the target is an idle pi seat.
## Durable comms
tmux delivery is host-local and does not survive a pane. Anything that must outlive the session
goes through the estate's durable comms protocol — a committed `comms/` tree in an estate repo,
with its own README. Use it for cross-host messages, verdicts, and anything a later session needs
to find.
## Handing work across seats
1. **A verdict handed to the requesting seat is a legitimate delivery path**, and the required one
for anything `pr-review.sh` would otherwise post (see `guides/CODE-REVIEW.md`).
2. **Address the seat, not the runtime.** A seat name is a session name; whether it runs claude,
pi or codex is not the sender's business.
3. **Say what you measured, not just what you concluded** — the receiving seat cannot see your
terminal.
@@ -219,7 +219,7 @@ Use the Cloudflare tools for any DNS configuration: pointing domains at services
# Update an existing record (get record ID from record-list first)
~/.config/mosaic/tools/cloudflare/record-update.sh \
-z example.com -r <record-id> -t A -n myapp -c 10.0.0.5 -p
-z example.com -r <record-id> -t A -n myapp -c 192.0.2.5 -p
```
**DNS + Deployment integration**: When deploying a new service via Coolify or Portainer that needs a public domain, the typical sequence is:
@@ -0,0 +1,133 @@
# Seat Identity & Credentials Guide
Every agent that touches a Mosaic-managed git host acts as a named seat with its own credential.
This guide is how that works on a host, and what an agent must never do with it.
The mechanism below is the framework's. The specific paths, seats and stores are per-host:
measure yours before trusting any of them.
## The rule
**One seat, one identity, one token file.** A seat never borrows another seat's credential, never
falls back to a shared owner account, and never carries a second copy of its own token. A second
copy is drift, and drift surfaces as the stale copy returning 401 — which reads as a revoked
token and sends whoever debugs it somewhere else entirely.
A credential refusal is correct behavior, not a bug to route around. If git refuses with a
fail-closed diagnostic, the fix is to provision or correct _your_ identity. Escalate; do not
substitute.
## How a credential is resolved
Find the helper the way **git** does, not with `command -v`. Git runs whatever
`credential.helper` names, and on a Mosaic host that is an absolute path — so a PATH lookup
answers a different question and the two disagree the moment the PATH copy is removed. It was
removed on hosts that have completed that migration.
```bash
git config --get-all credential.helper # every helper, in the order git tries them
```
Git tries **each** configured helper in turn until one supplies a credential. A fail-closed
helper supplies nothing, so a second helper configured behind it silently becomes the one that
answers. When you care which binary serves a credential, read the whole list.
Resolve all three forms git accepts — absolute path, `!command`, and a bare name looked up on
PATH — not just the one your host happens to use.
The helper resolves the identity in this order:
1. `$MOSAIC_GIT_IDENTITY`
2. `git config --get mosaic.gitIdentity`
3. the username git supplied on stdin
It maps the host to a store prefix — `git.mosaicstack.dev` to `gitea-mosaicstack`,
`git.uscllc.com` to `gitea-usc`. Any other host is declined quietly with rc=0, which is not an
error and raises no escalation.
Then it chooses **one** of two stores, and reads exactly one file:
```
brain_home = ${MOSAIC_BRAIN_HOME:-$HOME/.mosaic}
seat — when $brain_home/fleet/agents/<identity>/ EXISTS
$brain_home/fleet/agents/<identity>/secrets/<prefix>-<identity>.token
service — otherwise
~/.config/mosaic/secrets/gitea-tokens/<prefix>-<identity>.token
```
**There is no precedence between the two and no fallback from one to the other.** The existence
of the seat directory decides it. A seat that has a directory and an empty slot fails closed; it
does not reach the service store. That is the intended behavior — the alternative is an agent
silently acting as somebody else.
If the file is unreadable the helper **fails closed**: it refuses and writes a durable record to
the escalation spool. It does not fall back to a shared account. The record is what exists — any
alerting built on top of it is a separate, best-effort concern and is not performed by the helper,
so do not wait for a notification that nothing sends. That fallback is what made
`usc/uconnect#3084` unattributable, and it was removed deliberately.
Verify the helper you actually have:
```bash
h=$(git config --get credential.helper)
grep -c 'FAIL CLOSED' "$h" # expect >= 1
grep -c 'fleet/agents' "$h" # expect >= 1; 0 means it predates mosaicstack#1311
```
## Where a seat's token lives
The seat slot is the **only** copy:
```
~/.mosaic/fleet/agents/<seat>/secrets/<prefix>-<seat>.token real file, mode 600
```
The framework store at `~/.config/mosaic/secrets/gitea-tokens/` holds tokens for **service
identities only** — identities with no seat directory. A seat's token does not belong there.
Before mosaicstack#1311 the deployed helper knew only the service store, and seats were bridged
with a symlink from the store into the slot. **Those bridges must be removed once a seat-aware helper is deployed, and must not be
recreated.** Remove them only after the helper can reach the slot without them; the reverse order
takes every seat offline. A symlink
is not how a system finds a credential; the helper resolving the right store is.
`.principal` and `.scopes` beside the token are grant records, not secrets. They are tracked. The
`.token` never is.
### Provisioning a new seat
1. Create `~/.mosaic/fleet/agents/<seat>/secrets/` mode 700.
2. Write `.principal` (the Gitea login) and `.scopes` (the granted scopes), mode 600.
3. The estate operator mints the token into the seat slot, mode 600. Agents do not mint their
own, and do not ask another agent to mint one for them.
4. Verify with an authenticated `GET /user` and confirm the returned login is the seat, **not the
minting account**. Record the date in `ENTITY.md`. Never record the value.
There is no step that links the framework store to the slot. A seat-aware helper reads the slot
directly; a store entry pointing at a slot is the bridge described in **Where a seat's token lives** above,
and it is not part of provisioning.
Until step 3, the seat is unminted and its git writes fail closed. That is the designed state and
is safe to launch in — the seat is told at launch so it does not discover it mid-task.
## Acting as yourself
Name the identity on every invocation:
```bash
MOSAIC_GIT_IDENTITY=<seat> git push
git -c user.name=<seat> -c user.email=<seat>@mosaicstack.dev commit -m "..."
```
**Never persist `git config mosaic.gitIdentity` inside a `~/src/stack` worktree.** Every worktree
of that clone shares one `.git/config`, so a persisted identity there silently rewrites the
identity of every other seat working in that clone. The per-invocation form has no exception.
## Handling
1. **Never print a token value.** Compare by SHA-256 digest, or write `<REDACTED>`.
2. **Never stage a `.token`, `secrets.json`, or `ENTITY.md`.** Stage explicit paths and **never
`git add -A`** — `secrets/*.principal` and `secrets/*.scopes` are covered by no ignore rule.
3. **Never place a token in an environment variable** in an interactive session. A `declare -x`
dump has leaked the whole environment to a terminal before.
4. **No real credential or operator data on a sandbox VM, ever.**
@@ -11,22 +11,106 @@ All tool suites are located at `~/.config/mosaic/tools/`.
Mosaic wrappers at `~/.config/mosaic/tools/git/*.sh` handle platform detection and edge cases. Always use these before raw CLI commands.
This index is complete and is kept complete mechanically: `tools/quality/scripts/check-tools-index.sh`
fails CI when a wrapper ships without an entry here, or when an entry here names a wrapper that no
longer exists. A wrapper missing from this list is, from inside an agent session, indistinguishable
from a wrapper that was never written — which is how the APPROVE/APPROVED incident below happened.
Every command takes `--help`. All of them accept `--login <account>` to pin the acting identity;
supply it explicitly on any host where the provider CLI's default account is an admin.
| Issues | |
| ------------------ | --------------------------------- |
| `issue-create.sh` | Create an issue (Gitea or GitHub) |
| `issue-view.sh` | Show one issue |
| `issue-list.sh` | List issues |
| `issue-edit.sh` | Edit title/body/labels/milestone |
| `issue-comment.sh` | Add a comment |
| `issue-assign.sh` | Assign or unassign |
| `issue-close.sh` | Close an issue |
| `issue-reopen.sh` | Reopen a closed issue |
| Pull requests | |
| ---------------- | --------------------------------------------------------- |
| `pr-create.sh` | Open a pull request |
| `pr-edit.sh` | Edit PR title, body, base branch, or draft/ready state |
| `pr-view.sh` | Show one PR |
| `pr-list.sh` | List PRs |
| `pr-diff.sh` | Fetch a PR's diff |
| `pr-metadata.sh` | PR metadata as JSON (head SHA, base, state, mergeability) |
| `pr-review.sh` | **Place a review verdict — see the dialect note below** |
| `pr-ci-wait.sh` | Block until the PR's CI reaches a terminal state |
| `pr-merge.sh` | Merge a PR |
| `pr-close.sh` | Close a PR without merging |
| Milestones | |
| --------------------- | ------------------ |
| `milestone-create.sh` | Create a milestone |
| `milestone-list.sh` | List milestones |
| `milestone-close.sh` | Close a milestone |
| Gates and guards | |
| ----------------------- | --------------------------------------------------------------------------------------------------------- |
| `ci-queue-wait.sh` | CI queue guard — required before push/merge (see below) |
| `push-guard.sh` | Refuse verifications that pass for the wrong reason (e.g. green against an unpushed tree) |
| `mutate-push-guard.sh` | Regenerate the guard's mutation-coverage table from measurement, so the table cannot drift from the guard |
| `verify-clean-clone.sh` | Prove the **committed** artifact runs, from a clean clone — not the working tree |
| Context | |
| -------------------- | ---------------------------------------------------------------------------------------- |
| `detect-platform.sh` | Resolve the provider (Gitea vs GitHub) for the current repo; every other wrapper uses it |
| `lane-brief.sh` | Live dispatch brief for a repo "lane" (milestone/label) straight from the provider |
| Workspace | |
| -------------------- | ------------------------------------------------------------------------ |
| `mosaic-worktree.sh` | Create/list/remove git worktrees — **the only supported way**; see below |
| `wrapper-guard.sh` | PreToolUse hook that enforces the two rules above; not called by hand |
**Workspace placement is derived, not chosen.** `mosaic-worktree.sh new <branch>` takes a branch
name and nothing else. Every path comes out of `git worktree list --porcelain` — main worktree,
repo name, parent dir, then `<parent>/<repo>-worktrees/<branch-slug>`. There is no placement flag
because a decision an agent has to make is a decision that drifts: the rule "big work goes on a work
filesystem" already existed in prose and 255 GB accumulated in `$HOME` across 842 directories
anyway, under five simultaneous conventions on a single host.
```bash
# Issues
~/.config/mosaic/tools/git/issue-create.sh
~/.config/mosaic/tools/git/issue-close.sh
~/.config/mosaic/tools/git/mosaic-worktree.sh new <branch> [--from <base>]
~/.config/mosaic/tools/git/mosaic-worktree.sh path <branch> # derived path, no side effect
~/.config/mosaic/tools/git/mosaic-worktree.sh list # this repo's worktrees + state
~/.config/mosaic/tools/git/mosaic-worktree.sh rm <branch> # removal is part of the task
~/.config/mosaic/tools/git/mosaic-worktree.sh gc [--apply] # reclaim clean + fully-pushed ones
```
# PRs
~/.config/mosaic/tools/git/pr-create.sh
~/.config/mosaic/tools/git/pr-merge.sh
Worktrees rather than clones, because `git worktree list` makes every checkout enumerable — a bare
clone dropped somewhere on disk can never be safely reclaimed, so it is never reclaimed. `rm` and
`gc` decide by **evidence, never by size or age**: a worktree is reclaimable only when
`git status --porcelain` is empty _and_ `git rev-list --count HEAD --not --remotes` is 0. Anything
else is preserved and reported. `--force` exists and is yours to type deliberately.
# Milestones
~/.config/mosaic/tools/git/milestone-create.sh
`wrapper-guard.sh` is registered as a Claude Code `PreToolUse` hook on `Bash` (see
`runtime/claude/settings.json`). It blocks exactly three things and lets everything else through:
a `git clone`/`git worktree add` targeting `$HOME`; a raw provider-API **write** to an endpoint that
already has a wrapper above (reads are untouched — they are how you gather evidence); and the
literal `"event": "APPROVE"`. For a genuine gap no wrapper can express, prefix
`MOSAIC_WRAPPER_OVERRIDE=1`. Reaching for the override twice for the same call means the wrapper has
a missing flag — extend the wrapper.
```bash
~/.config/mosaic/tools/git/issue-create.sh --help
~/.config/mosaic/tools/git/pr-review.sh --pr 42 --event APPROVED --body "..."
# CI queue guard (required before push/merge; defaults to the checked-out branch)
~/.config/mosaic/tools/git/ci-queue-wait.sh --purpose push|merge
```
**Review dialect — the reason `pr-review.sh` is not optional.** Gitea's approve event is
`APPROVED`; GitHub's is `APPROVE`. Send GitHub's spelling to a Gitea host and it answers **HTTP
200**, files the review as PENDING, and then rejects the submit with `422 review stay pending` — the
verdict looks placed and is not. (`REQUEST_CHANGES` is spelled identically on both, so only the
approve path carries the trap.) `pr-review.sh` sends the correct token for the detected provider.
Whatever you use, re-read `GET /pulls/{n}/reviews` and assert the state before reporting a verdict
placed.
The guard exits nonzero for any provider-asserted non-green, missing, or malformed CI state. If credentials or the provider are unavailable, it emits `CANNOT_ASSERT` and writes a JSONL audit record. Push degrades to exit 0 so recovery work is not bricked; merge holds with retryable exit 75 until the provider recovers, then self-clears without manual reset. Neither outcome is evidence that CI was clear. `pr-merge.sh` automatically inspects the exact PR head repository and full commit SHA rather than its `main` base; this also handles fork PRs without branch-name ambiguity. Pass `--expect-head <approved-full-sha>` to bind a commit-specific review or merge-gate verdict; Gitea uses atomic `head_commit_id` and GitHub uses `--match-head-commit`.
### Code Review (Codex)
+2 -2
View File
@@ -17,7 +17,7 @@ set -Eeuo pipefail
# MOSAIC_HOME — target directory (default: ~/.config/mosaic)
# MOSAIC_INSTALL_MODE — prompt|keep|overwrite (default: prompt)
# MOSAIC_ALLOW_MISSING_SEQUENTIAL_THINKING — 1 to bypass MCP check
# MOSAIC_SKIP_SKILLS_SYNC — 1 to skip skill sync
# MOSAIC_SKIP_SKILLS_SYNC — 1 to skip linking skills into runtime homes
#
# Flags (CLI args, NOT environment variables — see #869 Point-1 C2):
# --allow-inactive-enforcement Explicit, per-invocation opt-out that lets the
@@ -828,7 +828,7 @@ if [[ -x "$SCRIPTS/mosaic-ensure-excalidraw" ]]; then
fi
if [[ "${MOSAIC_SKIP_SKILLS_SYNC:-0}" != "1" ]] && [[ -x "$SCRIPTS/mosaic-sync-skills" ]]; then
"$SCRIPTS/mosaic-sync-skills" >/dev/null 2>&1 && ok "Skills synced" || warn "Skills sync failed (non-fatal)"
"$SCRIPTS/mosaic-sync-skills" >/dev/null 2>&1 && ok "Skills linked into runtime homes" || warn "Skills linking failed (non-fatal)"
fi
if [[ -x "$SCRIPTS/mosaic-migrate-local-skills" ]]; then
@@ -64,6 +64,16 @@
"timeout": 10
}
]
},
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "~/.config/mosaic/tools/git/wrapper-guard.sh",
"timeout": 10
}
]
}
],
"PostToolUse": [
@@ -51,12 +51,26 @@ Skills are discovered from:
### Extensions
The Mosaic Pi extension (`~/.config/mosaic/runtime/pi/mosaic-extension.ts`) handles:
`mosaic pi` loads framework-owned extensions directly from `~/.config/mosaic/runtime/pi/` in this
order:
- Session start/end lifecycle hooks
- Active mission detection and context injection
- Memory routing to `~/.config/mosaic/memory/`
- MACP queue status reporting
1. `mosaic-extension.ts` — session lifecycle, mission context, memory routing, lease/mutator gates,
and fleet heartbeat reporting.
2. `goal-extension.ts` — optional persistent `/goal` controller with per-turn and post-compaction
checks.
The goal extension is deployed by Mosaic and MUST NOT be copied into `~/.pi/agent/extensions/`.
Use `/goal set <statement>` (or `/goal <statement>`) to start, then `/goal status`, `/goal pause`,
`/goal resume`, or `/goal cancel` to control it. An active goal is injected before every model
request, restored from branch-specific session entries, and considered achieved only after two
consecutive evidence-bearing reports. Common credential shapes are redacted before controller-owned
goal-state entries are persisted or
displayed; Pi's own model/tool-call history is separate. Goals and reports must contain references
and pass/fail summaries rather than secrets or raw sensitive output.
- `MOSAIC_GOAL_MAX_TURNS` — autonomous turn limit, default `40`, accepted range `1..500`.
- `MOSAIC_GOAL_MAX_NO_PROGRESS` — identical no-progress report limit, default `6`, accepted range
`1..100`.
### Sessions
File diff suppressed because it is too large Load Diff
+228
View File
@@ -0,0 +1,228 @@
# Agent Skills
Complete agent skill fleet for Mosaic Stack. 101 skills across 12 domains — coding, business development, design, marketing, writing, orchestration, document generation, Vue/Vite ecosystem, and more. Platform-aware — works with both GitHub (`gh`) and Gitea (`tea`) via our abstraction scripts.
This tree lives in the monorepo (`packages/mosaic/framework/skills/`) and ships inside the framework package; it is no longer a separate repository.
## Security Audit
All skills were reviewed on 2026-02-16. Findings:
| ID | Severity | Skill | Issue | Action |
| ----- | ------------- | ---------------------- | --------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| C-001 | **CRITICAL** | `vercel-deploy` | Uploads entire project to external endpoint via `curl` | **REMOVED** |
| C-002 | **ANNOTATED** | `docx`, `pptx`, `xlsx` | LD_PRELOAD shim compiles C at runtime to hook `socket()` | Security warnings added — legitimate sandbox workaround, should never activate on Docker Swarm |
| W-001 | WARNING | `using-superpowers` | Forces aggressive auto-loading via `<EXTREMELY-IMPORTANT>` tags | Awareness only — review before enabling |
| W-002 | WARNING | `mcp-builder` | Can connect to arbitrary MCP servers | Awareness only — review server URLs |
| W-003 | WARNING | `create-agent` | Uses `Function()` constructor (eval equivalent) | Awareness only — review generated code |
88 of 93 audited skills passed all checks as clean instruction-only SKILL.md files.
## Skills (95)
### Code Quality & Review (6)
| Skill | Purpose | Origin |
| -------------------------------- | ------------------------------------------------------------------------------- | ------------------------------- |
| `lint` | Zero-tolerance linting — detect linter, fix ALL violations, never disable rules | Mosaic Stack |
| `pr-reviewer` | Structured PR code review workflow (Gitea/GitHub) | Adapted from SpillwaveSolutions |
| `code-review-excellence` | Code review methodology and checklists | awesome-skills |
| `verification-before-completion` | Evidence-based completion claims | obra/superpowers |
| `receiving-code-review` | How to receive and respond to code reviews | obra/superpowers |
| `requesting-code-review` | How to request effective code reviews | obra/superpowers |
### Frontend & UI (8)
| Skill | Purpose | Origin |
| ----------------------------- | ----------------------------------------------------- | ----------------------- |
| `next-best-practices` | Next.js 15+ — RSC, async, self-hosting, data patterns | vercel-labs/next-skills |
| `vercel-react-best-practices` | React/Next.js performance (57 rules) | vercel-labs |
| `vercel-composition-patterns` | React composition and component patterns | vercel-labs |
| `vercel-react-native-skills` | React Native development patterns | vercel-labs |
| `shadcn-ui` | Component patterns — forms, dialogs, tables, charts | developer-kit |
| `tailwind-design-system` | Tailwind CSS v4 design system patterns | wshobson |
| `ui-animation` | Motion design — performance, accessibility, easing | mblode |
| `web-design-guidelines` | Web design principles and guidelines | vercel-labs |
### Backend & API (4)
| Skill | Purpose | Origin |
| --------------------------------- | ------------------------------------------------- | -------- |
| `nestjs-best-practices` | NestJS — 40 rules, 10 categories, priority-ranked | kadajett |
| `fastapi` | FastAPI + Pydantic v2 + async SQLAlchemy 2.0 | jezweb |
| `architecture-patterns` | Clean Architecture, Hexagonal, DDD | wshobson |
| `python-performance-optimization` | Profiling, memory, parallelization | wshobson |
### Authentication (5)
| Skill | Purpose | Origin |
| ------------------------------------------ | -------------------------------------------------- | ----------- |
| `better-auth-best-practices` | Better-Auth — Drizzle, sessions, plugins, security | better-auth |
| `create-auth-skill` | Creating custom Better-Auth skills | better-auth |
| `email-and-password-best-practices` | Email/password auth patterns | better-auth |
| `organization-best-practices` | Multi-org/team auth patterns | better-auth |
| `two-factor-authentication-best-practices` | 2FA implementation patterns | better-auth |
### AI & Agent Building (7)
| Skill | Purpose | Origin |
| ----------------------------- | --------------------------------------------------- | ---------------- |
| `ai-sdk` | Vercel AI SDK — streaming, multi-provider, agents | vercel/ai |
| `create-agent` | Modular agent with OpenRouter multi-model access | openrouterteam |
| `proactive-agent` | WAL Protocol, compaction recovery, self-improvement | halthelobster |
| `dispatching-parallel-agents` | Launching and managing parallel subagents | obra/superpowers |
| `subagent-driven-development` | Development workflow using subagents | obra/superpowers |
| `executing-plans` | Executing multi-step implementation plans | obra/superpowers |
| `using-superpowers` | Overview of the superpowers skill system | obra/superpowers |
### Development Workflow (6)
| Skill | Purpose | Origin |
| -------------------------------- | --------------------------------------- | ---------------- |
| `test-driven-development` | TDD Red-Green-Refactor discipline | obra/superpowers |
| `systematic-debugging` | Structured debugging methodology | obra/superpowers |
| `using-git-worktrees` | Git worktree patterns for parallel work | obra/superpowers |
| `finishing-a-development-branch` | Branch cleanup, squash, merge patterns | obra/superpowers |
| `writing-plans` | Writing effective implementation plans | obra/superpowers |
| `brainstorming` | Structured brainstorming methodology | obra/superpowers |
### Document Generation (6)
| Skill | Purpose | Origin |
| ----------------- | ---------------------------------- | ---------- |
| `pdf` | PDF document generation | anthropics |
| `docx` | Word document generation | anthropics |
| `pptx` | PowerPoint presentation generation | anthropics |
| `xlsx` | Excel spreadsheet generation | anthropics |
| `doc-coauthoring` | Collaborative document writing | anthropics |
| `internal-comms` | Internal communications drafting | anthropics |
### Design & Creative (7)
| Skill | Purpose | Origin |
| ----------------------- | --------------------------------------- | ---------- |
| `brand-guidelines` | Brand identity enforcement | anthropics |
| `frontend-design` | Frontend design patterns and principles | anthropics |
| `canvas-design` | Canvas/visual design patterns | anthropics |
| `algorithmic-art` | Generative/algorithmic art creation | anthropics |
| `theme-factory` | Theme generation and customization | anthropics |
| `slack-gif-creator` | Animated GIF creation for Slack | anthropics |
| `web-artifacts-builder` | Self-contained HTML artifact building | anthropics |
### Marketing & Business (25)
| Skill | Purpose | Origin |
| --------------------------- | --------------------------------------------- | ------------- |
| `marketing-ideas` | 139 ideas across 14 categories | coreyhaines31 |
| `pricing-strategy` | SaaS pricing — value metrics, tiers, research | coreyhaines31 |
| `programmatic-seo` | SEO at scale — templates, playbooks | coreyhaines31 |
| `competitor-alternatives` | Competitor comparison pages | coreyhaines31 |
| `referral-program` | Referral & affiliate programs | coreyhaines31 |
| `seo-audit` | Comprehensive SEO audit methodology | coreyhaines31 |
| `copywriting` | Marketing copywriting patterns | coreyhaines31 |
| `copy-editing` | Copy editing and proofreading | coreyhaines31 |
| `content-strategy` | Content strategy and planning | coreyhaines31 |
| `social-content` | Social media content creation | coreyhaines31 |
| `email-sequence` | Email sequence design and automation | coreyhaines31 |
| `launch-strategy` | Product launch planning | coreyhaines31 |
| `marketing-psychology` | Psychology-driven marketing | coreyhaines31 |
| `product-marketing-context` | Product marketing positioning | coreyhaines31 |
| `paid-ads` | Paid advertising campaigns | coreyhaines31 |
| `schema-markup` | Schema.org structured data | coreyhaines31 |
| `analytics-tracking` | Analytics setup and tracking | coreyhaines31 |
| `ab-test-setup` | A/B testing methodology | coreyhaines31 |
| `page-cro` | Landing page conversion optimization | coreyhaines31 |
| `form-cro` | Form conversion optimization | coreyhaines31 |
| `signup-flow-cro` | Signup flow conversion optimization | coreyhaines31 |
| `onboarding-cro` | User onboarding optimization | coreyhaines31 |
| `popup-cro` | Popup/modal conversion optimization | coreyhaines31 |
| `paywall-upgrade-cro` | Paywall/upgrade conversion optimization | coreyhaines31 |
| `free-tool-strategy` | Free tool as marketing strategy | coreyhaines31 |
### Vue/Vite Ecosystem (16)
| Skill | Purpose | Origin |
| ---------------------------- | ----------------------------------------- | ------ |
| `vue` | Vue.js development patterns | antfu |
| `vue-best-practices` | Vue.js best practices and conventions | antfu |
| `vue-router-best-practices` | Vue Router patterns and guards | antfu |
| `vue-testing-best-practices` | Vue component testing patterns | antfu |
| `vueuse-functions` | VueUse composable function patterns | antfu |
| `nuxt` | Nuxt.js framework patterns | antfu |
| `vite` | Vite build tool configuration and plugins | antfu |
| `vitest` | Vitest testing framework patterns | antfu |
| `vitepress` | VitePress documentation site patterns | antfu |
| `slidev` | Slidev presentation framework | antfu |
| `pnpm` | pnpm package manager patterns | antfu |
| `turborepo` | Turborepo monorepo patterns | antfu |
| `unocss` | UnoCSS atomic CSS engine | antfu |
| `tsdown` | tsdown TypeScript bundler | antfu |
| `pinia` | Pinia state management | antfu |
| `antfu` | Anthony Fu's coding conventions | antfu |
### Orchestration (1)
| Skill | Purpose | Origin |
| ----------- | ------------------------------------------------------------------------------------------ | ------------ |
| `kickstart` | Launch orchestrator for milestone/issue/task — auto-discovers context, bootstraps tracking | Mosaic Stack |
### Meta / Skill Authoring (4)
| Skill | Purpose | Origin |
| ---------------- | --------------------------------------------- | ---------------- |
| `writing-skills` | TDD-based skill authoring methodology | obra/superpowers |
| `skill-creator` | Anthropic's skill creation guide | anthropics |
| `mcp-builder` | Building MCP (Model Context Protocol) servers | anthropics |
| `webapp-testing` | Web application testing patterns | anthropics |
## Source Repositories
| Repository | Skills | Domain Focus |
| --------------------------------------------------------------------------------------------- | ------ | ---------------------------------------------- |
| [anthropics/skills](https://github.com/anthropics/skills) | 16 | Documents, design, MCP, testing |
| [obra/superpowers](https://github.com/obra/superpowers) | 14 | Agent workflows, TDD, code review, planning |
| [coreyhaines31/marketingskills](https://github.com/coreyhaines31/marketingskills) | 25 | Marketing, CRO, SEO, growth |
| [antfu/skills](https://github.com/antfu/skills) | 16 | Vue, Vite, Vitest, pnpm, Nuxt |
| [better-auth/skills](https://github.com/better-auth/skills) | 5 | Authentication patterns |
| [vercel-labs/agent-skills](https://github.com/vercel-labs/agent-skills) | 4 | React, design |
| [vercel-labs/next-skills](https://github.com/vercel-labs/next-skills) | 1 | Next.js 15+ |
| [vercel/ai](https://github.com/vercel/ai) | 1 | AI SDK |
| [halthelobster/proactive-agent](https://github.com/halthelobster/proactive-agent) | 1 | Agent architecture |
| [openrouterteam/agent-skills](https://github.com/openrouterteam/agent-skills) | 1 | Agent building |
| [kadajett/agent-nestjs-skills](https://github.com/kadajett/agent-nestjs-skills) | 1 | NestJS |
| [jezweb/claude-skills](https://github.com/jezweb/claude-skills) | 1 | FastAPI |
| [wshobson/agents](https://github.com/wshobson/agents) | 3 | Architecture, Python, Tailwind |
| [mblode/agent-skills](https://github.com/mblode/agent-skills) | 1 | UI animation |
| [giuseppe-trisciuoglio/developer-kit](https://github.com/giuseppe-trisciuoglio/developer-kit) | 1 | shadcn/ui |
| Mosaic Stack (original) | 4 | PR review, code review, orchestration, linting |
## Installation
The skills ship with the framework package. The framework installer installs them
into `~/.config/mosaic/skills/`, and the post-install step links them into each
runtime's skill directory:
```bash
# Install or upgrade the framework (skills arrive with it — no second repo)
./packages/mosaic/framework/install.sh
# Re-link installed skills into runtime homes (claude, codex, opencode, pi)
mosaic sync
```
Operators can override any canonical skill by copying it to
`~/.config/mosaic/skills-local/<name>/` — local skills take precedence during
linking.
## Adapting Skills
When adding skills from the community:
1. Replace raw `gh`/`tea` calls with our `~/.config/mosaic/rails/git/` scripts
2. Test on both GitHub and Gitea repos
3. Add Mosaic Stack context notes where upstream assumptions differ
4. Document any platform-specific limitations
## License
Individual skills retain their original licenses. Adaptations are MIT.
@@ -0,0 +1,287 @@
---
name: ab-test-setup
version: 1.0.0
description: When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," or "hypothesis." For tracking implementation, see analytics-tracking.
---
# A/B Test Setup
You are an expert in experimentation and A/B testing. Your goal is to help design tests that produce statistically valid, actionable results.
## Initial Assessment
**Check for product marketing context first:**
If `.mosaic/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before designing a test, understand:
1. **Test Context** - What are you trying to improve? What change are you considering?
2. **Current State** - Baseline conversion rate? Current traffic volume?
3. **Constraints** - Technical complexity? Timeline? Tools available?
---
## Core Principles
### 1. Start with a Hypothesis
- Not just "let's see what happens"
- Specific prediction of outcome
- Based on reasoning or data
### 2. Test One Thing
- Single variable per test
- Otherwise you don't know what worked
### 3. Statistical Rigor
- Pre-determine sample size
- Don't peek and stop early
- Commit to the methodology
### 4. Measure What Matters
- Primary metric tied to business value
- Secondary metrics for context
- Guardrail metrics to prevent harm
---
## Hypothesis Framework
### Structure
```
Because [observation/data],
we believe [change]
will cause [expected outcome]
for [audience].
We'll know this is true when [metrics].
```
### Example
**Weak**: "Changing the button color might increase clicks."
**Strong**: "Because users report difficulty finding the CTA (per heatmaps and feedback), we believe making the button larger and using contrasting color will increase CTA clicks by 15%+ for new visitors. We'll measure click-through rate from page view to signup start."
---
## Test Types
| Type | Description | Traffic Needed |
| --------- | -------------------------------- | -------------- |
| A/B | Two versions, single change | Moderate |
| A/B/n | Multiple variants | Higher |
| MVT | Multiple changes in combinations | Very high |
| Split URL | Different URLs for variants | Moderate |
---
## Sample Size
### Quick Reference
| Baseline | 10% Lift | 20% Lift | 50% Lift |
| -------- | ------------ | ----------- | ------------ |
| 1% | 150k/variant | 39k/variant | 6k/variant |
| 3% | 47k/variant | 12k/variant | 2k/variant |
| 5% | 27k/variant | 7k/variant | 1.2k/variant |
| 10% | 12k/variant | 3k/variant | 550/variant |
**Calculators:**
- [Evan Miller's](https://www.evanmiller.org/ab-testing/sample-size.html)
- [Optimizely's](https://www.optimizely.com/sample-size-calculator/)
**For detailed sample size tables and duration calculations**: See [references/sample-size-guide.md](references/sample-size-guide.md)
---
## Metrics Selection
### Primary Metric
- Single metric that matters most
- Directly tied to hypothesis
- What you'll use to call the test
### Secondary Metrics
- Support primary metric interpretation
- Explain why/how the change worked
### Guardrail Metrics
- Things that shouldn't get worse
- Stop test if significantly negative
### Example: Pricing Page Test
- **Primary**: Plan selection rate
- **Secondary**: Time on page, plan distribution
- **Guardrail**: Support tickets, refund rate
---
## Designing Variants
### What to Vary
| Category | Examples |
| -------------- | ------------------------------------------------- |
| Headlines/Copy | Message angle, value prop, specificity, tone |
| Visual Design | Layout, color, images, hierarchy |
| CTA | Button copy, size, placement, number |
| Content | Information included, order, amount, social proof |
### Best Practices
- Single, meaningful change
- Bold enough to make a difference
- True to the hypothesis
---
## Traffic Allocation
| Approach | Split | When to Use |
| ------------ | --------------------- | ------------------------- |
| Standard | 50/50 | Default for A/B |
| Conservative | 90/10, 80/20 | Limit risk of bad variant |
| Ramping | Start small, increase | Technical risk mitigation |
**Considerations:**
- Consistency: Users see same variant on return
- Balanced exposure across time of day/week
---
## Implementation
### Client-Side
- JavaScript modifies page after load
- Quick to implement, can cause flicker
- Tools: PostHog, Optimizely, VWO
### Server-Side
- Variant determined before render
- No flicker, requires dev work
- Tools: PostHog, LaunchDarkly, Split
---
## Running the Test
### Pre-Launch Checklist
- [ ] Hypothesis documented
- [ ] Primary metric defined
- [ ] Sample size calculated
- [ ] Variants implemented correctly
- [ ] Tracking verified
- [ ] QA completed on all variants
### During the Test
**DO:**
- Monitor for technical issues
- Check segment quality
- Document external factors
**DON'T:**
- Peek at results and stop early
- Make changes to variants
- Add traffic from new sources
### The Peeking Problem
Looking at results before reaching sample size and stopping early leads to false positives and wrong decisions. Pre-commit to sample size and trust the process.
---
## Analyzing Results
### Statistical Significance
- 95% confidence = p-value < 0.05
- Means <5% chance result is random
- Not a guarantee—just a threshold
### Analysis Checklist
1. **Reach sample size?** If not, result is preliminary
2. **Statistically significant?** Check confidence intervals
3. **Effect size meaningful?** Compare to MDE, project impact
4. **Secondary metrics consistent?** Support the primary?
5. **Guardrail concerns?** Anything get worse?
6. **Segment differences?** Mobile vs. desktop? New vs. returning?
### Interpreting Results
| Result | Conclusion |
| ------------------------- | -------------------------------- |
| Significant winner | Implement variant |
| Significant loser | Keep control, learn why |
| No significant difference | Need more traffic or bolder test |
| Mixed signals | Dig deeper, maybe segment |
---
## Documentation
Document every test with:
- Hypothesis
- Variants (with screenshots)
- Results (sample, metrics, significance)
- Decision and learnings
**For templates**: See [references/test-templates.md](references/test-templates.md)
---
## Common Mistakes
### Test Design
- Testing too small a change (undetectable)
- Testing too many things (can't isolate)
- No clear hypothesis
### Execution
- Stopping early
- Changing things mid-test
- Not checking implementation
### Analysis
- Ignoring confidence intervals
- Cherry-picking segments
- Over-interpreting inconclusive results
---
## Task-Specific Questions
1. What's your current conversion rate?
2. How much traffic does this page get?
3. What change are you considering and why?
4. What's the smallest improvement worth detecting?
5. What tools do you have for testing?
6. Have you tested this area before?
---
## Related Skills
- **page-cro**: For generating test ideas based on CRO principles
- **analytics-tracking**: For setting up test measurement
- **copywriting**: For creating variant copy
@@ -0,0 +1,272 @@
# Sample Size Guide
Reference for calculating sample sizes and test duration.
## Sample Size Fundamentals
### Required Inputs
1. **Baseline conversion rate**: Your current rate
2. **Minimum detectable effect (MDE)**: Smallest change worth detecting
3. **Statistical significance level**: Usually 95% (α = 0.05)
4. **Statistical power**: Usually 80% (β = 0.20)
### What These Mean
**Baseline conversion rate**: If your page converts at 5%, that's your baseline.
**MDE (Minimum Detectable Effect)**: The smallest improvement you care about detecting. Set this based on:
- Business impact (is a 5% lift meaningful?)
- Implementation cost (worth the effort?)
- Realistic expectations (what have past tests shown?)
**Statistical significance (95%)**: Means there's less than 5% chance the observed difference is due to random chance.
**Statistical power (80%)**: Means if there's a real effect of size MDE, you have 80% chance of detecting it.
---
## Sample Size Quick Reference Tables
### Conversion Rate: 1%
| Lift to Detect | Sample per Variant | Total Sample |
| --------------- | ------------------ | ------------ |
| 5% (1% → 1.05%) | 1,500,000 | 3,000,000 |
| 10% (1% → 1.1%) | 380,000 | 760,000 |
| 20% (1% → 1.2%) | 97,000 | 194,000 |
| 50% (1% → 1.5%) | 16,000 | 32,000 |
| 100% (1% → 2%) | 4,200 | 8,400 |
### Conversion Rate: 3%
| Lift to Detect | Sample per Variant | Total Sample |
| --------------- | ------------------ | ------------ |
| 5% (3% → 3.15%) | 480,000 | 960,000 |
| 10% (3% → 3.3%) | 120,000 | 240,000 |
| 20% (3% → 3.6%) | 31,000 | 62,000 |
| 50% (3% → 4.5%) | 5,200 | 10,400 |
| 100% (3% → 6%) | 1,400 | 2,800 |
### Conversion Rate: 5%
| Lift to Detect | Sample per Variant | Total Sample |
| --------------- | ------------------ | ------------ |
| 5% (5% → 5.25%) | 280,000 | 560,000 |
| 10% (5% → 5.5%) | 72,000 | 144,000 |
| 20% (5% → 6%) | 18,000 | 36,000 |
| 50% (5% → 7.5%) | 3,100 | 6,200 |
| 100% (5% → 10%) | 810 | 1,620 |
### Conversion Rate: 10%
| Lift to Detect | Sample per Variant | Total Sample |
| ---------------- | ------------------ | ------------ |
| 5% (10% → 10.5%) | 130,000 | 260,000 |
| 10% (10% → 11%) | 34,000 | 68,000 |
| 20% (10% → 12%) | 8,700 | 17,400 |
| 50% (10% → 15%) | 1,500 | 3,000 |
| 100% (10% → 20%) | 400 | 800 |
### Conversion Rate: 20%
| Lift to Detect | Sample per Variant | Total Sample |
| ---------------- | ------------------ | ------------ |
| 5% (20% → 21%) | 60,000 | 120,000 |
| 10% (20% → 22%) | 16,000 | 32,000 |
| 20% (20% → 24%) | 4,000 | 8,000 |
| 50% (20% → 30%) | 700 | 1,400 |
| 100% (20% → 40%) | 200 | 400 |
---
## Duration Calculator
### Formula
```
Duration (days) = (Sample per variant × Number of variants) / (Daily traffic × % exposed)
```
### Examples
**Scenario 1: High-traffic page**
- Need: 10,000 per variant (2 variants = 20,000 total)
- Daily traffic: 5,000 visitors
- 100% exposed to test
- Duration: 20,000 / 5,000 = **4 days**
**Scenario 2: Medium-traffic page**
- Need: 30,000 per variant (60,000 total)
- Daily traffic: 2,000 visitors
- 100% exposed
- Duration: 60,000 / 2,000 = **30 days**
**Scenario 3: Low-traffic with partial exposure**
- Need: 15,000 per variant (30,000 total)
- Daily traffic: 500 visitors
- 50% exposed to test
- Effective daily: 250
- Duration: 30,000 / 250 = **120 days** (too long!)
### Minimum Duration Rules
Even with sufficient sample size, run tests for at least:
- **1 full week**: To capture day-of-week variation
- **2 business cycles**: If B2B (weekday vs. weekend patterns)
- **Through paydays**: If e-commerce (beginning/end of month)
### Maximum Duration Guidelines
Avoid running tests longer than 4-8 weeks:
- Novelty effects wear off
- External factors intervene
- Opportunity cost of other tests
---
## Online Calculators
### Recommended Tools
**Evan Miller's Calculator**
https://www.evanmiller.org/ab-testing/sample-size.html
- Simple interface
- Bookmark-worthy
**Optimizely's Calculator**
https://www.optimizely.com/sample-size-calculator/
- Business-friendly language
- Duration estimates
**AB Test Guide Calculator**
https://www.abtestguide.com/calc/
- Includes Bayesian option
- Multiple test types
**VWO Duration Calculator**
https://vwo.com/tools/ab-test-duration-calculator/
- Duration-focused
- Good for planning
---
## Adjusting for Multiple Variants
With more than 2 variants (A/B/n tests), you need more sample:
| Variants | Multiplier |
| ----------- | -------------------------- |
| 2 (A/B) | 1x |
| 3 (A/B/C) | ~1.5x |
| 4 (A/B/C/D) | ~2x |
| 5+ | Consider reducing variants |
**Why?** More comparisons increase chance of false positives. You're comparing:
- A vs B
- A vs C
- B vs C (sometimes)
Apply Bonferroni correction or use tools that handle this automatically.
---
## Common Sample Size Mistakes
### 1. Underpowered tests
**Problem**: Not enough sample to detect realistic effects
**Fix**: Be realistic about MDE, get more traffic, or don't test
### 2. Overpowered tests
**Problem**: Waiting for sample size when you already have significance
**Fix**: This is actually fine—you committed to sample size, honor it
### 3. Wrong baseline rate
**Problem**: Using wrong conversion rate for calculation
**Fix**: Use the specific metric and page, not site-wide averages
### 4. Ignoring segments
**Problem**: Calculating for full traffic, then analyzing segments
**Fix**: If you plan segment analysis, calculate sample for smallest segment
### 5. Testing too many things
**Problem**: Dividing traffic too many ways
**Fix**: Prioritize ruthlessly, run fewer concurrent tests
---
## When Sample Size Requirements Are Too High
Options when you can't get enough traffic:
1. **Increase MDE**: Accept only detecting larger effects (20%+ lift)
2. **Lower confidence**: Use 90% instead of 95% (risky, document it)
3. **Reduce variants**: Test only the most promising variant
4. **Combine traffic**: Test across multiple similar pages
5. **Test upstream**: Test earlier in funnel where traffic is higher
6. **Don't test**: Make decision based on qualitative data instead
7. **Longer test**: Accept longer duration (weeks/months)
---
## Sequential Testing
If you must check results before reaching sample size:
### What is it?
Statistical method that adjusts for multiple looks at data.
### When to use
- High-risk changes
- Need to stop bad variants early
- Time-sensitive decisions
### Tools that support it
- Optimizely (Stats Accelerator)
- VWO (SmartStats)
- PostHog (Bayesian approach)
### Tradeoff
- More flexibility to stop early
- Slightly larger sample size requirement
- More complex analysis
---
## Quick Decision Framework
### Can I run this test?
```
Daily traffic to page: _____
Baseline conversion rate: _____
MDE I care about: _____
Sample needed per variant: _____ (from tables above)
Days to run: Sample / Daily traffic = _____
If days > 60: Consider alternatives
If days > 30: Acceptable for high-impact tests
If days < 14: Likely feasible
If days < 7: Easy to run, consider running longer anyway
```
@@ -0,0 +1,292 @@
# A/B Test Templates Reference
Templates for planning, documenting, and analyzing experiments.
## Test Plan Template
```markdown
# A/B Test: [Name]
## Overview
- **Owner**: [Name]
- **Test ID**: [ID in testing tool]
- **Page/Feature**: [What's being tested]
- **Planned dates**: [Start] - [End]
## Hypothesis
Because [observation/data],
we believe [change]
will cause [expected outcome]
for [audience].
We'll know this is true when [metrics].
## Test Design
| Element | Details |
| ------------------ | ------------------------- |
| Test type | A/B / A/B/n / MVT |
| Duration | X weeks |
| Sample size | X per variant |
| Traffic allocation | 50/50 |
| Tool | [Tool name] |
| Implementation | Client-side / Server-side |
## Variants
### Control (A)
[Screenshot]
- Current experience
- [Key details about current state]
### Variant (B)
[Screenshot or mockup]
- [Specific change #1]
- [Specific change #2]
- Rationale: [Why we think this will win]
## Metrics
### Primary
- **Metric**: [metric name]
- **Definition**: [how it's calculated]
- **Current baseline**: [X%]
- **Minimum detectable effect**: [X%]
### Secondary
- [Metric 1]: [what it tells us]
- [Metric 2]: [what it tells us]
- [Metric 3]: [what it tells us]
### Guardrails
- [Metric that shouldn't get worse]
- [Another safety metric]
## Segment Analysis Plan
- Mobile vs. desktop
- New vs. returning visitors
- Traffic source
- [Other relevant segments]
## Success Criteria
- Winner: [Primary metric improves by X% with 95% confidence]
- Loser: [Primary metric decreases significantly]
- Inconclusive: [What we'll do if no significant result]
## Pre-Launch Checklist
- [ ] Hypothesis documented and reviewed
- [ ] Primary metric defined and trackable
- [ ] Sample size calculated
- [ ] Test duration estimated
- [ ] Variants implemented correctly
- [ ] Tracking verified in all variants
- [ ] QA completed on all variants
- [ ] Stakeholders informed
- [ ] Calendar hold for analysis date
```
---
## Results Documentation Template
```markdown
# A/B Test Results: [Name]
## Summary
| Element | Value |
| -------- | ----------------------------- |
| Test ID | [ID] |
| Dates | [Start] - [End] |
| Duration | X days |
| Result | Winner / Loser / Inconclusive |
| Decision | [What we're doing] |
## Hypothesis (Reminder)
[Copy from test plan]
## Results
### Sample Size
| Variant | Target | Actual | % of target |
| ------- | ------ | ------ | ----------- |
| Control | X | Y | Z% |
| Variant | X | Y | Z% |
### Primary Metric: [Metric Name]
| Variant | Value | 95% CI | vs. Control |
| ------- | ----- | -------- | ----------- |
| Control | X% | [X%, Y%] | — |
| Variant | X% | [X%, Y%] | +X% |
**Statistical significance**: p = X.XX (95% = sig / not sig)
**Practical significance**: [Is this lift meaningful for the business?]
### Secondary Metrics
| Metric | Control | Variant | Change | Significant? |
| ---------- | ------- | ------- | ------ | ------------ |
| [Metric 1] | X | Y | +Z% | Yes/No |
| [Metric 2] | X | Y | +Z% | Yes/No |
### Guardrail Metrics
| Metric | Control | Variant | Change | Concern? |
| ---------- | ------- | ------- | ------ | -------- |
| [Metric 1] | X | Y | +Z% | Yes/No |
### Segment Analysis
**Mobile vs. Desktop**
| Segment | Control | Variant | Lift |
|---------|---------|---------|------|
| Mobile | X% | Y% | +Z% |
| Desktop | X% | Y% | +Z% |
**New vs. Returning**
| Segment | Control | Variant | Lift |
|---------|---------|---------|------|
| New | X% | Y% | +Z% |
| Returning | X% | Y% | +Z% |
## Interpretation
### What happened?
[Explanation of results in plain language]
### Why do we think this happened?
[Analysis and reasoning]
### Caveats
[Any limitations, external factors, or concerns]
## Decision
**Winner**: [Control / Variant]
**Action**: [Implement variant / Keep control / Re-test]
**Timeline**: [When changes will be implemented]
## Learnings
### What we learned
- [Key insight 1]
- [Key insight 2]
### What to test next
- [Follow-up test idea 1]
- [Follow-up test idea 2]
### Impact
- **Projected lift**: [X% improvement in Y metric]
- **Business impact**: [Revenue, conversions, etc.]
```
---
## Test Repository Entry Template
For tracking all tests in a central location:
```markdown
| Test ID | Name | Page | Dates | Primary Metric | Result | Lift | Link |
| ------- | -------------------- | -------- | --------- | -------------- | ------------ | ---- | ------ |
| 001 | Hero headline test | Homepage | 1/1-1/15 | CTR | Winner | +12% | [Link] |
| 002 | Pricing table layout | Pricing | 1/10-1/31 | Plan selection | Loser | -5% | [Link] |
| 003 | Signup form fields | Signup | 2/1-2/14 | Completion | Inconclusive | +2% | [Link] |
```
---
## Quick Test Brief Template
For simple tests that don't need full documentation:
```markdown
## [Test Name]
**What**: [One sentence description]
**Why**: [One sentence hypothesis]
**Metric**: [Primary metric]
**Duration**: [X weeks]
**Result**: [TBD / Winner / Loser / Inconclusive]
**Learnings**: [Key takeaway]
```
---
## Stakeholder Update Template
```markdown
## A/B Test Update: [Name]
**Status**: Running / Complete
**Days remaining**: X (or complete)
**Current sample**: X% of target
### Preliminary observations
[What we're seeing - without making decisions yet]
### Next steps
[What happens next]
### Timeline
- [Date]: Analysis complete
- [Date]: Decision and recommendation
- [Date]: Implementation 'if winner'
```
---
## Experiment Prioritization Scorecard
For deciding which tests to run:
| Factor | Weight | Test A | Test B | Test C |
| ------------------------ | ------ | ------ | ------ | ------ |
| Potential impact | 30% | | | |
| Confidence in hypothesis | 25% | | | |
| Ease of implementation | 20% | | | |
| Risk if wrong | 15% | | | |
| Strategic alignment | 10% | | | |
| **Total** | | | | |
Scoring: 1-5 (5 = best)
---
## Hypothesis Bank Template
For collecting test ideas:
```markdown
| ID | Page/Area | Observation | Hypothesis | Potential Impact | Status |
| --- | --------- | ------------------- | ------------------------------------- | ---------------- | ------- |
| H1 | Homepage | Low scroll depth | Shorter hero will increase scroll | High | Testing |
| H2 | Pricing | Users compare plans | Comparison table will help | Medium | Backlog |
| H3 | Signup | Drop-off at email | Social login will increase completion | Medium | Backlog |
```
@@ -0,0 +1,78 @@
---
name: ai-sdk
description: 'Answer questions about the AI SDK and help build AI-powered features. Use when developers: (1) Ask about AI SDK functions like generateText, streamText, ToolLoopAgent, embed, or tools, (2) Want to build AI agents, chatbots, RAG systems, or text generation features, (3) Have questions about AI providers (OpenAI, Anthropic, Google, etc.), streaming, tool calling, structured output, or embeddings, (4) Use React hooks like useChat or useCompletion. Triggers on: "AI SDK", "Vercel AI SDK", "generateText", "streamText", "add AI to my app", "build an agent", "tool calling", "structured output", "useChat".'
---
## Prerequisites
Before searching docs, check if `node_modules/ai/docs/` exists. If not, install **only** the `ai` package using the project's package manager (e.g., `pnpm add ai`).
Do not install other packages at this stage. Provider packages (e.g., `@ai-sdk/openai`) and client packages (e.g., `@ai-sdk/react`) should be installed later when needed based on user requirements.
## Critical: Do Not Trust Internal Knowledge
Everything you know about the AI SDK is outdated or wrong. Your training data contains obsolete APIs, deprecated patterns, and incorrect usage.
**When working with the AI SDK:**
1. Ensure `ai` package is installed (see Prerequisites)
2. Search `node_modules/ai/docs/` and `node_modules/ai/src/` for current APIs
3. If not found locally, search ai-sdk.dev documentation (instructions below)
4. Never rely on memory - always verify against source code or docs
5. **`useChat` has changed significantly** - check [Common Errors](references/common-errors.md) before writing client code
6. When deciding which model and provider to use (e.g. OpenAI, Anthropic, Gemini), use the Vercel AI Gateway provider unless the user specifies otherwise. See [AI Gateway Reference](references/ai-gateway.md) for usage details.
7. **Always fetch current model IDs** - Never use model IDs from memory. Before writing code that uses a model, run `curl -s https://ai-gateway.vercel.sh/v1/models | jq -r '[.data[] | select(.id | startswith("provider/")) | .id] | reverse | .[]'` (replacing `provider` with the relevant provider like `anthropic`, `openai`, or `google`) to get the full list with newest models first. Use the model with the highest version number (e.g., `claude-sonnet-4-5` over `claude-sonnet-4` over `claude-3-5-sonnet`).
8. Run typecheck after changes to ensure code is correct
9. **Be minimal** - Only specify options that differ from defaults. When unsure of defaults, check docs or source rather than guessing or over-specifying.
If you cannot find documentation to support your answer, state that explicitly.
## Finding Documentation
### ai@6.0.34+
Search bundled docs and source in `node_modules/ai/`:
- **Docs**: `grep "query" node_modules/ai/docs/`
- **Source**: `grep "query" node_modules/ai/src/`
Provider packages include docs at `node_modules/@ai-sdk/<provider>/docs/`.
### Earlier versions
1. Search: `https://ai-sdk.dev/api/search-docs?q=your_query`
2. Fetch `.md` URLs from results (e.g., `https://ai-sdk.dev/docs/agents/building-agents.md`)
## When Typecheck Fails
**Before searching source code**, grep [Common Errors](references/common-errors.md) for the failing property or function name. Many type errors are caused by deprecated APIs documented there.
If not found in common-errors.md:
1. Search `node_modules/ai/src/` and `node_modules/ai/docs/`
2. Search ai-sdk.dev (for earlier versions or if not found locally)
## Building and Consuming Agents
### Creating Agents
Always use the `ToolLoopAgent` pattern. Search `node_modules/ai/docs/` for current agent creation APIs.
**File conventions**: See [type-safe-agents.md](references/type-safe-agents.md) for where to save agents and tools.
**Type Safety**: When consuming agents with `useChat`, always use `InferAgentUIMessage<typeof agent>` for type-safe tool results. See [reference](references/type-safe-agents.md).
### Consuming Agents (Framework-Specific)
Before implementing agent consumption:
1. Check `package.json` to detect the project's framework/stack
2. Search documentation for the framework's quickstart guide
3. Follow the framework-specific patterns for streaming, API routes, and client integration
## References
- [Common Errors](references/common-errors.md) - Renamed parameters reference (parameters → inputSchema, etc.)
- [AI Gateway](references/ai-gateway.md) - Gateway setup and usage
- [Type-Safe Agents with useChat](references/type-safe-agents.md) - End-to-end type safety with InferAgentUIMessage
- [DevTools](references/devtools.md) - Set up local debugging and observability (development only)
@@ -0,0 +1,66 @@
---
title: Vercel AI Gateway
description: Reference for using Vercel AI Gateway with the AI SDK.
---
# Vercel AI Gateway
The Vercel AI Gateway is the fastest way to get started with the AI SDK. It provides access to models from OpenAI, Anthropic, Google, and other providers through a single API.
## Authentication
Authenticate with OIDC (for Vercel deployments) or an [AI Gateway API key](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai-gateway%2Fapi-keys&title=AI+Gateway+API+Keys):
```env filename=".env.local"
AI_GATEWAY_API_KEY=your_api_key_here
```
## Usage
The AI Gateway is the default global provider, so you can access models using a simple string:
```ts
import { generateText } from 'ai';
const { text } = await generateText({
model: 'anthropic/claude-sonnet-4.5',
prompt: 'What is love?',
});
```
You can also explicitly import and use the gateway provider:
```ts
// Option 1: Import from 'ai' package (included by default)
import { gateway } from 'ai';
model: gateway('anthropic/claude-sonnet-4.5');
// Option 2: Install and import from '@ai-sdk/gateway' package
import { gateway } from '@ai-sdk/gateway';
model: gateway('anthropic/claude-sonnet-4.5');
```
## Find Available Models
**Important**: Always fetch the current model list before writing code. Never use model IDs from memory - they may be outdated.
List all available models through the gateway API:
```bash
curl https://ai-gateway.vercel.sh/v1/models
```
Filter by provider using `jq`. **Do not truncate with `head`** - always fetch the full list to find the latest models:
```bash
# Anthropic models
curl -s https://ai-gateway.vercel.sh/v1/models | jq -r '[.data[] | select(.id | startswith("anthropic/")) | .id] | reverse | .[]'
# OpenAI models
curl -s https://ai-gateway.vercel.sh/v1/models | jq -r '[.data[] | select(.id | startswith("openai/")) | .id] | reverse | .[]'
# Google models
curl -s https://ai-gateway.vercel.sh/v1/models | jq -r '[.data[] | select(.id | startswith("google/")) | .id] | reverse | .[]'
```
When multiple versions of a model exist, use the one with the highest version number (e.g., prefer `claude-sonnet-4-5` over `claude-sonnet-4` over `claude-3-5-sonnet`).
@@ -0,0 +1,439 @@
---
title: Common Errors
description: Reference for common AI SDK errors and how to resolve them.
---
# Common Errors
## `maxTokens``maxOutputTokens`
```typescript
// ❌ Incorrect
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
maxTokens: 512, // deprecated: use `maxOutputTokens` instead
prompt: 'Write a short story',
});
// ✅ Correct
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
maxOutputTokens: 512,
prompt: 'Write a short story',
});
```
## `maxSteps``stopWhen: stepCountIs(n)`
```typescript
// ❌ Incorrect
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
tools: { weather },
maxSteps: 5, // deprecated: use `stopWhen: stepCountIs(n)` instead
prompt: 'What is the weather in NYC?',
});
// ✅ Correct
import { generateText, stepCountIs } from 'ai';
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
tools: { weather },
stopWhen: stepCountIs(5),
prompt: 'What is the weather in NYC?',
});
```
## `parameters``inputSchema` (in tool definition)
```typescript
// ❌ Incorrect
const weatherTool = tool({
description: 'Get weather for a location',
parameters: z.object({
// deprecated: use `inputSchema` instead
location: z.string(),
}),
execute: async ({ location }) => ({ location, temp: 72 }),
});
// ✅ Correct
const weatherTool = tool({
description: 'Get weather for a location',
inputSchema: z.object({
location: z.string(),
}),
execute: async ({ location }) => ({ location, temp: 72 }),
});
```
## `generateObject``generateText` with `output`
`generateObject` is deprecated. Use `generateText` with the `output` option instead.
```typescript
// ❌ Deprecated
import { generateObject } from 'ai'; // deprecated: use `generateText` with `output` instead
const result = await generateObject({
// deprecated function
model: 'anthropic/claude-opus-4.5',
schema: z.object({
// deprecated: use `Output.object({ schema })` instead
recipe: z.object({
name: z.string(),
ingredients: z.array(z.string()),
}),
}),
prompt: 'Generate a recipe for chocolate cake',
});
// ✅ Correct
import { generateText, Output } from 'ai';
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
output: Output.object({
schema: z.object({
recipe: z.object({
name: z.string(),
ingredients: z.array(z.string()),
}),
}),
}),
prompt: 'Generate a recipe for chocolate cake',
});
console.log(result.output); // typed object
```
## Manual JSON parsing → `generateText` with `output`
```typescript
// ❌ Incorrect
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
prompt: `Extract the user info as JSON: { "name": string, "age": number }
Input: John is 25 years old`,
});
const parsed = JSON.parse(result.text);
// ✅ Correct
import { generateText, Output } from 'ai';
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
output: Output.object({
schema: z.object({
name: z.string(),
age: z.number(),
}),
}),
prompt: 'Extract the user info: John is 25 years old',
});
console.log(result.output); // { name: 'John', age: 25 }
```
## Other `output` options
```typescript
// Output.array - for generating arrays of items
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
output: Output.array({
element: z.object({
city: z.string(),
country: z.string(),
}),
}),
prompt: 'List 5 capital cities',
});
// Output.choice - for selecting from predefined options
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
output: Output.choice({
options: ['positive', 'negative', 'neutral'] as const,
}),
prompt: 'Classify the sentiment: I love this product!',
});
// Output.json - for untyped JSON output
const result = await generateText({
model: 'anthropic/claude-opus-4.5',
output: Output.json(),
prompt: 'Return some JSON data',
});
```
## `toDataStreamResponse``toUIMessageStreamResponse`
When using `useChat` on the frontend, use `toUIMessageStreamResponse()` instead of `toDataStreamResponse()`. The UI message stream format is designed to work with the chat UI components and handles message state correctly.
```typescript
// ❌ Incorrect (when using useChat)
const result = streamText({
// config
});
return result.toDataStreamResponse(); // deprecated for useChat: use toUIMessageStreamResponse
// ✅ Correct
const result = streamText({
// config
});
return result.toUIMessageStreamResponse();
```
## Removed managed input state in `useChat`
The `useChat` hook no longer manages input state internally. You must now manage input state manually.
```tsx
// ❌ Deprecated
import { useChat } from '@ai-sdk/react';
export default function Page() {
const {
input, // deprecated: manage input state manually with useState
handleInputChange, // deprecated: use custom onChange handler
handleSubmit, // deprecated: use sendMessage() instead
} = useChat({
api: '/api/chat', // deprecated: use `transport: new DefaultChatTransport({ api })` instead
});
return (
<form onSubmit={handleSubmit}>
<input value={input} onChange={handleInputChange} />
<button type="submit">Send</button>
</form>
);
}
// ✅ Correct
import { useChat } from '@ai-sdk/react';
import { DefaultChatTransport } from 'ai';
import { useState } from 'react';
export default function Page() {
const [input, setInput] = useState('');
const { sendMessage } = useChat({
transport: new DefaultChatTransport({ api: '/api/chat' }),
});
const handleSubmit = (e) => {
e.preventDefault();
sendMessage({ text: input });
setInput('');
};
return (
<form onSubmit={handleSubmit}>
<input value={input} onChange={(e) => setInput(e.target.value)} />
<button type="submit">Send</button>
</form>
);
}
```
## `tool-invocation``tool-{toolName}` (typed tool parts)
When rendering messages with `useChat`, use the typed tool part names (`tool-{toolName}`) instead of the generic `tool-invocation` type. This provides better type safety and access to tool-specific input/output types.
> For end-to-end type-safety, see [Type-Safe Agents](type-safe-agents.md).
Typed tool parts also use different property names:
- `part.args``part.input`
- `part.result``part.output`
```tsx
// ❌ Incorrect - using generic tool-invocation
{
message.parts.map((part, i) => {
switch (part.type) {
case 'text':
return <div key={`${message.id}-${i}`}>{part.text}</div>;
case 'tool-invocation': // deprecated: use typed tool parts instead
return <pre key={`${message.id}-${i}`}>{JSON.stringify(part.toolInvocation, null, 2)}</pre>;
}
});
}
// ✅ Correct - using typed tool parts (recommended)
{
message.parts.map((part) => {
switch (part.type) {
case 'text':
return part.text;
case 'tool-askForConfirmation':
// handle askForConfirmation tool
break;
case 'tool-getWeatherInformation':
// handle getWeatherInformation tool
break;
}
});
}
// ✅ Alternative - using isToolUIPart as a catch-all
import { isToolUIPart } from 'ai';
{
message.parts.map((part) => {
if (part.type === 'text') {
return part.text;
}
if (isToolUIPart(part)) {
// handle any tool part generically
return (
<div key={part.toolCallId}>
{part.toolName}: {part.state}
</div>
);
}
});
}
```
## `useChat` state-dependent property access
Tool part properties are only available in certain states. TypeScript will error if you access them without checking state first.
```tsx
// ❌ Incorrect - input may be undefined during streaming
// TS18048: 'part.input' is possibly 'undefined'
if (part.type === 'tool-getWeather') {
const location = part.input.location;
}
// ✅ Correct - check for input-available or output-available
if (
part.type === 'tool-getWeather' &&
(part.state === 'input-available' || part.state === 'output-available')
) {
const location = part.input.location;
}
// ❌ Incorrect - output is only available after execution
// TS18048: 'part.output' is possibly 'undefined'
if (part.type === 'tool-getWeather') {
const weather = part.output;
}
// ✅ Correct - check for output-available
if (part.type === 'tool-getWeather' && part.state === 'output-available') {
const location = part.input.location;
const weather = part.output;
}
```
## `part.toolInvocation.args``part.input`
```tsx
// ❌ Incorrect
if (part.type === 'tool-invocation') {
// deprecated: use `part.input` on typed tool parts instead
const location = part.toolInvocation.args.location;
}
// ✅ Correct
if (
part.type === 'tool-getWeather' &&
(part.state === 'input-available' || part.state === 'output-available')
) {
const location = part.input.location;
}
```
## `part.toolInvocation.result``part.output`
```tsx
// ❌ Incorrect
if (part.type === 'tool-invocation') {
// deprecated: use `part.output` on typed tool parts instead
const weather = part.toolInvocation.result;
}
// ✅ Correct
if (part.type === 'tool-getWeather' && part.state === 'output-available') {
const weather = part.output;
}
```
## `part.toolInvocation.toolCallId``part.toolCallId`
```tsx
// ❌ Incorrect
if (part.type === 'tool-invocation') {
// deprecated: use `part.toolCallId` on typed tool parts instead
const id = part.toolInvocation.toolCallId;
}
// ✅ Correct
if (part.type === 'tool-getWeather') {
const id = part.toolCallId;
}
```
## Tool invocation states renamed
```tsx
// ❌ Incorrect
switch (part.toolInvocation.state) {
case 'partial-call': // deprecated: use `input-streaming` instead
return <div>Loading...</div>;
case 'call': // deprecated: use `input-available` instead
return <div>Executing...</div>;
case 'result': // deprecated: use `output-available` instead
return <div>Done</div>;
}
// ✅ Correct
switch (part.state) {
case 'input-streaming':
return <div>Loading...</div>;
case 'input-available':
return <div>Executing...</div>;
case 'output-available':
return <div>Done</div>;
}
```
## `addToolResult``addToolOutput`
```tsx
// ❌ Incorrect
addToolResult({
// deprecated: use `addToolOutput` instead
toolCallId: part.toolInvocation.toolCallId,
result: 'Yes, confirmed.', // deprecated: use `output` instead
});
// ✅ Correct
addToolOutput({
tool: 'askForConfirmation',
toolCallId: part.toolCallId,
output: 'Yes, confirmed.',
});
```
## `messages``uiMessages` in `createAgentUIStreamResponse`
```typescript
// ❌ Incorrect
return createAgentUIStreamResponse({
agent: myAgent,
messages, // incorrect: use `uiMessages` instead
});
// ✅ Correct
return createAgentUIStreamResponse({
agent: myAgent,
uiMessages: messages,
});
```
@@ -0,0 +1,52 @@
---
title: AI SDK DevTools
description: Debug AI SDK calls by inspecting captured runs and steps.
---
# AI SDK DevTools
## Why Use DevTools
DevTools captures all AI SDK calls (`generateText`, `streamText`, `ToolLoopAgent`) to a local JSON file. This lets you inspect LLM requests, responses, tool calls, and multi-step interactions without manually logging.
## Setup
Requires AI SDK 6. Install `@ai-sdk/devtools` using your project's package manager.
Wrap your model with the middleware:
```ts
import { wrapLanguageModel, gateway } from 'ai';
import { devToolsMiddleware } from '@ai-sdk/devtools';
const model = wrapLanguageModel({
model: gateway('anthropic/claude-sonnet-4.5'),
middleware: devToolsMiddleware(),
});
```
## Viewing Captured Data
All runs and steps are saved to:
```
.devtools/generations.json
```
Read this file directly to inspect captured data:
```bash
cat .devtools/generations.json | jq
```
Or launch the web UI:
```bash
npx @ai-sdk/devtools
# Open http://localhost:4983
```
## Data Structure
- **Run**: A complete multi-step interaction grouped by initial prompt
- **Step**: A single LLM call within a run (includes input, output, tool calls, token usage)
@@ -0,0 +1,200 @@
---
title: Type-Safe useChat with Agents
description: Build end-to-end type-safe agents by inferring UIMessage types from your agent definition.
---
# Type-Safe useChat with Agents
Build end-to-end type-safe agents by inferring `UIMessage` types from your agent definition for type-safe UI rendering with `useChat`.
## Recommended Structure
```
lib/
agents/
my-agent.ts # Agent definition + type export
tools/
weather-tool.ts # Individual tool definitions
calculator-tool.ts
```
## Define Tools
```ts
// lib/tools/weather-tool.ts
import { tool } from 'ai';
import { z } from 'zod';
export const weatherTool = tool({
description: 'Get current weather for a location',
inputSchema: z.object({
location: z.string().describe('City name'),
}),
execute: async ({ location }) => {
return { temperature: 72, condition: 'sunny', location };
},
});
```
## Define Agent and Export Type
```ts
// lib/agents/my-agent.ts
import { ToolLoopAgent, InferAgentUIMessage } from 'ai';
import { weatherTool } from '../tools/weather-tool';
import { calculatorTool } from '../tools/calculator-tool';
export const myAgent = new ToolLoopAgent({
model: 'anthropic/claude-sonnet-4',
instructions: 'You are a helpful assistant.',
tools: {
weather: weatherTool,
calculator: calculatorTool,
},
});
// Infer the UIMessage type from the agent
export type MyAgentUIMessage = InferAgentUIMessage<typeof myAgent>;
```
### With Custom Metadata
```ts
// lib/agents/my-agent.ts
import { z } from 'zod';
const metadataSchema = z.object({
createdAt: z.number(),
model: z.string().optional(),
});
type MyMetadata = z.infer<typeof metadataSchema>;
export type MyAgentUIMessage = InferAgentUIMessage<typeof myAgent, MyMetadata>;
```
## Use with `useChat`
```tsx
// app/chat.tsx
import { useChat } from '@ai-sdk/react';
import type { MyAgentUIMessage } from '@/lib/agents/my-agent';
export function Chat() {
const { messages } = useChat<MyAgentUIMessage>();
return (
<div>
{messages.map((message) => (
<Message key={message.id} message={message} />
))}
</div>
);
}
```
## Rendering Parts with Type Safety
Tool parts are typed as `tool-{toolName}` based on your agent's tools:
```tsx
function Message({ message }: { message: MyAgentUIMessage }) {
return (
<div>
{message.parts.map((part, i) => {
switch (part.type) {
case 'text':
return <p key={i}>{part.text}</p>;
case 'tool-weather':
// part.input and part.output are fully typed
if (part.state === 'output-available') {
return (
<div key={i}>
Weather in {part.input.location}: {part.output.temperature}F
</div>
);
}
return <div key={i}>Loading weather...</div>;
case 'tool-calculator':
// TypeScript knows this is the calculator tool
return <div key={i}>Calculating...</div>;
default:
return null;
}
})}
</div>
);
}
```
The `part.type` discriminant narrows the type, giving you autocomplete and type checking for `input` and `output` based on each tool's schema.
## Splitting Tool Rendering into Components
When rendering many tools, you may want to split each tool into its own component. Use `UIToolInvocation<TOOL>` to derive a typed invocation from your tool and export it alongside the tool definition:
```ts
// lib/tools/weather-tool.ts
import { tool, UIToolInvocation } from 'ai';
import { z } from 'zod';
export const weatherTool = tool({
description: 'Get current weather for a location',
inputSchema: z.object({
location: z.string().describe('City name'),
}),
execute: async ({ location }) => {
return { temperature: 72, condition: 'sunny', location };
},
});
// Export the invocation type for use in UI components
export type WeatherToolInvocation = UIToolInvocation<typeof weatherTool>;
```
Then import only the type in your component:
```tsx
// components/weather-tool.tsx
import type { WeatherToolInvocation } from '@/lib/tools/weather-tool';
export function WeatherToolComponent({ invocation }: { invocation: WeatherToolInvocation }) {
// invocation.input and invocation.output are fully typed
if (invocation.state === 'output-available') {
return (
<div>
Weather in {invocation.input.location}: {invocation.output.temperature}F
</div>
);
}
return <div>Loading weather for {invocation.input?.location}...</div>;
}
```
Use the component in your message renderer:
```tsx
function Message({ message }: { message: MyAgentUIMessage }) {
return (
<div>
{message.parts.map((part, i) => {
switch (part.type) {
case 'text':
return <p key={i}>{part.text}</p>;
case 'tool-weather':
return <WeatherToolComponent key={i} invocation={part} />;
case 'tool-calculator':
return <CalculatorToolComponent key={i} invocation={part} />;
default:
return null;
}
})}
</div>
);
}
```
This approach keeps your tool rendering logic organized while maintaining full type safety, without needing to import the tool implementation into your UI components.
@@ -0,0 +1,202 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [yyyy] [name of copyright owner]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
@@ -0,0 +1,443 @@
---
name: algorithmic-art
description: Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
license: Complete terms in LICENSE.txt
---
Algorithmic philosophies are computational aesthetic movements that are then expressed through code. Output .md files (philosophy), .html files (interactive viewer), and .js files (generative algorithms).
This happens in two steps:
1. Algorithmic Philosophy Creation (.md file)
2. Express by creating p5.js generative art (.html + .js files)
First, undertake this task:
## ALGORITHMIC PHILOSOPHY CREATION
To begin, create an ALGORITHMIC PHILOSOPHY (not static images or templates) that will be interpreted through:
- Computational processes, emergent behavior, mathematical beauty
- Seeded randomness, noise fields, organic systems
- Particles, flows, fields, forces
- Parametric variation and controlled chaos
### THE CRITICAL UNDERSTANDING
- What is received: Some subtle input or instructions by the user to take into account, but use as a foundation; it should not constrain creative freedom.
- What is created: An algorithmic philosophy/generative aesthetic movement.
- What happens next: The same version receives the philosophy and EXPRESSES IT IN CODE - creating p5.js sketches that are 90% algorithmic generation, 10% essential parameters.
Consider this approach:
- Write a manifesto for a generative art movement
- The next phase involves writing the algorithm that brings it to life
The philosophy must emphasize: Algorithmic expression. Emergent behavior. Computational beauty. Seeded variation.
### HOW TO GENERATE AN ALGORITHMIC PHILOSOPHY
**Name the movement** (1-2 words): "Organic Turbulence" / "Quantum Harmonics" / "Emergent Stillness"
**Articulate the philosophy** (4-6 paragraphs - concise but complete):
To capture the ALGORITHMIC essence, express how this philosophy manifests through:
- Computational processes and mathematical relationships?
- Noise functions and randomness patterns?
- Particle behaviors and field dynamics?
- Temporal evolution and system states?
- Parametric variation and emergent complexity?
**CRITICAL GUIDELINES:**
- **Avoid redundancy**: Each algorithmic aspect should be mentioned once. Avoid repeating concepts about noise theory, particle dynamics, or mathematical principles unless adding new depth.
- **Emphasize craftsmanship REPEATEDLY**: The philosophy MUST stress multiple times that the final algorithm should appear as though it took countless hours to develop, was refined with care, and comes from someone at the absolute top of their field. This framing is essential - repeat phrases like "meticulously crafted algorithm," "the product of deep computational expertise," "painstaking optimization," "master-level implementation."
- **Leave creative space**: Be specific about the algorithmic direction, but concise enough that the next Claude has room to make interpretive implementation choices at an extremely high level of craftsmanship.
The philosophy must guide the next version to express ideas ALGORITHMICALLY, not through static images. Beauty lives in the process, not the final frame.
### PHILOSOPHY EXAMPLES
**"Organic Turbulence"**
Philosophy: Chaos constrained by natural law, order emerging from disorder.
Algorithmic expression: Flow fields driven by layered Perlin noise. Thousands of particles following vector forces, their trails accumulating into organic density maps. Multiple noise octaves create turbulent regions and calm zones. Color emerges from velocity and density - fast particles burn bright, slow ones fade to shadow. The algorithm runs until equilibrium - a meticulously tuned balance where every parameter was refined through countless iterations by a master of computational aesthetics.
**"Quantum Harmonics"**
Philosophy: Discrete entities exhibiting wave-like interference patterns.
Algorithmic expression: Particles initialized on a grid, each carrying a phase value that evolves through sine waves. When particles are near, their phases interfere - constructive interference creates bright nodes, destructive creates voids. Simple harmonic motion generates complex emergent mandalas. The result of painstaking frequency calibration where every ratio was carefully chosen to produce resonant beauty.
**"Recursive Whispers"**
Philosophy: Self-similarity across scales, infinite depth in finite space.
Algorithmic expression: Branching structures that subdivide recursively. Each branch slightly randomized but constrained by golden ratios. L-systems or recursive subdivision generate tree-like forms that feel both mathematical and organic. Subtle noise perturbations break perfect symmetry. Line weights diminish with each recursion level. Every branching angle the product of deep mathematical exploration.
**"Field Dynamics"**
Philosophy: Invisible forces made visible through their effects on matter.
Algorithmic expression: Vector fields constructed from mathematical functions or noise. Particles born at edges, flowing along field lines, dying when they reach equilibrium or boundaries. Multiple fields can attract, repel, or rotate particles. The visualization shows only the traces - ghost-like evidence of invisible forces. A computational dance meticulously choreographed through force balance.
**"Stochastic Crystallization"**
Philosophy: Random processes crystallizing into ordered structures.
Algorithmic expression: Randomized circle packing or Voronoi tessellation. Start with random points, let them evolve through relaxation algorithms. Cells push apart until equilibrium. Color based on cell size, neighbor count, or distance from center. The organic tiling that emerges feels both random and inevitable. Every seed produces unique crystalline beauty - the mark of a master-level generative algorithm.
_These are condensed examples. The actual algorithmic philosophy should be 4-6 substantial paragraphs._
### ESSENTIAL PRINCIPLES
- **ALGORITHMIC PHILOSOPHY**: Creating a computational worldview to be expressed through code
- **PROCESS OVER PRODUCT**: Always emphasize that beauty emerges from the algorithm's execution - each run is unique
- **PARAMETRIC EXPRESSION**: Ideas communicate through mathematical relationships, forces, behaviors - not static composition
- **ARTISTIC FREEDOM**: The next Claude interprets the philosophy algorithmically - provide creative implementation room
- **PURE GENERATIVE ART**: This is about making LIVING ALGORITHMS, not static images with randomness
- **EXPERT CRAFTSMANSHIP**: Repeatedly emphasize the final algorithm must feel meticulously crafted, refined through countless iterations, the product of deep expertise by someone at the absolute top of their field in computational aesthetics
**The algorithmic philosophy should be 4-6 paragraphs long.** Fill it with poetic computational philosophy that brings together the intended vision. Avoid repeating the same points. Output this algorithmic philosophy as a .md file.
---
## DEDUCING THE CONCEPTUAL SEED
**CRITICAL STEP**: Before implementing the algorithm, identify the subtle conceptual thread from the original request.
**THE ESSENTIAL PRINCIPLE**:
The concept is a **subtle, niche reference embedded within the algorithm itself** - not always literal, always sophisticated. Someone familiar with the subject should feel it intuitively, while others simply experience a masterful generative composition. The algorithmic philosophy provides the computational language. The deduced concept provides the soul - the quiet conceptual DNA woven invisibly into parameters, behaviors, and emergence patterns.
This is **VERY IMPORTANT**: The reference must be so refined that it enhances the work's depth without announcing itself. Think like a jazz musician quoting another song through algorithmic harmony - only those who know will catch it, but everyone appreciates the generative beauty.
---
## P5.JS IMPLEMENTATION
With the philosophy AND conceptual framework established, express it through code. Pause to gather thoughts before proceeding. Use only the algorithmic philosophy created and the instructions below.
### ⚠️ STEP 0: READ THE TEMPLATE FIRST ⚠️
**CRITICAL: BEFORE writing any HTML:**
1. **Read** `templates/viewer.html` using the Read tool
2. **Study** the exact structure, styling, and Anthropic branding
3. **Use that file as the LITERAL STARTING POINT** - not just inspiration
4. **Keep all FIXED sections exactly as shown** (header, sidebar structure, Anthropic colors/fonts, seed controls, action buttons)
5. **Replace only the VARIABLE sections** marked in the file's comments (algorithm, parameters, UI controls for parameters)
**Avoid:**
- ❌ Creating HTML from scratch
- ❌ Inventing custom styling or color schemes
- ❌ Using system fonts or dark themes
- ❌ Changing the sidebar structure
**Follow these practices:**
- ✅ Copy the template's exact HTML structure
- ✅ Keep Anthropic branding (Poppins/Lora fonts, light colors, gradient backdrop)
- ✅ Maintain the sidebar layout (Seed → Parameters → Colors? → Actions)
- ✅ Replace only the p5.js algorithm and parameter controls
The template is the foundation. Build on it, don't rebuild it.
---
To create gallery-quality computational art that lives and breathes, use the algorithmic philosophy as the foundation.
### TECHNICAL REQUIREMENTS
**Seeded Randomness (Art Blocks Pattern)**:
```javascript
// ALWAYS use a seed for reproducibility
let seed = 12345; // or hash from user input
randomSeed(seed);
noiseSeed(seed);
```
**Parameter Structure - FOLLOW THE PHILOSOPHY**:
To establish parameters that emerge naturally from the algorithmic philosophy, consider: "What qualities of this system can be adjusted?"
```javascript
let params = {
seed: 12345, // Always include seed for reproducibility
// colors
// Add parameters that control YOUR algorithm:
// - Quantities (how many?)
// - Scales (how big? how fast?)
// - Probabilities (how likely?)
// - Ratios (what proportions?)
// - Angles (what direction?)
// - Thresholds (when does behavior change?)
};
```
**To design effective parameters, focus on the properties the system needs to be tunable rather than thinking in terms of "pattern types".**
**Core Algorithm - EXPRESS THE PHILOSOPHY**:
**CRITICAL**: The algorithmic philosophy should dictate what to build.
To express the philosophy through code, avoid thinking "which pattern should I use?" and instead think "how to express this philosophy through code?"
If the philosophy is about **organic emergence**, consider using:
- Elements that accumulate or grow over time
- Random processes constrained by natural rules
- Feedback loops and interactions
If the philosophy is about **mathematical beauty**, consider using:
- Geometric relationships and ratios
- Trigonometric functions and harmonics
- Precise calculations creating unexpected patterns
If the philosophy is about **controlled chaos**, consider using:
- Random variation within strict boundaries
- Bifurcation and phase transitions
- Order emerging from disorder
**The algorithm flows from the philosophy, not from a menu of options.**
To guide the implementation, let the conceptual essence inform creative and original choices. Build something that expresses the vision for this particular request.
**Canvas Setup**: Standard p5.js structure:
```javascript
function setup() {
createCanvas(1200, 1200);
// Initialize your system
}
function draw() {
// Your generative algorithm
// Can be static (noLoop) or animated
}
```
### CRAFTSMANSHIP REQUIREMENTS
**CRITICAL**: To achieve mastery, create algorithms that feel like they emerged through countless iterations by a master generative artist. Tune every parameter carefully. Ensure every pattern emerges with purpose. This is NOT random noise - this is CONTROLLED CHAOS refined through deep expertise.
- **Balance**: Complexity without visual noise, order without rigidity
- **Color Harmony**: Thoughtful palettes, not random RGB values
- **Composition**: Even in randomness, maintain visual hierarchy and flow
- **Performance**: Smooth execution, optimized for real-time if animated
- **Reproducibility**: Same seed ALWAYS produces identical output
### OUTPUT FORMAT
Output:
1. **Algorithmic Philosophy** - As markdown or text explaining the generative aesthetic
2. **Single HTML Artifact** - Self-contained interactive generative art built from `templates/viewer.html` (see STEP 0 and next section)
The HTML artifact contains everything: p5.js (from CDN), the algorithm, parameter controls, and UI - all in one file that works immediately in claude.ai artifacts or any browser. Start from the template file, not from scratch.
---
## INTERACTIVE ARTIFACT CREATION
**REMINDER: `templates/viewer.html` should have already been read (see STEP 0). Use that file as the starting point.**
To allow exploration of the generative art, create a single, self-contained HTML artifact. Ensure this artifact works immediately in claude.ai or any browser - no setup required. Embed everything inline.
### CRITICAL: WHAT'S FIXED VS VARIABLE
The `templates/viewer.html` file is the foundation. It contains the exact structure and styling needed.
**FIXED (always include exactly as shown):**
- Layout structure (header, sidebar, main canvas area)
- Anthropic branding (UI colors, fonts, gradients)
- Seed section in sidebar:
- Seed display
- Previous/Next buttons
- Random button
- Jump to seed input + Go button
- Actions section in sidebar:
- Regenerate button
- Reset button
**VARIABLE (customize for each artwork):**
- The entire p5.js algorithm (setup/draw/classes)
- The parameters object (define what the art needs)
- The Parameters section in sidebar:
- Number of parameter controls
- Parameter names
- Min/max/step values for sliders
- Control types (sliders, inputs, etc.)
- Colors section (optional):
- Some art needs color pickers
- Some art might use fixed colors
- Some art might be monochrome (no color controls needed)
- Decide based on the art's needs
**Every artwork should have unique parameters and algorithm!** The fixed parts provide consistent UX - everything else expresses the unique vision.
### REQUIRED FEATURES
**1. Parameter Controls**
- Sliders for numeric parameters (particle count, noise scale, speed, etc.)
- Color pickers for palette colors
- Real-time updates when parameters change
- Reset button to restore defaults
**2. Seed Navigation**
- Display current seed number
- "Previous" and "Next" buttons to cycle through seeds
- "Random" button for random seed
- Input field to jump to specific seed
- Generate 100 variations when requested (seeds 1-100)
**3. Single Artifact Structure**
```html
<!DOCTYPE html>
<html>
<head>
<!-- p5.js from CDN - always available -->
<script src="https://cdnjs.cloudflare.com/ajax/libs/p5.js/1.7.0/p5.min.js"></script>
<style>
/* All styling inline - clean, minimal */
/* Canvas on top, controls below */
</style>
</head>
<body>
<div id="canvas-container"></div>
<div id="controls">
<!-- All parameter controls -->
</div>
<script>
// ALL p5.js code inline here
// Parameter objects, classes, functions
// setup() and draw()
// UI handlers
// Everything self-contained
</script>
</body>
</html>
```
**CRITICAL**: This is a single artifact. No external files, no imports (except p5.js CDN). Everything inline.
**4. Implementation Details - BUILD THE SIDEBAR**
The sidebar structure:
**1. Seed (FIXED)** - Always include exactly as shown:
- Seed display
- Prev/Next/Random/Jump buttons
**2. Parameters (VARIABLE)** - Create controls for the art:
```html
<div class="control-group">
<label>Parameter Name</label>
<input
type="range"
id="param"
min="..."
max="..."
step="..."
value="..."
oninput="updateParam('param', this.value)"
/>
<span class="value-display" id="param-value">...</span>
</div>
```
Add as many control-group divs as there are parameters.
**3. Colors (OPTIONAL/VARIABLE)** - Include if the art needs adjustable colors:
- Add color pickers if users should control palette
- Skip this section if the art uses fixed colors
- Skip if the art is monochrome
**4. Actions (FIXED)** - Always include exactly as shown:
- Regenerate button
- Reset button
- Download PNG button
**Requirements**:
- Seed controls must work (prev/next/random/jump/display)
- All parameters must have UI controls
- Regenerate, Reset, Download buttons must work
- Keep Anthropic branding (UI styling, not art colors)
### USING THE ARTIFACT
The HTML artifact works immediately:
1. **In claude.ai**: Displayed as an interactive artifact - runs instantly
2. **As a file**: Save and open in any browser - no server needed
3. **Sharing**: Send the HTML file - it's completely self-contained
---
## VARIATIONS & EXPLORATION
The artifact includes seed navigation by default (prev/next/random buttons), allowing users to explore variations without creating multiple files. If the user wants specific variations highlighted:
- Include seed presets (buttons for "Variation 1: Seed 42", "Variation 2: Seed 127", etc.)
- Add a "Gallery Mode" that shows thumbnails of multiple seeds side-by-side
- All within the same single artifact
This is like creating a series of prints from the same plate - the algorithm is consistent, but each seed reveals different facets of its potential. The interactive nature means users discover their own favorites by exploring the seed space.
---
## THE CREATIVE PROCESS
**User request** → **Algorithmic philosophy** → **Implementation**
Each request is unique. The process involves:
1. **Interpret the user's intent** - What aesthetic is being sought?
2. **Create an algorithmic philosophy** (4-6 paragraphs) describing the computational approach
3. **Implement it in code** - Build the algorithm that expresses this philosophy
4. **Design appropriate parameters** - What should be tunable?
5. **Build matching UI controls** - Sliders/inputs for those parameters
**The constants**:
- Anthropic branding (colors, fonts, layout)
- Seed navigation (always present)
- Self-contained HTML artifact
**Everything else is variable**:
- The algorithm itself
- The parameters
- The UI controls
- The visual outcome
To achieve the best results, trust creativity and let the philosophy guide the implementation.
---
## RESOURCES
This skill includes helpful templates and documentation:
- **templates/viewer.html**: REQUIRED STARTING POINT for all HTML artifacts.
- This is the foundation - contains the exact structure and Anthropic branding
- **Keep unchanged**: Layout structure, sidebar organization, Anthropic colors/fonts, seed controls, action buttons
- **Replace**: The p5.js algorithm, parameter definitions, and UI controls in Parameters section
- The extensive comments in the file mark exactly what to keep vs replace
- **templates/generator_template.js**: Reference for p5.js best practices and code structure principles.
- Shows how to organize parameters, use seeded randomness, structure classes
- NOT a pattern menu - use these principles to build unique algorithms
- Embed algorithms inline in the HTML artifact (don't create separate .js files)
**Critical reminder**:
- The **template is the STARTING POINT**, not inspiration
- The **algorithm is where to create** something unique
- Don't copy the flow field example - build what the philosophy demands
- But DO keep the exact UI structure and Anthropic branding from the template
@@ -0,0 +1,223 @@
/**
*
* P5.JS GENERATIVE ART - BEST PRACTICES
*
*
* This file shows STRUCTURE and PRINCIPLES for p5.js generative art.
* It does NOT prescribe what art you should create.
*
* Your algorithmic philosophy should guide what you build.
* These are just best practices for how to structure your code.
*
*
*/
// ============================================================================
// 1. PARAMETER ORGANIZATION
// ============================================================================
// Keep all tunable parameters in one object
// This makes it easy to:
// - Connect to UI controls
// - Reset to defaults
// - Serialize/save configurations
let params = {
// Define parameters that match YOUR algorithm
// Examples (customize for your art):
// - Counts: how many elements (particles, circles, branches, etc.)
// - Scales: size, speed, spacing
// - Probabilities: likelihood of events
// - Angles: rotation, direction
// - Colors: palette arrays
seed: 12345,
// define colorPalette as an array -- choose whatever colors you'd like ['#d97757', '#6a9bcc', '#788c5d', '#b0aea5']
// Add YOUR parameters here based on your algorithm
};
// ============================================================================
// 2. SEEDED RANDOMNESS (Critical for reproducibility)
// ============================================================================
// ALWAYS use seeded random for Art Blocks-style reproducible output
function initializeSeed(seed) {
randomSeed(seed);
noiseSeed(seed);
// Now all random() and noise() calls will be deterministic
}
// ============================================================================
// 3. P5.JS LIFECYCLE
// ============================================================================
function setup() {
createCanvas(800, 800);
// Initialize seed first
initializeSeed(params.seed);
// Set up your generative system
// This is where you initialize:
// - Arrays of objects
// - Grid structures
// - Initial positions
// - Starting states
// For static art: call noLoop() at the end of setup
// For animated art: let draw() keep running
}
function draw() {
// Option 1: Static generation (runs once, then stops)
// - Generate everything in setup()
// - Call noLoop() in setup()
// - draw() doesn't do much or can be empty
// Option 2: Animated generation (continuous)
// - Update your system each frame
// - Common patterns: particle movement, growth, evolution
// - Can optionally call noLoop() after N frames
// Option 3: User-triggered regeneration
// - Use noLoop() by default
// - Call redraw() when parameters change
}
// ============================================================================
// 4. CLASS STRUCTURE (When you need objects)
// ============================================================================
// Use classes when your algorithm involves multiple entities
// Examples: particles, agents, cells, nodes, etc.
class Entity {
constructor() {
// Initialize entity properties
// Use random() here - it will be seeded
}
update() {
// Update entity state
// This might involve:
// - Physics calculations
// - Behavioral rules
// - Interactions with neighbors
}
display() {
// Render the entity
// Keep rendering logic separate from update logic
}
}
// ============================================================================
// 5. PERFORMANCE CONSIDERATIONS
// ============================================================================
// For large numbers of elements:
// - Pre-calculate what you can
// - Use simple collision detection (spatial hashing if needed)
// - Limit expensive operations (sqrt, trig) when possible
// - Consider using p5 vectors efficiently
// For smooth animation:
// - Aim for 60fps
// - Profile if things are slow
// - Consider reducing particle counts or simplifying calculations
// ============================================================================
// 6. UTILITY FUNCTIONS
// ============================================================================
// Color utilities
function hexToRgb(hex) {
const result = /^#?([a-f\d]{2})([a-f\d]{2})([a-f\d]{2})$/i.exec(hex);
return result
? {
r: parseInt(result[1], 16),
g: parseInt(result[2], 16),
b: parseInt(result[3], 16),
}
: null;
}
function colorFromPalette(index) {
return params.colorPalette[index % params.colorPalette.length];
}
// Mapping and easing
function mapRange(value, inMin, inMax, outMin, outMax) {
return outMin + (outMax - outMin) * ((value - inMin) / (inMax - inMin));
}
function easeInOutCubic(t) {
return t < 0.5 ? 4 * t * t * t : 1 - Math.pow(-2 * t + 2, 3) / 2;
}
// Constrain to bounds
function wrapAround(value, max) {
if (value < 0) return max;
if (value > max) return 0;
return value;
}
// ============================================================================
// 7. PARAMETER UPDATES (Connect to UI)
// ============================================================================
function updateParameter(paramName, value) {
params[paramName] = value;
// Decide if you need to regenerate or just update
// Some params can update in real-time, others need full regeneration
}
function regenerate() {
// Reinitialize your generative system
// Useful when parameters change significantly
initializeSeed(params.seed);
// Then regenerate your system
}
// ============================================================================
// 8. COMMON P5.JS PATTERNS
// ============================================================================
// Drawing with transparency for trails/fading
function fadeBackground(opacity) {
fill(250, 249, 245, opacity); // Anthropic light with alpha
noStroke();
rect(0, 0, width, height);
}
// Using noise for organic variation
function getNoiseValue(x, y, scale = 0.01) {
return noise(x * scale, y * scale);
}
// Creating vectors from angles
function vectorFromAngle(angle, magnitude = 1) {
return createVector(cos(angle), sin(angle)).mult(magnitude);
}
// ============================================================================
// 9. EXPORT FUNCTIONS
// ============================================================================
function exportImage() {
saveCanvas('generative-art-' + params.seed, 'png');
}
// ============================================================================
// REMEMBER
// ============================================================================
//
// These are TOOLS and PRINCIPLES, not a recipe.
// Your algorithmic philosophy should guide WHAT you create.
// This structure helps you create it WELL.
//
// Focus on:
// - Clean, readable code
// - Parameterized for exploration
// - Seeded for reproducibility
// - Performant execution
//
// The art itself is entirely up to you!
//
// ============================================================================
@@ -0,0 +1,599 @@
<!DOCTYPE html>
<!--
THIS IS A TEMPLATE THAT SHOULD BE USED EVERY TIME AND MODIFIED.
WHAT TO KEEP:
✓ Overall structure (header, sidebar, main content)
✓ Anthropic branding (colors, fonts, layout)
✓ Seed navigation section (always include this)
✓ Self-contained artifact (everything inline)
WHAT TO CREATIVELY EDIT:
✗ The p5.js algorithm (implement YOUR vision)
✗ The parameters (define what YOUR art needs)
✗ The UI controls (match YOUR parameters)
Let your philosophy guide the implementation.
The world is your oyster - be creative!
-->
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Generative Art Viewer</title>
<script src="https://cdnjs.cloudflare.com/ajax/libs/p5.js/1.7.0/p5.min.js"></script>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Poppins:wght@400;500;600&family=Lora:wght@400;500&display=swap" rel="stylesheet">
<style>
/* Anthropic Brand Colors */
:root {
--anthropic-dark: #141413;
--anthropic-light: #faf9f5;
--anthropic-mid-gray: #b0aea5;
--anthropic-light-gray: #e8e6dc;
--anthropic-orange: #d97757;
--anthropic-blue: #6a9bcc;
--anthropic-green: #788c5d;
}
* {
margin: 0;
padding: 0;
box-sizing: border-box;
}
body {
font-family: 'Poppins', sans-serif;
background: linear-gradient(135deg, var(--anthropic-light) 0%, #f5f3ee 100%);
min-height: 100vh;
color: var(--anthropic-dark);
}
.container {
display: flex;
min-height: 100vh;
padding: 20px;
gap: 20px;
}
/* Sidebar */
.sidebar {
width: 320px;
flex-shrink: 0;
background: rgba(255, 255, 255, 0.95);
backdrop-filter: blur(10px);
padding: 24px;
border-radius: 12px;
box-shadow: 0 10px 30px rgba(20, 20, 19, 0.1);
overflow-y: auto;
overflow-x: hidden;
}
.sidebar h1 {
font-family: 'Lora', serif;
font-size: 24px;
font-weight: 500;
color: var(--anthropic-dark);
margin-bottom: 8px;
}
.sidebar .subtitle {
color: var(--anthropic-mid-gray);
font-size: 14px;
margin-bottom: 32px;
line-height: 1.4;
}
/* Control Sections */
.control-section {
margin-bottom: 32px;
}
.control-section h3 {
font-size: 16px;
font-weight: 600;
color: var(--anthropic-dark);
margin-bottom: 16px;
display: flex;
align-items: center;
gap: 8px;
}
.control-section h3::before {
content: '•';
color: var(--anthropic-orange);
font-weight: bold;
}
/* Seed Controls */
.seed-input {
width: 100%;
background: var(--anthropic-light);
padding: 12px;
border-radius: 8px;
font-family: 'Courier New', monospace;
font-size: 14px;
margin-bottom: 12px;
border: 1px solid var(--anthropic-light-gray);
text-align: center;
}
.seed-input:focus {
outline: none;
border-color: var(--anthropic-orange);
box-shadow: 0 0 0 2px rgba(217, 119, 87, 0.1);
background: white;
}
.seed-controls {
display: grid;
grid-template-columns: 1fr 1fr;
gap: 8px;
margin-bottom: 8px;
}
.regen-button {
margin-bottom: 0;
}
/* Parameter Controls */
.control-group {
margin-bottom: 20px;
}
.control-group label {
display: block;
font-size: 14px;
font-weight: 500;
color: var(--anthropic-dark);
margin-bottom: 8px;
}
.slider-container {
display: flex;
align-items: center;
gap: 12px;
}
.slider-container input[type="range"] {
flex: 1;
height: 4px;
background: var(--anthropic-light-gray);
border-radius: 2px;
outline: none;
-webkit-appearance: none;
}
.slider-container input[type="range"]::-webkit-slider-thumb {
-webkit-appearance: none;
width: 16px;
height: 16px;
background: var(--anthropic-orange);
border-radius: 50%;
cursor: pointer;
transition: all 0.2s ease;
}
.slider-container input[type="range"]::-webkit-slider-thumb:hover {
transform: scale(1.1);
background: #c86641;
}
.slider-container input[type="range"]::-moz-range-thumb {
width: 16px;
height: 16px;
background: var(--anthropic-orange);
border-radius: 50%;
border: none;
cursor: pointer;
transition: all 0.2s ease;
}
.value-display {
font-family: 'Courier New', monospace;
font-size: 12px;
color: var(--anthropic-mid-gray);
min-width: 60px;
text-align: right;
}
/* Color Pickers */
.color-group {
margin-bottom: 16px;
}
.color-group label {
display: block;
font-size: 12px;
color: var(--anthropic-mid-gray);
margin-bottom: 4px;
}
.color-picker-container {
display: flex;
align-items: center;
gap: 8px;
}
.color-picker-container input[type="color"] {
width: 32px;
height: 32px;
border: none;
border-radius: 6px;
cursor: pointer;
background: none;
padding: 0;
}
.color-value {
font-family: 'Courier New', monospace;
font-size: 12px;
color: var(--anthropic-mid-gray);
}
/* Buttons */
.button {
background: var(--anthropic-orange);
color: white;
border: none;
padding: 10px 16px;
border-radius: 6px;
font-size: 14px;
font-weight: 500;
cursor: pointer;
transition: all 0.2s ease;
width: 100%;
}
.button:hover {
background: #c86641;
transform: translateY(-1px);
}
.button:active {
transform: translateY(0);
}
.button.secondary {
background: var(--anthropic-blue);
}
.button.secondary:hover {
background: #5a8bb8;
}
.button.tertiary {
background: var(--anthropic-green);
}
.button.tertiary:hover {
background: #6b7b52;
}
.button-row {
display: flex;
gap: 8px;
}
.button-row .button {
flex: 1;
}
/* Canvas Area */
.canvas-area {
flex: 1;
display: flex;
align-items: center;
justify-content: center;
min-width: 0;
}
#canvas-container {
width: 100%;
max-width: 1000px;
border-radius: 12px;
overflow: hidden;
box-shadow: 0 20px 40px rgba(20, 20, 19, 0.1);
background: white;
}
#canvas-container canvas {
display: block;
width: 100% !important;
height: auto !important;
}
/* Loading State */
.loading {
display: flex;
align-items: center;
justify-content: center;
font-size: 18px;
color: var(--anthropic-mid-gray);
}
/* Responsive - Stack on mobile */
@media (max-width: 600px) {
.container {
flex-direction: column;
}
.sidebar {
width: 100%;
}
.canvas-area {
padding: 20px;
}
}
</style>
</head>
<body>
<div class="container">
<!-- Control Sidebar -->
<div class="sidebar">
<!-- Headers (CUSTOMIZE THIS FOR YOUR ART) -->
<h1>TITLE - EDIT</h1>
<div class="subtitle">SUBHEADER - EDIT</div>
<!-- Seed Section (ALWAYS KEEP THIS) -->
<div class="control-section">
<h3>Seed</h3>
<input type="number" id="seed-input" class="seed-input" value="12345" onchange="updateSeed()">
<div class="seed-controls">
<button class="button secondary" onclick="previousSeed()">← Prev</button>
<button class="button secondary" onclick="nextSeed()">Next →</button>
</div>
<button class="button tertiary regen-button" onclick="randomSeedAndUpdate()">↻ Random</button>
</div>
<!-- Parameters Section (CUSTOMIZE THIS FOR YOUR ART) -->
<div class="control-section">
<h3>Parameters</h3>
<!-- Particle Count -->
<div class="control-group">
<label>Particle Count</label>
<div class="slider-container">
<input type="range" id="particleCount" min="1000" max="10000" step="500" value="5000" oninput="updateParam('particleCount', this.value)">
<span class="value-display" id="particleCount-value">5000</span>
</div>
</div>
<!-- Flow Speed -->
<div class="control-group">
<label>Flow Speed</label>
<div class="slider-container">
<input type="range" id="flowSpeed" min="0.1" max="2.0" step="0.1" value="0.5" oninput="updateParam('flowSpeed', this.value)">
<span class="value-display" id="flowSpeed-value">0.5</span>
</div>
</div>
<!-- Noise Scale -->
<div class="control-group">
<label>Noise Scale</label>
<div class="slider-container">
<input type="range" id="noiseScale" min="0.001" max="0.02" step="0.001" value="0.005" oninput="updateParam('noiseScale', this.value)">
<span class="value-display" id="noiseScale-value">0.005</span>
</div>
</div>
<!-- Trail Length -->
<div class="control-group">
<label>Trail Length</label>
<div class="slider-container">
<input type="range" id="trailLength" min="2" max="20" step="1" value="8" oninput="updateParam('trailLength', this.value)">
<span class="value-display" id="trailLength-value">8</span>
</div>
</div>
</div>
<!-- Colors Section (OPTIONAL - CUSTOMIZE OR REMOVE) -->
<div class="control-section">
<h3>Colors</h3>
<!-- Color 1 -->
<div class="color-group">
<label>Primary Color</label>
<div class="color-picker-container">
<input type="color" id="color1" value="#d97757" onchange="updateColor('color1', this.value)">
<span class="color-value" id="color1-value">#d97757</span>
</div>
</div>
<!-- Color 2 -->
<div class="color-group">
<label>Secondary Color</label>
<div class="color-picker-container">
<input type="color" id="color2" value="#6a9bcc" onchange="updateColor('color2', this.value)">
<span class="color-value" id="color2-value">#6a9bcc</span>
</div>
</div>
<!-- Color 3 -->
<div class="color-group">
<label>Accent Color</label>
<div class="color-picker-container">
<input type="color" id="color3" value="#788c5d" onchange="updateColor('color3', this.value)">
<span class="color-value" id="color3-value">#788c5d</span>
</div>
</div>
</div>
<!-- Actions Section (ALWAYS KEEP THIS) -->
<div class="control-section">
<h3>Actions</h3>
<div class="button-row">
<button class="button" onclick="resetParameters()">Reset</button>
</div>
</div>
</div>
<!-- Main Canvas Area -->
<div class="canvas-area">
<div id="canvas-container">
<div class="loading">Initializing generative art...</div>
</div>
</div>
</div>
<script>
// ═══════════════════════════════════════════════════════════════════════
// GENERATIVE ART PARAMETERS - CUSTOMIZE FOR YOUR ALGORITHM
// ═══════════════════════════════════════════════════════════════════════
let params = {
seed: 12345,
particleCount: 5000,
flowSpeed: 0.5,
noiseScale: 0.005,
trailLength: 8,
colorPalette: ['#d97757', '#6a9bcc', '#788c5d']
};
let defaultParams = {...params}; // Store defaults for reset
// ═══════════════════════════════════════════════════════════════════════
// P5.JS GENERATIVE ART ALGORITHM - REPLACE WITH YOUR VISION
// ═══════════════════════════════════════════════════════════════════════
let particles = [];
let flowField = [];
let cols, rows;
let scl = 10; // Flow field resolution
function setup() {
let canvas = createCanvas(1200, 1200);
canvas.parent('canvas-container');
initializeSystem();
// Remove loading message
document.querySelector('.loading').style.display = 'none';
}
function initializeSystem() {
// Seed the randomness for reproducibility
randomSeed(params.seed);
noiseSeed(params.seed);
// Clear particles and recreate
particles = [];
// Initialize particles
for (let i = 0; i < params.particleCount; i++) {
particles.push(new Particle());
}
// Calculate flow field dimensions
cols = floor(width / scl);
rows = floor(height / scl);
// Generate flow field
generateFlowField();
// Clear background
background(250, 249, 245); // Anthropic light background
}
function generateFlowField() {
// fill this in
}
function draw() {
// fill this in
}
// ═══════════════════════════════════════════════════════════════════════
// PARTICLE SYSTEM - CUSTOMIZE FOR YOUR ALGORITHM
// ═══════════════════════════════════════════════════════════════════════
class Particle {
constructor() {
// fill this in
}
// fill this in
}
// ═══════════════════════════════════════════════════════════════════════
// UI CONTROL HANDLERS - CUSTOMIZE FOR YOUR PARAMETERS
// ═══════════════════════════════════════════════════════════════════════
function updateParam(paramName, value) {
// fill this in
}
function updateColor(colorId, value) {
// fill this in
}
// ═══════════════════════════════════════════════════════════════════════
// SEED CONTROL FUNCTIONS - ALWAYS KEEP THESE
// ═══════════════════════════════════════════════════════════════════════
function updateSeedDisplay() {
document.getElementById('seed-input').value = params.seed;
}
function updateSeed() {
let input = document.getElementById('seed-input');
let newSeed = parseInt(input.value);
if (newSeed && newSeed > 0) {
params.seed = newSeed;
initializeSystem();
} else {
// Reset to current seed if invalid
updateSeedDisplay();
}
}
function previousSeed() {
params.seed = Math.max(1, params.seed - 1);
updateSeedDisplay();
initializeSystem();
}
function nextSeed() {
params.seed = params.seed + 1;
updateSeedDisplay();
initializeSystem();
}
function randomSeedAndUpdate() {
params.seed = Math.floor(Math.random() * 999999) + 1;
updateSeedDisplay();
initializeSystem();
}
function resetParameters() {
params = {...defaultParams};
// Update UI elements
document.getElementById('particleCount').value = params.particleCount;
document.getElementById('particleCount-value').textContent = params.particleCount;
document.getElementById('flowSpeed').value = params.flowSpeed;
document.getElementById('flowSpeed-value').textContent = params.flowSpeed;
document.getElementById('noiseScale').value = params.noiseScale;
document.getElementById('noiseScale-value').textContent = params.noiseScale;
document.getElementById('trailLength').value = params.trailLength;
document.getElementById('trailLength-value').textContent = params.trailLength;
// Reset colors
document.getElementById('color1').value = params.colorPalette[0];
document.getElementById('color1-value').textContent = params.colorPalette[0];
document.getElementById('color2').value = params.colorPalette[1];
document.getElementById('color2-value').textContent = params.colorPalette[1];
document.getElementById('color3').value = params.colorPalette[2];
document.getElementById('color3-value').textContent = params.colorPalette[2];
updateSeedDisplay();
initializeSystem();
}
// Initialize UI on load
window.addEventListener('load', function() {
updateSeedDisplay();
});
</script>
</body>
</html>
@@ -0,0 +1,317 @@
---
name: analytics-tracking
version: 1.0.0
description: When the user wants to set up, improve, or audit analytics tracking and measurement. Also use when the user mentions "set up tracking," "GA4," "Google Analytics," "conversion tracking," "event tracking," "UTM parameters," "tag manager," "GTM," "analytics implementation," or "tracking plan." For A/B test measurement, see ab-test-setup.
---
# Analytics Tracking
You are an expert in analytics implementation and measurement. Your goal is to help set up tracking that provides actionable insights for marketing and product decisions.
## Initial Assessment
**Check for product marketing context first:**
If `.mosaic/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before implementing tracking, understand:
1. **Business Context** - What decisions will this data inform? What are key conversions?
2. **Current State** - What tracking exists? What tools are in use?
3. **Technical Context** - What's the tech stack? Any privacy/compliance requirements?
---
## Core Principles
### 1. Track for Decisions, Not Data
- Every event should inform a decision
- Avoid vanity metrics
- Quality > quantity of events
### 2. Start with the Questions
- What do you need to know?
- What actions will you take based on this data?
- Work backwards to what you need to track
### 3. Name Things Consistently
- Naming conventions matter
- Establish patterns before implementing
- Document everything
### 4. Maintain Data Quality
- Validate implementation
- Monitor for issues
- Clean data > more data
---
## Tracking Plan Framework
### Structure
```
Event Name | Category | Properties | Trigger | Notes
---------- | -------- | ---------- | ------- | -----
```
### Event Types
| Type | Examples |
| ------------------ | ------------------------------------------------ |
| Pageviews | Automatic, enhanced with metadata |
| User Actions | Button clicks, form submissions, feature usage |
| System Events | Signup completed, purchase, subscription changed |
| Custom Conversions | Goal completions, funnel stages |
**For comprehensive event lists**: See [references/event-library.md](references/event-library.md)
---
## Event Naming Conventions
### Recommended Format: Object-Action
```
signup_completed
button_clicked
form_submitted
article_read
checkout_payment_completed
```
### Best Practices
- Lowercase with underscores
- Be specific: `cta_hero_clicked` vs. `button_clicked`
- Include context in properties, not event name
- Avoid spaces and special characters
- Document decisions
---
## Essential Events
### Marketing Site
| Event | Properties |
| ---------------- | --------------------- |
| cta_clicked | button_text, location |
| form_submitted | form_type |
| signup_completed | method, source |
| demo_requested | - |
### Product/App
| Event | Properties |
| ------------------------- | ---------------------- |
| onboarding_step_completed | step_number, step_name |
| feature_used | feature_name |
| purchase_completed | plan, value |
| subscription_cancelled | reason |
**For full event library by business type**: See [references/event-library.md](references/event-library.md)
---
## Event Properties
### Standard Properties
| Category | Properties |
| -------- | ----------------------------------------- |
| Page | page_title, page_location, page_referrer |
| User | user_id, user_type, account_id, plan_type |
| Campaign | source, medium, campaign, content, term |
| Product | product_id, product_name, category, price |
### Best Practices
- Use consistent property names
- Include relevant context
- Don't duplicate automatic properties
- Avoid PII in properties
---
## GA4 Implementation
### Quick Setup
1. Create GA4 property and data stream
2. Install gtag.js or GTM
3. Enable enhanced measurement
4. Configure custom events
5. Mark conversions in Admin
### Custom Event Example
```javascript
gtag('event', 'signup_completed', {
method: 'email',
plan: 'free',
});
```
**For detailed GA4 implementation**: See [references/ga4-implementation.md](references/ga4-implementation.md)
---
## Google Tag Manager
### Container Structure
| Component | Purpose |
| --------- | --------------------------------------- |
| Tags | Code that executes (GA4, pixels) |
| Triggers | When tags fire (page view, click) |
| Variables | Dynamic values (click text, data layer) |
### Data Layer Pattern
```javascript
dataLayer.push({
event: 'form_submitted',
form_name: 'contact',
form_location: 'footer',
});
```
**For detailed GTM implementation**: See [references/gtm-implementation.md](references/gtm-implementation.md)
---
## UTM Parameter Strategy
### Standard Parameters
| Parameter | Purpose | Example |
| ------------ | ---------------------- | ------------------ |
| utm_source | Traffic source | google, newsletter |
| utm_medium | Marketing medium | cpc, email, social |
| utm_campaign | Campaign name | spring_sale |
| utm_content | Differentiate versions | hero_cta |
| utm_term | Paid search keywords | running+shoes |
### Naming Conventions
- Lowercase everything
- Use underscores or hyphens consistently
- Be specific but concise: `blog_footer_cta`, not `cta1`
- Document all UTMs in a spreadsheet
---
## Debugging and Validation
### Testing Tools
| Tool | Use For |
| ------------------ | ---------------------------------- |
| GA4 DebugView | Real-time event monitoring |
| GTM Preview Mode | Test triggers before publish |
| Browser Extensions | Tag Assistant, dataLayer Inspector |
### Validation Checklist
- [ ] Events firing on correct triggers
- [ ] Property values populating correctly
- [ ] No duplicate events
- [ ] Works across browsers and mobile
- [ ] Conversions recorded correctly
- [ ] No PII leaking
### Common Issues
| Issue | Check |
| ----------------- | ----------------------------------------- |
| Events not firing | Trigger config, GTM loaded |
| Wrong values | Variable path, data layer structure |
| Duplicate events | Multiple containers, trigger firing twice |
---
## Privacy and Compliance
### Considerations
- Cookie consent required in EU/UK/CA
- No PII in analytics properties
- Data retention settings
- User deletion capabilities
### Implementation
- Use consent mode (wait for consent)
- IP anonymization
- Only collect what you need
- Integrate with consent management platform
---
## Output Format
### Tracking Plan Document
```markdown
# [Site/Product] Tracking Plan
## Overview
- Tools: GA4, GTM
- Last updated: [Date]
## Events
| Event Name | Description | Properties | Trigger |
| ---------------- | --------------------- | ------------ | ------------ |
| signup_completed | User completes signup | method, plan | Success page |
## Custom Dimensions
| Name | Scope | Parameter |
| --------- | ----- | --------- |
| user_type | User | user_type |
## Conversions
| Conversion | Event | Counting |
| ---------- | ---------------- | ---------------- |
| Signup | signup_completed | Once per session |
```
---
## Task-Specific Questions
1. What tools are you using (GA4, Mixpanel, etc.)?
2. What key actions do you want to track?
3. What decisions will this data inform?
4. Who implements - dev team or marketing?
5. Are there privacy/consent requirements?
6. What's already tracked?
---
## Tool Integrations
For implementation, see the [tools registry](../../tools/REGISTRY.md). Key analytics tools:
| Tool | Best For | MCP | Guide |
| ------------- | ------------------------------------- | :-: | ----------------------------------------------------- |
| **GA4** | Web analytics, Google ecosystem | ✓ | [ga4.md](../../tools/integrations/ga4.md) |
| **Mixpanel** | Product analytics, event tracking | - | [mixpanel.md](../../tools/integrations/mixpanel.md) |
| **Amplitude** | Product analytics, cohort analysis | - | [amplitude.md](../../tools/integrations/amplitude.md) |
| **PostHog** | Open-source analytics, session replay | - | [posthog.md](../../tools/integrations/posthog.md) |
| **Segment** | Customer data platform, routing | - | [segment.md](../../tools/integrations/segment.md) |
---
## Related Skills
- **ab-test-setup**: For experiment tracking
- **seo-audit**: For organic traffic analysis
- **page-cro**: For conversion optimization (uses this data)
@@ -0,0 +1,259 @@
# Event Library Reference
Comprehensive list of events to track by business type and context.
## Marketing Site Events
### Navigation & Engagement
| Event Name | Description | Properties |
| --------------------- | -------------------------- | ---------------------------------------- |
| page_view | Page loaded (enhanced) | page_title, page_location, content_group |
| scroll_depth | User scrolled to threshold | depth (25, 50, 75, 100) |
| outbound_link_clicked | Click to external site | link_url, link_text |
| internal_link_clicked | Click within site | link_url, link_text, location |
| video_played | Video started | video_id, video_title, duration |
| video_completed | Video finished | video_id, video_title, duration |
### CTA & Form Interactions
| Event Name | Description | Properties |
| -------------------- | ---------------------- | ------------------------------- |
| cta_clicked | Call to action clicked | button_text, cta_location, page |
| form_started | User began form | form_name, form_location |
| form_field_completed | Field filled | form_name, field_name |
| form_submitted | Form successfully sent | form_name, form_location |
| form_error | Form validation failed | form_name, error_type |
| resource_downloaded | Asset downloaded | resource_name, resource_type |
### Conversion Events
| Event Name | Description | Properties |
| --------------------- | ------------------- | ---------------------- |
| signup_started | Initiated signup | source, page |
| signup_completed | Finished signup | method, plan, source |
| demo_requested | Demo form submitted | company_size, industry |
| contact_submitted | Contact form sent | inquiry_type |
| newsletter_subscribed | Email list signup | source, list_name |
| trial_started | Free trial began | plan, source |
---
## Product/App Events
### Onboarding
| Event Name | Description | Properties |
| -------------------------- | ----------------------- | --------------------------------- |
| signup_completed | Account created | method, referral_source |
| onboarding_started | Began onboarding | - |
| onboarding_step_completed | Step finished | step_number, step_name |
| onboarding_completed | All steps done | steps_completed, time_to_complete |
| onboarding_skipped | User skipped onboarding | step_skipped_at |
| first_key_action_completed | Aha moment reached | action_type |
### Core Usage
| Event Name | Description | Properties |
| ---------------- | --------------------- | ------------------------------ |
| session_started | App session began | session_number |
| feature_used | Feature interaction | feature_name, feature_category |
| action_completed | Core action done | action_type, count |
| content_created | User created content | content_type |
| content_edited | User modified content | content_type |
| content_deleted | User removed content | content_type |
| search_performed | In-app search | query, results_count |
| settings_changed | Settings modified | setting_name, new_value |
| invite_sent | User invited others | invite_type, count |
### Errors & Support
| Event Name | Description | Properties |
| ------------------ | -------------------- | ------------------------------- |
| error_occurred | Error experienced | error_type, error_message, page |
| help_opened | Help accessed | help_type, page |
| support_contacted | Support request made | contact_method, issue_type |
| feedback_submitted | User feedback given | feedback_type, rating |
---
## Monetization Events
### Pricing & Checkout
| Event Name | Description | Properties |
| -------------------- | ------------------- | ------------------------------------- |
| pricing_viewed | Pricing page seen | source |
| plan_selected | Plan chosen | plan_name, billing_cycle |
| checkout_started | Began checkout | plan, value |
| payment_info_entered | Payment submitted | payment_method |
| purchase_completed | Purchase successful | plan, value, currency, transaction_id |
| purchase_failed | Purchase failed | error_reason, plan |
### Subscription Management
| Event Name | Description | Properties |
| ----------------------- | ---------------------- | ------------------------- |
| trial_started | Trial began | plan, trial_length |
| trial_ended | Trial expired | plan, converted (bool) |
| subscription_upgraded | Plan upgraded | from_plan, to_plan, value |
| subscription_downgraded | Plan downgraded | from_plan, to_plan |
| subscription_cancelled | Cancelled | plan, reason, tenure |
| subscription_renewed | Renewed | plan, value |
| billing_updated | Payment method changed | - |
---
## E-commerce Events
### Browsing
| Event Name | Description | Properties |
| ------------------- | -------------------- | ----------------------------------------- |
| product_viewed | Product page viewed | product_id, product_name, category, price |
| product_list_viewed | Category/list viewed | list_name, products[] |
| product_searched | Search performed | query, results_count |
| product_filtered | Filters applied | filter_type, filter_value |
| product_sorted | Sort applied | sort_by, sort_order |
### Cart
| Event Name | Description | Properties |
| ------------------------- | ---------------- | ----------------------------------------- |
| product_added_to_cart | Item added | product_id, product_name, price, quantity |
| product_removed_from_cart | Item removed | product_id, product_name, price, quantity |
| cart_viewed | Cart page viewed | cart_value, items_count |
### Checkout
| Event Name | Description | Properties |
| ----------------------- | --------------- | ---------------------------------------- |
| checkout_started | Checkout began | cart_value, items_count |
| checkout_step_completed | Step finished | step_number, step_name |
| shipping_info_entered | Address entered | shipping_method |
| payment_info_entered | Payment entered | payment_method |
| coupon_applied | Coupon used | coupon_code, discount_value |
| purchase_completed | Order placed | transaction_id, value, currency, items[] |
### Post-Purchase
| Event Name | Description | Properties |
| ---------------- | ------------------- | ---------------------- |
| order_confirmed | Confirmation viewed | transaction_id |
| refund_requested | Refund initiated | transaction_id, reason |
| refund_completed | Refund processed | transaction_id, value |
| review_submitted | Product reviewed | product_id, rating |
---
## B2B / SaaS Specific Events
### Team & Collaboration
| Event Name | Description | Properties |
| ------------------- | ------------------- | --------------------------- |
| team_created | New team/org made | team_size, plan |
| team_member_invited | Invite sent | role, invite_method |
| team_member_joined | Member accepted | role |
| team_member_removed | Member removed | role |
| role_changed | Permissions updated | user_id, old_role, new_role |
### Integration Events
| Event Name | Description | Properties |
| ------------------------ | ---------------------- | ------------------------ |
| integration_viewed | Integration page seen | integration_name |
| integration_started | Setup began | integration_name |
| integration_connected | Successfully connected | integration_name |
| integration_disconnected | Removed integration | integration_name, reason |
### Account Events
| Event Name | Description | Properties |
| ------------------- | ----------------- | ------------------------- |
| account_created | New account | source, plan |
| account_upgraded | Plan upgrade | from_plan, to_plan |
| account_churned | Account closed | reason, tenure, mrr_lost |
| account_reactivated | Returned customer | previous_tenure, new_plan |
---
## Event Properties (Parameters)
### Standard Properties to Include
**User Context:**
```
user_id: "12345"
user_type: "free" | "trial" | "paid"
account_id: "acct_123"
plan_type: "starter" | "pro" | "enterprise"
```
**Session Context:**
```
session_id: "sess_abc"
session_number: 5
page: "/pricing"
referrer: "https://google.com"
```
**Campaign Context:**
```
source: "google"
medium: "cpc"
campaign: "spring_sale"
content: "hero_cta"
```
**Product Context (E-commerce):**
```
product_id: "SKU123"
product_name: "Product Name"
category: "Category"
price: 99.99
quantity: 1
currency: "USD"
```
**Timing:**
```
timestamp: "2024-01-15T10:30:00Z"
time_on_page: 45
session_duration: 300
```
---
## Funnel Event Sequences
### Signup Funnel
1. signup_started
2. signup_step_completed (email)
3. signup_step_completed (password)
4. signup_completed
5. onboarding_started
### Purchase Funnel
1. pricing_viewed
2. plan_selected
3. checkout_started
4. payment_info_entered
5. purchase_completed
### E-commerce Funnel
1. product_viewed
2. product_added_to_cart
3. cart_viewed
4. checkout_started
5. shipping_info_entered
6. payment_info_entered
7. purchase_completed
@@ -0,0 +1,309 @@
# GA4 Implementation Reference
Detailed implementation guide for Google Analytics 4.
## Configuration
### Data Streams
- One stream per platform (web, iOS, Android)
- Enable enhanced measurement for automatic tracking
- Configure data retention (2 months default, 14 months max)
- Enable Google Signals (for cross-device, if consented)
### Enhanced Measurement Events (Automatic)
| Event | Description | Configuration |
| ---------------- | ------------------------ | ----------------------- |
| page_view | Page loads | Automatic |
| scroll | 90% scroll depth | Toggle on/off |
| outbound_click | Click to external domain | Automatic |
| site_search | Search query used | Configure parameter |
| video_engagement | YouTube video plays | Toggle on/off |
| file_download | PDF, docs, etc. | Configurable extensions |
### Recommended Events
Use Google's predefined events when possible for enhanced reporting:
**All properties:**
- login, sign_up
- share
- search
**E-commerce:**
- view_item, view_item_list
- add_to_cart, remove_from_cart
- begin_checkout
- add_payment_info
- purchase, refund
**Games:**
- level_up, unlock_achievement
- post_score, spend_virtual_currency
Reference: https://support.google.com/analytics/answer/9267735
---
## Custom Events
### gtag.js Implementation
```javascript
// Basic event
gtag('event', 'signup_completed', {
method: 'email',
plan: 'free',
});
// Event with value
gtag('event', 'purchase', {
transaction_id: 'T12345',
value: 99.99,
currency: 'USD',
items: [
{
item_id: 'SKU123',
item_name: 'Product Name',
price: 99.99,
},
],
});
// User properties
gtag('set', 'user_properties', {
user_type: 'premium',
plan_name: 'pro',
});
// User ID (for logged-in users)
gtag('config', 'GA_MEASUREMENT_ID', {
user_id: 'USER_ID',
});
```
### Google Tag Manager (dataLayer)
```javascript
// Custom event
dataLayer.push({
event: 'signup_completed',
method: 'email',
plan: 'free',
});
// Set user properties
dataLayer.push({
user_id: '12345',
user_type: 'premium',
});
// E-commerce purchase
dataLayer.push({
event: 'purchase',
ecommerce: {
transaction_id: 'T12345',
value: 99.99,
currency: 'USD',
items: [
{
item_id: 'SKU123',
item_name: 'Product Name',
price: 99.99,
quantity: 1,
},
],
},
});
// Clear ecommerce before sending (best practice)
dataLayer.push({ ecommerce: null });
dataLayer.push({
event: 'view_item',
ecommerce: {
// ...
},
});
```
---
## Conversions Setup
### Creating Conversions
1. **Collect the event** - Ensure event is firing in GA4
2. **Mark as conversion** - Admin > Events > Mark as conversion
3. **Set counting method**:
- Once per session (leads, signups)
- Every event (purchases)
4. **Import to Google Ads** - For conversion-optimized bidding
### Conversion Values
```javascript
// Event with conversion value
gtag('event', 'purchase', {
value: 99.99,
currency: 'USD',
});
```
Or set default value in GA4 Admin when marking conversion.
---
## Custom Dimensions and Metrics
### When to Use
**Custom dimensions:**
- Properties you want to segment/filter by
- User attributes (plan type, industry)
- Content attributes (author, category)
**Custom metrics:**
- Numeric values to aggregate
- Scores, counts, durations
### Setup Steps
1. Admin > Data display > Custom definitions
2. Create dimension or metric
3. Choose scope:
- **Event**: Per event (content_type)
- **User**: Per user (account_type)
- **Item**: Per product (product_category)
4. Enter parameter name (must match event parameter)
### Examples
| Dimension | Scope | Parameter | Description |
| ---------------- | ----- | ------------- | ------------------- |
| User Type | User | user_type | Free, trial, paid |
| Content Author | Event | author | Blog post author |
| Product Category | Item | item_category | E-commerce category |
---
## Audiences
### Creating Audiences
Admin > Data display > Audiences
**Use cases:**
- Remarketing audiences (export to Ads)
- Segment analysis
- Trigger-based events
### Audience Examples
**High-intent visitors:**
- Viewed pricing page
- Did not convert
- In last 7 days
**Engaged users:**
- 3+ sessions
- Or 5+ minutes total engagement
**Purchasers:**
- Purchase event
- For exclusion or lookalike
---
## Debugging
### DebugView
Enable with:
- URL parameter: `?debug_mode=true`
- Chrome extension: GA Debugger
- gtag: `'debug_mode': true` in config
View at: Reports > Configure > DebugView
### Real-Time Reports
Check events within 30 minutes:
Reports > Real-time
### Common Issues
**Events not appearing:**
- Check DebugView first
- Verify gtag/GTM firing
- Check filter exclusions
**Parameter values missing:**
- Custom dimension not created
- Parameter name mismatch
- Data still processing (24-48 hrs)
**Conversions not recording:**
- Event not marked as conversion
- Event name doesn't match
- Counting method (once vs. every)
---
## Data Quality
### Filters
Admin > Data streams > [Stream] > Configure tag settings > Define internal traffic
**Exclude:**
- Internal IP addresses
- Developer traffic
- Testing environments
### Cross-Domain Tracking
For multiple domains sharing analytics:
1. Admin > Data streams > [Stream] > Configure tag settings
2. Configure your domains
3. List all domains that should share sessions
### Session Settings
Admin > Data streams > [Stream] > Configure tag settings
- Session timeout (default 30 min)
- Engaged session duration (10 sec default)
---
## Integration with Google Ads
### Linking
1. Admin > Product links > Google Ads links
2. Enable auto-tagging in Google Ads
3. Import conversions in Google Ads
### Audience Export
Audiences created in GA4 can be used in Google Ads for:
- Remarketing campaigns
- Customer match
- Similar audiences
@@ -0,0 +1,410 @@
# Google Tag Manager Implementation Reference
Detailed guide for implementing tracking via Google Tag Manager.
## Container Structure
### Tags
Tags are code snippets that execute when triggered.
**Common tag types:**
- GA4 Configuration (base setup)
- GA4 Event (custom events)
- Google Ads Conversion
- Facebook Pixel
- LinkedIn Insight Tag
- Custom HTML (for other pixels)
### Triggers
Triggers define when tags fire.
**Built-in triggers:**
- Page View: All Pages, DOM Ready, Window Loaded
- Click: All Elements, Just Links
- Form Submission
- Scroll Depth
- Timer
- Element Visibility
**Custom triggers:**
- Custom Event (from dataLayer)
- Trigger Groups (multiple conditions)
### Variables
Variables capture dynamic values.
**Built-in (enable as needed):**
- Click Text, Click URL, Click ID, Click Classes
- Page Path, Page URL, Page Hostname
- Referrer
- Form Element, Form ID
**User-defined:**
- Data Layer variables
- JavaScript variables
- Lookup tables
- RegEx tables
- Constants
---
## Naming Conventions
### Recommended Format
```
[Type] - [Description] - [Detail]
Tags:
GA4 - Event - Signup Completed
GA4 - Config - Base Configuration
FB - Pixel - Page View
HTML - LiveChat Widget
Triggers:
Click - CTA Button
Submit - Contact Form
View - Pricing Page
Custom - signup_completed
Variables:
DL - user_id
JS - Current Timestamp
LT - Campaign Source Map
```
---
## Data Layer Patterns
### Basic Structure
```javascript
// Initialize (in <head> before GTM)
window.dataLayer = window.dataLayer || [];
// Push event
dataLayer.push({
event: 'event_name',
property1: 'value1',
property2: 'value2',
});
```
### Page Load Data
```javascript
// Set on page load (before GTM container)
window.dataLayer = window.dataLayer || [];
dataLayer.push({
pageType: 'product',
contentGroup: 'products',
user: {
loggedIn: true,
userId: '12345',
userType: 'premium',
},
});
```
### Form Submission
```javascript
document.querySelector('#contact-form').addEventListener('submit', function () {
dataLayer.push({
event: 'form_submitted',
formName: 'contact',
formLocation: 'footer',
});
});
```
### Button Click
```javascript
document.querySelector('.cta-button').addEventListener('click', function () {
dataLayer.push({
event: 'cta_clicked',
ctaText: this.innerText,
ctaLocation: 'hero',
});
});
```
### E-commerce Events
```javascript
// Product view
dataLayer.push({ ecommerce: null }); // Clear previous
dataLayer.push({
event: 'view_item',
ecommerce: {
items: [
{
item_id: 'SKU123',
item_name: 'Product Name',
price: 99.99,
item_category: 'Category',
quantity: 1,
},
],
},
});
// Add to cart
dataLayer.push({ ecommerce: null });
dataLayer.push({
event: 'add_to_cart',
ecommerce: {
items: [
{
item_id: 'SKU123',
item_name: 'Product Name',
price: 99.99,
quantity: 1,
},
],
},
});
// Purchase
dataLayer.push({ ecommerce: null });
dataLayer.push({
event: 'purchase',
ecommerce: {
transaction_id: 'T12345',
value: 99.99,
currency: 'USD',
tax: 5.0,
shipping: 10.0,
items: [
{
item_id: 'SKU123',
item_name: 'Product Name',
price: 99.99,
quantity: 1,
},
],
},
});
```
---
## Common Tag Configurations
### GA4 Configuration Tag
**Tag Type:** Google Analytics: GA4 Configuration
**Settings:**
- Measurement ID: G-XXXXXXXX
- Send page view: Checked (for pageviews)
- User Properties: Add any user-level dimensions
**Trigger:** All Pages
### GA4 Event Tag
**Tag Type:** Google Analytics: GA4 Event
**Settings:**
- Configuration Tag: Select your config tag
- Event Name: {{DL - event_name}} or hardcode
- Event Parameters: Add parameters from dataLayer
**Trigger:** Custom Event with event name match
### Facebook Pixel - Base
**Tag Type:** Custom HTML
```html
<script>
!(function (f, b, e, v, n, t, s) {
if (f.fbq) return;
n = f.fbq = function () {
n.callMethod ? n.callMethod.apply(n, arguments) : n.queue.push(arguments);
};
if (!f._fbq) f._fbq = n;
n.push = n;
n.loaded = !0;
n.version = '2.0';
n.queue = [];
t = b.createElement(e);
t.async = !0;
t.src = v;
s = b.getElementsByTagName(e)[0];
s.parentNode.insertBefore(t, s);
})(window, document, 'script', 'https://connect.facebook.net/en_US/fbevents.js');
fbq('init', 'YOUR_PIXEL_ID');
fbq('track', 'PageView');
</script>
```
**Trigger:** All Pages
### Facebook Pixel - Event
**Tag Type:** Custom HTML
```html
<script>
fbq('track', 'Lead', {
content_name: '{{DL - form_name}}',
});
</script>
```
**Trigger:** Custom Event - form_submitted
---
## Preview and Debug
### Preview Mode
1. Click "Preview" in GTM
2. Enter site URL
3. GTM debug panel opens at bottom
**What to check:**
- Tags fired on this event
- Tags not fired (and why)
- Variables and their values
- Data layer contents
### Debug Tips
**Tag not firing:**
- Check trigger conditions
- Verify data layer push
- Check tag sequencing
**Wrong variable value:**
- Check data layer structure
- Verify variable path (nested objects)
- Check timing (data may not exist yet)
**Multiple firings:**
- Check trigger uniqueness
- Look for duplicate tags
- Check tag firing options
---
## Workspaces and Versioning
### Workspaces
Use workspaces for team collaboration:
- Default workspace for production
- Separate workspaces for large changes
- Merge when ready
### Version Management
**Best practices:**
- Name every version descriptively
- Add notes explaining changes
- Review changes before publish
- Keep production version noted
**Version notes example:**
```
v15: Added purchase conversion tracking
- New tag: GA4 - Event - Purchase
- New trigger: Custom Event - purchase
- New variables: DL - transaction_id, DL - value
- Tested: Chrome, Safari, Mobile
```
---
## Consent Management
### Consent Mode Integration
```javascript
// Default state (before consent)
gtag('consent', 'default', {
analytics_storage: 'denied',
ad_storage: 'denied',
});
// Update on consent
function grantConsent() {
gtag('consent', 'update', {
analytics_storage: 'granted',
ad_storage: 'granted',
});
}
```
### GTM Consent Overview
1. Enable Consent Overview in Admin
2. Configure consent for each tag
3. Tags respect consent state automatically
---
## Advanced Patterns
### Tag Sequencing
**Setup tags to fire in order:**
Tag Configuration > Advanced Settings > Tag Sequencing
**Use cases:**
- Config tag before event tags
- Pixel initialization before tracking
- Cleanup after conversion
### Exception Handling
**Trigger exceptions** - Prevent tag from firing:
- Exclude certain pages
- Exclude internal traffic
- Exclude during testing
### Custom JavaScript Variables
```javascript
// Get URL parameter
function() {
var params = new URLSearchParams(window.location.search);
return params.get('campaign') || '(not set)';
}
// Get cookie value
function() {
var match = document.cookie.match('(^|;) ?user_id=([^;]*)(;|$)');
return match ? match[2] : null;
}
// Get data from page
function() {
var el = document.querySelector('.product-price');
return el ? parseFloat(el.textContent.replace('$', '')) : 0;
}
```
@@ -0,0 +1,129 @@
---
name: antfu
description: Anthony Fu's opinionated tooling and conventions for JavaScript/TypeScript projects. Use when setting up new projects, configuring ESLint/Prettier alternatives, monorepos, library publishing, or when the user mentions Anthony Fu's preferences.
metadata:
author: Anthony Fu
version: '2026.02.03'
---
## Coding Practices
### Code Organization
- **Single responsibility**: Each source file should have a clear, focused scope/purpose
- **Split large files**: Break files when they become large or handle too many concerns
- **Type separation**: Always separate types and interfaces into `types.ts` or `types/*.ts`
- **Constants extraction**: Move constants to a dedicated `constants.ts` file
### Runtime Environment
- **Prefer isomorphic code**: Write runtime-agnostic code that works in Node, browser, and workers whenever possible
- **Clear runtime indicators**: When code is environment-specific, add a comment at the top of the file:
```ts
// @env node
// @env browser
```
### TypeScript
- **Explicit return types**: Declare return types explicitly when possible
- **Avoid complex inline types**: Extract complex types into dedicated `type` or `interface` declarations
### Comments
- **Avoid unnecessary comments**: Code should be self-explanatory
- **Explain "why" not "how"**: Comments should describe the reasoning or intent, not what the code does
### Testing (Vitest)
- Test files: `foo.ts``foo.test.ts` (same directory)
- Use `describe`/`it` API (not `test`)
- Use `toMatchSnapshot` for complex outputs
- Use `toMatchFileSnapshot` with explicit path for language-specific snapshots
---
## Tooling Choices
### @antfu/ni Commands
| Command | Description |
| -------------------------- | ------------------------------------------ |
| `ni` | Install dependencies |
| `ni <pkg>` / `ni -D <pkg>` | Add dependency / dev dependency |
| `nr <script>` | Run script |
| `nu` | Upgrade dependencies |
| `nun <pkg>` | Uninstall dependency |
| `nci` | Clean install (`pnpm i --frozen-lockfile`) |
| `nlx <pkg>` | Execute package (`npx`) |
### TypeScript Config
```json
{
"compilerOptions": {
"target": "ESNext",
"module": "ESNext",
"moduleResolution": "bundler",
"strict": true,
"esModuleInterop": true,
"skipLibCheck": true,
"resolveJsonModule": true,
"isolatedModules": true,
"noEmit": true
}
}
```
### ESLint Setup
```js
// eslint.config.mjs
import antfu from '@antfu/eslint-config';
export default antfu();
```
When completing tasks, run `pnpm run lint --fix` to format the code and fix coding style.
For detailed configuration options: [antfu-eslint-config](references/antfu-eslint-config.md)
### Git Hooks
```json
{
"simple-git-hooks": {
"pre-commit": "pnpm i --frozen-lockfile --ignore-scripts --offline && npx lint-staged"
},
"lint-staged": { "*": "eslint --fix" },
"scripts": {
"prepare": "npx simple-git-hooks"
}
}
```
### pnpm Catalogs
Use named catalogs in `pnpm-workspace.yaml` for version management:
| Catalog | Purpose |
| ---------- | ------------------------------------ |
| `prod` | Production dependencies |
| `inlined` | Bundler-inlined dependencies |
| `dev` | Dev tools (linter, bundler, testing) |
| `frontend` | Frontend libraries |
Avoid the default catalog. Catalog names can be adjusted per project needs.
---
## References
| Topic | Description | Reference |
| ------------------- | --------------------------------------------------------------- | -------------------------------------------------------- |
| ESLint Config | Framework support, formatters, rule overrides, VS Code settings | [antfu-eslint-config](references/antfu-eslint-config.md) |
| Project Setup | .gitignore, GitHub Actions, VS Code extensions | [setting-up](references/setting-up.md) |
| App Development | Vue/Nuxt/UnoCSS conventions and patterns | [app-development](references/app-development.md) |
| Library Development | tsdown bundling, pure ESM publishing | [library-development](references/library-development.md) |
| Monorepo | pnpm workspaces, centralized alias, Turborepo | [monorepo](references/monorepo.md) |
@@ -0,0 +1,301 @@
---
name: antfu-eslint-config
description: Configuring @antfu/eslint-config for framework support, formatters, and rule overrides. Use when adding React/Vue/Svelte/Astro support, customizing rules, or setting up VS Code integration.
---
# @antfu/eslint-config
Handles both linting and formatting (no Prettier needed). Auto-detects TypeScript and Vue.
**Style**: Single quotes, no semicolons, sorted imports, dangling commas.
## Configuration Options
```js
import antfu from '@antfu/eslint-config';
export default antfu({
// Project type: 'lib' for libraries, 'app' (default) for applications
type: 'lib',
// Global ignores (extends defaults, doesn't override)
ignores: ['**/fixtures', '**/dist'],
// Stylistic options
stylistic: {
indent: 2, // 2, 4, or 'tab'
quotes: 'single', // or 'double'
},
// Framework support (auto-detected, but can be explicit)
typescript: true,
vue: true,
// Disable specific language support
jsonc: false,
yaml: false,
});
```
## Framework Support
### Vue
Vue accessibility:
```js
export default antfu({
vue: {
a11y: true,
},
});
// Requires: pnpm add -D eslint-plugin-vuejs-accessibility
```
### React
```js
export default antfu({
react: true,
});
// Requires: pnpm add -D @eslint-react/eslint-plugin eslint-plugin-react-hooks eslint-plugin-react-refresh
```
### Next.js
```js
export default antfu({
nextjs: true,
});
// Requires: pnpm add -D @next/eslint-plugin-next
```
### Svelte
```js
export default antfu({
svelte: true,
});
// Requires: pnpm add -D eslint-plugin-svelte
```
### Astro
```js
export default antfu({
astro: true,
});
// Requires: pnpm add -D eslint-plugin-astro
```
### Solid
```js
export default antfu({
solid: true,
});
// Requires: pnpm add -D eslint-plugin-solid
```
### UnoCSS
```js
export default antfu({
unocss: true,
});
// Requires: pnpm add -D @unocss/eslint-plugin
```
## Formatters (CSS, HTML, Markdown)
For files ESLint doesn't handle natively:
```js
export default antfu({
formatters: {
css: true, // Format CSS, LESS, SCSS (uses Prettier)
html: true, // Format HTML (uses Prettier)
markdown: 'prettier', // or 'dprint'
},
});
// Requires: pnpm add -D eslint-plugin-format
```
## Rule Overrides
### Global overrides
```js
export default antfu(
{
// First argument: antfu config options
},
// Additional arguments: ESLint flat configs
{
rules: {
'style/semi': ['error', 'never'],
},
},
);
```
### Per-integration overrides
```js
export default antfu({
vue: {
overrides: {
'vue/operator-linebreak': ['error', 'before'],
},
},
typescript: {
overrides: {
'ts/consistent-type-definitions': ['error', 'interface'],
},
},
});
```
### File-specific overrides
```js
export default antfu(
{ vue: true, typescript: true },
{
files: ['**/*.vue'],
rules: {
'vue/operator-linebreak': ['error', 'before'],
},
},
);
```
## Plugin Prefix Renaming
The config renames plugin prefixes for consistency:
| New Prefix | Original |
| ---------- | ---------------------- |
| `ts/*` | `@typescript-eslint/*` |
| `style/*` | `@stylistic/*` |
| `import/*` | `import-lite/*` |
| `node/*` | `n/*` |
| `yaml/*` | `yml/*` |
| `test/*` | `vitest/*` |
| `next/*` | `@next/next` |
Use the new prefix when overriding or disabling rules:
```ts
// eslint-disable-next-line ts/consistent-type-definitions
type Foo = { bar: 2 };
```
## Type-Aware Rules
Enable TypeScript type checking:
```js
export default antfu({
typescript: {
tsconfigPath: 'tsconfig.json',
},
});
```
## Config Composer API
Chain methods for flexible composition:
```js
export default antfu()
.prepend(/* configs before main */)
.override('antfu/stylistic/rules', {
rules: {
'style/generator-star-spacing': ['error', { after: true, before: false }],
},
})
.renamePlugins({
'old-prefix': 'new-prefix',
});
```
## Less Opinionated Mode
Disable Anthony's most opinionated rules:
```js
export default antfu({
lessOpinionated: true,
});
```
## Lint-Staged Setup
```json
{
"simple-git-hooks": {
"pre-commit": "pnpm lint-staged"
},
"lint-staged": {
"*": "eslint --fix"
}
}
```
```bash
pnpm add -D lint-staged simple-git-hooks
npx simple-git-hooks
```
## VS Code Settings
Add to `.vscode/settings.json`:
```jsonc
{
"prettier.enable": false,
"editor.formatOnSave": false,
"editor.codeActionsOnSave": {
"source.fixAll.eslint": "explicit",
"source.organizeImports": "never",
},
"eslint.rules.customizations": [
{ "rule": "style/*", "severity": "off", "fixable": true },
{ "rule": "format/*", "severity": "off", "fixable": true },
{ "rule": "*-indent", "severity": "off", "fixable": true },
{ "rule": "*-spacing", "severity": "off", "fixable": true },
{ "rule": "*-spaces", "severity": "off", "fixable": true },
{ "rule": "*-order", "severity": "off", "fixable": true },
{ "rule": "*-dangle", "severity": "off", "fixable": true },
{ "rule": "*-newline", "severity": "off", "fixable": true },
{ "rule": "*quotes", "severity": "off", "fixable": true },
{ "rule": "*semi", "severity": "off", "fixable": true },
],
"eslint.validate": [
"javascript",
"javascriptreact",
"typescript",
"typescriptreact",
"vue",
"html",
"markdown",
"json",
"jsonc",
"yaml",
"toml",
"xml",
"astro",
"svelte",
"css",
"less",
"scss",
],
}
```
<!--
Source references:
- https://github.com/antfu/eslint-config
- https://raw.githubusercontent.com/antfu/eslint-config/refs/heads/main/README.md
-->
@@ -0,0 +1,45 @@
---
name: app-development
description: Vue/Nuxt/UnoCSS application conventions. Use when building web apps, choosing between Vite and Nuxt, or writing Vue components.
---
# App Development
## Framework Selection
| Use Case | Choice |
| ------------------------------------------------------ | ---------- |
| SPA, client-only, library playgrounds | Vite + Vue |
| SSR, SSG, SEO-critical, file-based routing, API routes | Nuxt |
## Vue Conventions
| Convention | Preference |
| ------------- | ---------------------------------- |
| Script syntax | Always `<script setup lang="ts">` |
| State | Prefer `shallowRef()` over `ref()` |
| Objects | Use `ref()`, avoid `reactive()` |
| Styling | UnoCSS |
| Utilities | VueUse |
### Props and Emits
```vue
<script setup lang="ts">
interface Props {
title: string;
count?: number;
}
interface Emits {
(e: 'update', value: number): void;
(e: 'close'): void;
}
const props = withDefaults(defineProps<Props>(), {
count: 0,
});
const emit = defineEmits<Emits>();
</script>
```
@@ -0,0 +1,76 @@
---
name: library-development
description: Building and publishing TypeScript libraries with tsdown. Use when creating npm packages, configuring library bundling, or setting up package.json exports.
---
# Library Development
| Aspect | Choice |
| ------- | ------------------------- |
| Bundler | tsdown |
| Output | Pure ESM only (no CJS) |
| DTS | Generated via tsdown |
| Exports | Auto-generated via tsdown |
## tsdown Configuration
Use tsdown with these options enabled:
```ts
// tsdown.config.ts
import { defineConfig } from 'tsdown';
export default defineConfig({
entry: ['src/index.ts'],
format: ['esm'],
dts: true,
exports: true,
});
```
| Option | Value | Purpose |
| --------- | --------- | --------------------------------------------- |
| `format` | `['esm']` | Pure ESM, no CommonJS |
| `dts` | `true` | Generate `.d.ts` files |
| `exports` | `true` | Auto-update `exports` field in `package.json` |
### Multiple Entry Points
```ts
export default defineConfig({
entry: ['src/index.ts', 'src/utils.ts'],
format: ['esm'],
dts: true,
exports: true,
});
```
The `exports: true` option auto-generates the `exports` field in `package.json` when running `tsdown`.
---
## package.json
Required fields for pure ESM library:
```json
{
"type": "module",
"main": "./dist/index.mjs",
"module": "./dist/index.mjs",
"types": "./dist/index.d.mts",
"files": ["dist"],
"scripts": {
"build": "tsdown",
"prepack": "pnpm build",
"test": "vitest",
"release": "bumpp -r"
}
}
```
The `exports` field is managed by tsdown when `exports: true`.
### prepack Script
For each public package, add `"prepack": "pnpm build"` to `scripts`. This ensures the package is automatically built before publishing (e.g., when running `npm publish` or `pnpm publish`). This prevents accidentally publishing stale or missing build artifacts.
@@ -0,0 +1,120 @@
---
name: monorepo
description: Monorepo setup with pnpm workspaces, centralized aliases, and Turborepo. Use when creating or managing multi-package repositories.
---
# Monorepo Setup
## pnpm Workspaces
Use pnpm workspaces for monorepo management:
```yaml
# pnpm-workspace.yaml
packages:
- 'packages/*'
```
## Scripts Convention
Have scripts in each package, and use `-r` (recursive) flag at root,
Enable ESLint cache for faster linting in monorepos.
```json
// root package.json
{
"scripts": {
"build": "pnpm run -r build",
"test": "vitest",
"lint": "eslint . --cache --concurrency=auto"
}
}
```
In each package's `package.json`, add the scripts.
```json
// packages/*/package.json
{
"scripts": {
"build": "tsdown",
"prepack": "pnpm build"
}
}
```
## ESLint Cache
```json
{
"scripts": {
"lint": "eslint . --cache --concurrency=auto"
}
}
```
## Turborepo (Optional)
For monorepos with many packages or long build times, use Turborepo for task orchestration and caching.
See the dedicated Turborepo skill for detailed configuration.
## Centralized Alias
For better DX across Vite, Nuxt, Vitest configs, create a centralized `alias.ts` at project root:
```ts
// alias.ts
import fs from 'node:fs';
import { fileURLToPath } from 'node:url';
import { join, relative } from 'pathe';
const root = fileURLToPath(new URL('.', import.meta.url));
const r = (path: string) => fileURLToPath(new URL(`./packages/${path}`, import.meta.url));
export const alias = {
'@myorg/core': r('core/src/index.ts'),
'@myorg/utils': r('utils/src/index.ts'),
'@myorg/ui': r('ui/src/index.ts'),
// Add more aliases as needed
};
// Auto-update tsconfig.alias.json paths
const raw = fs.readFileSync(join(root, 'tsconfig.alias.json'), 'utf-8').trim();
const tsconfig = JSON.parse(raw);
tsconfig.compilerOptions.paths = Object.fromEntries(
Object.entries(alias).map(([key, value]) => [key, [`./${relative(root, value)}`]]),
);
const newRaw = JSON.stringify(tsconfig, null, 2);
if (newRaw !== raw) fs.writeFileSync(join(root, 'tsconfig.alias.json'), `${newRaw}\n`, 'utf-8');
```
Then update the `tsconfig.json` to use the alias file:
```json
{
"extends": ["./tsconfig.alias.json"]
}
```
### Using Alias in Configs
Reference the centralized alias in all config files:
```ts
// vite.config.ts
import { alias } from './alias';
export default defineConfig({
resolve: { alias },
});
```
```ts
// nuxt.config.ts
import { alias } from './alias';
export default defineNuxtConfig({
alias,
});
```
@@ -0,0 +1,119 @@
---
name: setting-up
description: Project setup files including .gitignore, GitHub Actions workflows, and VS Code extensions. Use when initializing new projects or adding CI/editor config.
---
# Project Setup
## .gitignore
Create when `.gitignore` is not present:
```
*.log
*.tgz
.cache
.DS_Store
.eslintcache
.idea
.env
.nuxt
.temp
.output
.turbo
cache
coverage
dist
lib-cov
logs
node_modules
temp
```
## GitHub Actions
Add these workflows when setting up a new project. Skip if workflows already exist. All use [sxzz/workflows](https://github.com/sxzz/workflows) reusable workflows.
### Autofix Workflow
**`.github/workflows/autofix.yml`** - Auto-fix linting on PRs:
```yaml
name: autofix.ci
on: [pull_request]
jobs:
autofix:
uses: sxzz/workflows/.github/workflows/autofix.yml@v1
permissions:
contents: read
```
### Unit Test Workflow
**`.github/workflows/unit-test.yml`** - Run tests on push/PR:
```yaml
name: Unit Test
on:
push:
branches: [main]
pull_request:
branches: [main]
permissions: {}
jobs:
unit-test:
uses: sxzz/workflows/.github/workflows/unit-test.yml@v1
```
### Release Workflow
**`.github/workflows/release.yml`** - Publish on tag (library projects only):
```yaml
name: Release
on:
push:
tags:
- 'v*'
jobs:
release:
uses: sxzz/workflows/.github/workflows/release.yml@v1
with:
publish: true
permissions:
contents: write
id-token: write
```
## VS Code Extensions
Configure in `.vscode/extensions.json`:
```json
{
"recommendations": [
"dbaeumer.vscode-eslint",
"antfu.pnpm-catalog-lens",
"antfu.iconify",
"antfu.unocss",
"antfu.slidev",
"vue.volar"
]
}
```
| Extension | Description |
| ------------------------- | --------------------------------------------- |
| `dbaeumer.vscode-eslint` | ESLint integration for linting and formatting |
| `antfu.pnpm-catalog-lens` | Shows pnpm catalog version hints inline |
| `antfu.iconify` | Iconify icon preview and autocomplete |
| `antfu.unocss` | UnoCSS IntelliSense and syntax highlighting |
| `antfu.slidev` | Slidev preview and syntax highlighting |
| `vue.volar` | Vue Language Features |
@@ -0,0 +1,494 @@
---
name: architecture-patterns
description: Implement proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design. Use when architecting complex backend systems or refactoring existing applications for better maintainability.
---
# Architecture Patterns
Master proven backend architecture patterns including Clean Architecture, Hexagonal Architecture, and Domain-Driven Design to build maintainable, testable, and scalable systems.
## When to Use This Skill
- Designing new backend systems from scratch
- Refactoring monolithic applications for better maintainability
- Establishing architecture standards for your team
- Migrating from tightly coupled to loosely coupled architectures
- Implementing domain-driven design principles
- Creating testable and mockable codebases
- Planning microservices decomposition
## Core Concepts
### 1. Clean Architecture (Uncle Bob)
**Layers (dependency flows inward):**
- **Entities**: Core business models
- **Use Cases**: Application business rules
- **Interface Adapters**: Controllers, presenters, gateways
- **Frameworks & Drivers**: UI, database, external services
**Key Principles:**
- Dependencies point inward
- Inner layers know nothing about outer layers
- Business logic independent of frameworks
- Testable without UI, database, or external services
### 2. Hexagonal Architecture (Ports and Adapters)
**Components:**
- **Domain Core**: Business logic
- **Ports**: Interfaces defining interactions
- **Adapters**: Implementations of ports (database, REST, message queue)
**Benefits:**
- Swap implementations easily (mock for testing)
- Technology-agnostic core
- Clear separation of concerns
### 3. Domain-Driven Design (DDD)
**Strategic Patterns:**
- **Bounded Contexts**: Separate models for different domains
- **Context Mapping**: How contexts relate
- **Ubiquitous Language**: Shared terminology
**Tactical Patterns:**
- **Entities**: Objects with identity
- **Value Objects**: Immutable objects defined by attributes
- **Aggregates**: Consistency boundaries
- **Repositories**: Data access abstraction
- **Domain Events**: Things that happened
## Clean Architecture Pattern
### Directory Structure
```
app/
├── domain/ # Entities & business rules
│ ├── entities/
│ │ ├── user.py
│ │ └── order.py
│ ├── value_objects/
│ │ ├── email.py
│ │ └── money.py
│ └── interfaces/ # Abstract interfaces
│ ├── user_repository.py
│ └── payment_gateway.py
├── use_cases/ # Application business rules
│ ├── create_user.py
│ ├── process_order.py
│ └── send_notification.py
├── adapters/ # Interface implementations
│ ├── repositories/
│ │ ├── postgres_user_repository.py
│ │ └── redis_cache_repository.py
│ ├── controllers/
│ │ └── user_controller.py
│ └── gateways/
│ ├── stripe_payment_gateway.py
│ └── sendgrid_email_gateway.py
└── infrastructure/ # Framework & external concerns
├── database.py
├── config.py
└── logging.py
```
### Implementation Example
```python
# domain/entities/user.py
from dataclasses import dataclass
from datetime import datetime
from typing import Optional
@dataclass
class User:
"""Core user entity - no framework dependencies."""
id: str
email: str
name: str
created_at: datetime
is_active: bool = True
def deactivate(self):
"""Business rule: deactivating user."""
self.is_active = False
def can_place_order(self) -> bool:
"""Business rule: active users can order."""
return self.is_active
# domain/interfaces/user_repository.py
from abc import ABC, abstractmethod
from typing import Optional, List
from domain.entities.user import User
class IUserRepository(ABC):
"""Port: defines contract, no implementation."""
@abstractmethod
async def find_by_id(self, user_id: str) -> Optional[User]:
pass
@abstractmethod
async def find_by_email(self, email: str) -> Optional[User]:
pass
@abstractmethod
async def save(self, user: User) -> User:
pass
@abstractmethod
async def delete(self, user_id: str) -> bool:
pass
# use_cases/create_user.py
from domain.entities.user import User
from domain.interfaces.user_repository import IUserRepository
from dataclasses import dataclass
from datetime import datetime
import uuid
@dataclass
class CreateUserRequest:
email: str
name: str
@dataclass
class CreateUserResponse:
user: User
success: bool
error: Optional[str] = None
class CreateUserUseCase:
"""Use case: orchestrates business logic."""
def __init__(self, user_repository: IUserRepository):
self.user_repository = user_repository
async def execute(self, request: CreateUserRequest) -> CreateUserResponse:
# Business validation
existing = await self.user_repository.find_by_email(request.email)
if existing:
return CreateUserResponse(
user=None,
success=False,
error="Email already exists"
)
# Create entity
user = User(
id=str(uuid.uuid4()),
email=request.email,
name=request.name,
created_at=datetime.now(),
is_active=True
)
# Persist
saved_user = await self.user_repository.save(user)
return CreateUserResponse(
user=saved_user,
success=True
)
# adapters/repositories/postgres_user_repository.py
from domain.interfaces.user_repository import IUserRepository
from domain.entities.user import User
from typing import Optional
import asyncpg
class PostgresUserRepository(IUserRepository):
"""Adapter: PostgreSQL implementation."""
def __init__(self, pool: asyncpg.Pool):
self.pool = pool
async def find_by_id(self, user_id: str) -> Optional[User]:
async with self.pool.acquire() as conn:
row = await conn.fetchrow(
"SELECT * FROM users WHERE id = $1", user_id
)
return self._to_entity(row) if row else None
async def find_by_email(self, email: str) -> Optional[User]:
async with self.pool.acquire() as conn:
row = await conn.fetchrow(
"SELECT * FROM users WHERE email = $1", email
)
return self._to_entity(row) if row else None
async def save(self, user: User) -> User:
async with self.pool.acquire() as conn:
await conn.execute(
"""
INSERT INTO users (id, email, name, created_at, is_active)
VALUES ($1, $2, $3, $4, $5)
ON CONFLICT (id) DO UPDATE
SET email = $2, name = $3, is_active = $5
""",
user.id, user.email, user.name, user.created_at, user.is_active
)
return user
async def delete(self, user_id: str) -> bool:
async with self.pool.acquire() as conn:
result = await conn.execute(
"DELETE FROM users WHERE id = $1", user_id
)
return result == "DELETE 1"
def _to_entity(self, row) -> User:
"""Map database row to entity."""
return User(
id=row["id"],
email=row["email"],
name=row["name"],
created_at=row["created_at"],
is_active=row["is_active"]
)
# adapters/controllers/user_controller.py
from fastapi import APIRouter, Depends, HTTPException
from use_cases.create_user import CreateUserUseCase, CreateUserRequest
from pydantic import BaseModel
router = APIRouter()
class CreateUserDTO(BaseModel):
email: str
name: str
@router.post("/users")
async def create_user(
dto: CreateUserDTO,
use_case: CreateUserUseCase = Depends(get_create_user_use_case)
):
"""Controller: handles HTTP concerns only."""
request = CreateUserRequest(email=dto.email, name=dto.name)
response = await use_case.execute(request)
if not response.success:
raise HTTPException(status_code=400, detail=response.error)
return {"user": response.user}
```
## Hexagonal Architecture Pattern
```python
# Core domain (hexagon center)
class OrderService:
"""Domain service - no infrastructure dependencies."""
def __init__(
self,
order_repository: OrderRepositoryPort,
payment_gateway: PaymentGatewayPort,
notification_service: NotificationPort
):
self.orders = order_repository
self.payments = payment_gateway
self.notifications = notification_service
async def place_order(self, order: Order) -> OrderResult:
# Business logic
if not order.is_valid():
return OrderResult(success=False, error="Invalid order")
# Use ports (interfaces)
payment = await self.payments.charge(
amount=order.total,
customer=order.customer_id
)
if not payment.success:
return OrderResult(success=False, error="Payment failed")
order.mark_as_paid()
saved_order = await self.orders.save(order)
await self.notifications.send(
to=order.customer_email,
subject="Order confirmed",
body=f"Order {order.id} confirmed"
)
return OrderResult(success=True, order=saved_order)
# Ports (interfaces)
class OrderRepositoryPort(ABC):
@abstractmethod
async def save(self, order: Order) -> Order:
pass
class PaymentGatewayPort(ABC):
@abstractmethod
async def charge(self, amount: Money, customer: str) -> PaymentResult:
pass
class NotificationPort(ABC):
@abstractmethod
async def send(self, to: str, subject: str, body: str):
pass
# Adapters (implementations)
class StripePaymentAdapter(PaymentGatewayPort):
"""Primary adapter: connects to Stripe API."""
def __init__(self, api_key: str):
self.stripe = stripe
self.stripe.api_key = api_key
async def charge(self, amount: Money, customer: str) -> PaymentResult:
try:
charge = self.stripe.Charge.create(
amount=amount.cents,
currency=amount.currency,
customer=customer
)
return PaymentResult(success=True, transaction_id=charge.id)
except stripe.error.CardError as e:
return PaymentResult(success=False, error=str(e))
class MockPaymentAdapter(PaymentGatewayPort):
"""Test adapter: no external dependencies."""
async def charge(self, amount: Money, customer: str) -> PaymentResult:
return PaymentResult(success=True, transaction_id="mock-123")
```
## Domain-Driven Design Pattern
```python
# Value Objects (immutable)
from dataclasses import dataclass
from typing import Optional
@dataclass(frozen=True)
class Email:
"""Value object: validated email."""
value: str
def __post_init__(self):
if "@" not in self.value:
raise ValueError("Invalid email")
@dataclass(frozen=True)
class Money:
"""Value object: amount with currency."""
amount: int # cents
currency: str
def add(self, other: "Money") -> "Money":
if self.currency != other.currency:
raise ValueError("Currency mismatch")
return Money(self.amount + other.amount, self.currency)
# Entities (with identity)
class Order:
"""Entity: has identity, mutable state."""
def __init__(self, id: str, customer: Customer):
self.id = id
self.customer = customer
self.items: List[OrderItem] = []
self.status = OrderStatus.PENDING
self._events: List[DomainEvent] = []
def add_item(self, product: Product, quantity: int):
"""Business logic in entity."""
item = OrderItem(product, quantity)
self.items.append(item)
self._events.append(ItemAddedEvent(self.id, item))
def total(self) -> Money:
"""Calculated property."""
return sum(item.subtotal() for item in self.items)
def submit(self):
"""State transition with business rules."""
if not self.items:
raise ValueError("Cannot submit empty order")
if self.status != OrderStatus.PENDING:
raise ValueError("Order already submitted")
self.status = OrderStatus.SUBMITTED
self._events.append(OrderSubmittedEvent(self.id))
# Aggregates (consistency boundary)
class Customer:
"""Aggregate root: controls access to entities."""
def __init__(self, id: str, email: Email):
self.id = id
self.email = email
self._addresses: List[Address] = []
self._orders: List[str] = [] # Order IDs, not full objects
def add_address(self, address: Address):
"""Aggregate enforces invariants."""
if len(self._addresses) >= 5:
raise ValueError("Maximum 5 addresses allowed")
self._addresses.append(address)
@property
def primary_address(self) -> Optional[Address]:
return next((a for a in self._addresses if a.is_primary), None)
# Domain Events
@dataclass
class OrderSubmittedEvent:
order_id: str
occurred_at: datetime = field(default_factory=datetime.now)
# Repository (aggregate persistence)
class OrderRepository:
"""Repository: persist/retrieve aggregates."""
async def find_by_id(self, order_id: str) -> Optional[Order]:
"""Reconstitute aggregate from storage."""
pass
async def save(self, order: Order):
"""Persist aggregate and publish events."""
await self._persist(order)
await self._publish_events(order._events)
order._events.clear()
```
## Resources
- **references/clean-architecture-guide.md**: Detailed layer breakdown
- **references/hexagonal-architecture-guide.md**: Ports and adapters patterns
- **references/ddd-tactical-patterns.md**: Entities, value objects, aggregates
- **assets/clean-architecture-template/**: Complete project structure
- **assets/ddd-examples/**: Domain modeling examples
## Best Practices
1. **Dependency Rule**: Dependencies always point inward
2. **Interface Segregation**: Small, focused interfaces
3. **Business Logic in Domain**: Keep frameworks out of core
4. **Test Independence**: Core testable without infrastructure
5. **Bounded Contexts**: Clear domain boundaries
6. **Ubiquitous Language**: Consistent terminology
7. **Thin Controllers**: Delegate to use cases
8. **Rich Domain Models**: Behavior with data
## Common Pitfalls
- **Anemic Domain**: Entities with only data, no behavior
- **Framework Coupling**: Business logic depends on frameworks
- **Fat Controllers**: Business logic in controllers
- **Repository Leakage**: Exposing ORM objects
- **Missing Abstractions**: Concrete dependencies in core
- **Over-Engineering**: Clean architecture for simple CRUD
@@ -0,0 +1,174 @@
---
name: better-auth-best-practices
description: Skill for integrating Better Auth - the comprehensive TypeScript authentication framework.
---
# Better Auth Integration Guide
**Always consult [better-auth.com/docs](https://better-auth.com/docs) for code examples and latest API.**
Better Auth is a TypeScript-first, framework-agnostic auth framework supporting email/password, OAuth, magic links, passkeys, and more via plugins.
---
## Quick Reference
### Environment Variables
- `BETTER_AUTH_SECRET` - Encryption secret (min 32 chars). Generate: `openssl rand -base64 32`
- `BETTER_AUTH_URL` - Base URL (e.g., `https://example.com`)
Only define `baseURL`/`secret` in config if env vars are NOT set.
### File Location
CLI looks for `auth.ts` in: `./`, `./lib`, `./utils`, or under `./src`. Use `--config` for custom path.
### CLI Commands
- `npx @better-auth/cli@latest migrate` - Apply schema (built-in adapter)
- `npx @better-auth/cli@latest generate` - Generate schema for Prisma/Drizzle
- `npx @better-auth/cli mcp --cursor` - Add MCP to AI tools
**Re-run after adding/changing plugins.**
---
## Core Config Options
| Option | Notes |
| ------------------ | ---------------------------------------------- |
| `appName` | Optional display name |
| `baseURL` | Only if `BETTER_AUTH_URL` not set |
| `basePath` | Default `/api/auth`. Set `/` for root. |
| `secret` | Only if `BETTER_AUTH_SECRET` not set |
| `database` | Required for most features. See adapters docs. |
| `secondaryStorage` | Redis/KV for sessions & rate limits |
| `emailAndPassword` | `{ enabled: true }` to activate |
| `socialProviders` | `{ google: { clientId, clientSecret }, ... }` |
| `plugins` | Array of plugins |
| `trustedOrigins` | CSRF whitelist |
---
## Database
**Direct connections:** Pass `pg.Pool`, `mysql2` pool, `better-sqlite3`, or `bun:sqlite` instance.
**ORM adapters:** Import from `better-auth/adapters/drizzle`, `better-auth/adapters/prisma`, `better-auth/adapters/mongodb`.
**Critical:** Better Auth uses adapter model names, NOT underlying table names. If Prisma model is `User` mapping to table `users`, use `modelName: "user"` (Prisma reference), not `"users"`.
---
## Session Management
**Storage priority:**
1. If `secondaryStorage` defined → sessions go there (not DB)
2. Set `session.storeSessionInDatabase: true` to also persist to DB
3. No database + `cookieCache` → fully stateless mode
**Cookie cache strategies:**
- `compact` (default) - Base64url + HMAC. Smallest.
- `jwt` - Standard JWT. Readable but signed.
- `jwe` - Encrypted. Maximum security.
**Key options:** `session.expiresIn` (default 7 days), `session.updateAge` (refresh interval), `session.cookieCache.maxAge`, `session.cookieCache.version` (change to invalidate all sessions).
---
## User & Account Config
**User:** `user.modelName`, `user.fields` (column mapping), `user.additionalFields`, `user.changeEmail.enabled` (disabled by default), `user.deleteUser.enabled` (disabled by default).
**Account:** `account.modelName`, `account.accountLinking.enabled`, `account.storeAccountCookie` (for stateless OAuth).
**Required for registration:** `email` and `name` fields.
---
## Email Flows
- `emailVerification.sendVerificationEmail` - Must be defined for verification to work
- `emailVerification.sendOnSignUp` / `sendOnSignIn` - Auto-send triggers
- `emailAndPassword.sendResetPassword` - Password reset email handler
---
## Security
**In `advanced`:**
- `useSecureCookies` - Force HTTPS cookies
- `disableCSRFCheck` - ⚠️ Security risk
- `disableOriginCheck` - ⚠️ Security risk
- `crossSubDomainCookies.enabled` - Share cookies across subdomains
- `ipAddress.ipAddressHeaders` - Custom IP headers for proxies
- `database.generateId` - Custom ID generation or `"serial"`/`"uuid"`/`false`
**Rate limiting:** `rateLimit.enabled`, `rateLimit.window`, `rateLimit.max`, `rateLimit.storage` ("memory" | "database" | "secondary-storage").
---
## Hooks
**Endpoint hooks:** `hooks.before` / `hooks.after` - Array of `{ matcher, handler }`. Use `createAuthMiddleware`. Access `ctx.path`, `ctx.context.returned` (after), `ctx.context.session`.
**Database hooks:** `databaseHooks.user.create.before/after`, same for `session`, `account`. Useful for adding default values or post-creation actions.
**Hook context (`ctx.context`):** `session`, `secret`, `authCookies`, `password.hash()`/`verify()`, `adapter`, `internalAdapter`, `generateId()`, `tables`, `baseURL`.
---
## Plugins
**Import from dedicated paths for tree-shaking:**
```
import { twoFactor } from "better-auth/plugins/two-factor"
```
NOT `from "better-auth/plugins"`.
**Popular plugins:** `twoFactor`, `organization`, `passkey`, `magicLink`, `emailOtp`, `username`, `phoneNumber`, `admin`, `apiKey`, `bearer`, `jwt`, `multiSession`, `sso`, `oauthProvider`, `oidcProvider`, `openAPI`, `genericOAuth`.
Client plugins go in `createAuthClient({ plugins: [...] })`.
---
## Client
Import from: `better-auth/client` (vanilla), `better-auth/react`, `better-auth/vue`, `better-auth/svelte`, `better-auth/solid`.
Key methods: `signUp.email()`, `signIn.email()`, `signIn.social()`, `signOut()`, `useSession()`, `getSession()`, `revokeSession()`, `revokeSessions()`.
---
## Type Safety
Infer types: `typeof auth.$Infer.Session`, `typeof auth.$Infer.Session.user`.
For separate client/server projects: `createAuthClient<typeof auth>()`.
---
## Common Gotchas
1. **Model vs table name** - Config uses ORM model name, not DB table name
2. **Plugin schema** - Re-run CLI after adding plugins
3. **Secondary storage** - Sessions go there by default, not DB
4. **Cookie cache** - Custom session fields NOT cached, always re-fetched
5. **Stateless mode** - No DB = session in cookie only, logout on cache expiry
6. **Change email flow** - Sends to current email first, then new email
---
## Resources
- [Docs](https://better-auth.com/docs)
- [Options Reference](https://better-auth.com/docs/reference/options)
- [LLMs.txt](https://better-auth.com/llms.txt)
- [GitHub](https://github.com/better-auth/better-auth)
- [Init Options Source](https://github.com/better-auth/better-auth/blob/main/packages/core/src/types/init-options.ts)
@@ -0,0 +1,101 @@
---
name: brainstorming
description: 'You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.'
---
# Brainstorming Ideas Into Designs
## Overview
Help turn ideas into fully formed designs and specs through natural collaborative dialogue.
Start by understanding the current project context, then ask questions one at a time to refine the idea. Once you understand what you're building, present the design and get user approval.
<HARD-GATE>
Do NOT invoke any implementation skill, write any code, scaffold any project, or take any implementation action until you have presented a design and the user has approved it. This applies to EVERY project regardless of perceived simplicity.
</HARD-GATE>
## Anti-Pattern: "This Is Too Simple To Need A Design"
Every project goes through this process. A todo list, a single-function utility, a config change — all of them. "Simple" projects are where unexamined assumptions cause the most wasted work. The design can be short (a few sentences for truly simple projects), but you MUST present it and get approval.
## Checklist
You MUST create a task for each of these items and complete them in order:
1. **Explore project context** — check files, docs, recent commits
2. **Ask clarifying questions** — one at a time, understand purpose/constraints/success criteria
3. **Propose 2-3 approaches** — with trade-offs and your recommendation
4. **Present design** — in sections scaled to their complexity, get user approval after each section
5. **Write design doc** — save to `docs/plans/YYYY-MM-DD-<topic>-design.md` and commit
6. **Transition to implementation** — invoke writing-plans skill to create implementation plan
## Process Flow
```dot
digraph brainstorming {
"Explore project context" [shape=box];
"Ask clarifying questions" [shape=box];
"Propose 2-3 approaches" [shape=box];
"Present design sections" [shape=box];
"User approves design?" [shape=diamond];
"Write design doc" [shape=box];
"Invoke writing-plans skill" [shape=doublecircle];
"Explore project context" -> "Ask clarifying questions";
"Ask clarifying questions" -> "Propose 2-3 approaches";
"Propose 2-3 approaches" -> "Present design sections";
"Present design sections" -> "User approves design?";
"User approves design?" -> "Present design sections" [label="no, revise"];
"User approves design?" -> "Write design doc" [label="yes"];
"Write design doc" -> "Invoke writing-plans skill";
}
```
**The terminal state is invoking writing-plans.** Do NOT invoke frontend-design, mcp-builder, or any other implementation skill. The ONLY skill you invoke after brainstorming is writing-plans.
## The Process
**Understanding the idea:**
- Check out the current project state first (files, docs, recent commits)
- Ask questions one at a time to refine the idea
- Prefer multiple choice questions when possible, but open-ended is fine too
- Only one question per message - if a topic needs more exploration, break it into multiple questions
- Focus on understanding: purpose, constraints, success criteria
**Exploring approaches:**
- Propose 2-3 different approaches with trade-offs
- Present options conversationally with your recommendation and reasoning
- Lead with your recommended option and explain why
**Presenting the design:**
- Once you believe you understand what you're building, present the design
- Scale each section to its complexity: a few sentences if straightforward, up to 200-300 words if nuanced
- Ask after each section whether it looks right so far
- Cover: architecture, components, data flow, error handling, testing
- Be ready to go back and clarify if something doesn't make sense
## After the Design
**Documentation:**
- Write the validated design to `docs/plans/YYYY-MM-DD-<topic>-design.md`
- Use elements-of-style:writing-clearly-and-concisely skill if available
- Commit the design document to git
**Implementation:**
- Invoke the writing-plans skill to create a detailed implementation plan
- Do NOT invoke any other skill. writing-plans is the next step.
## Key Principles
- **One question at a time** - Don't overwhelm with multiple questions
- **Multiple choice preferred** - Easier to answer than open-ended when possible
- **YAGNI ruthlessly** - Remove unnecessary features from all designs
- **Explore alternatives** - Always propose 2-3 approaches before settling
- **Incremental validation** - Present design, get approval before moving on
- **Be flexible** - Go back and clarify when something doesn't make sense
@@ -0,0 +1,202 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [yyyy] [name of copyright owner]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
@@ -0,0 +1,73 @@
---
name: brand-guidelines
description: Applies Anthropic's official brand colors and typography to any sort of artifact that may benefit from having Anthropic's look-and-feel. Use it when brand colors or style guidelines, visual formatting, or company design standards apply.
license: Complete terms in LICENSE.txt
---
# Anthropic Brand Styling
## Overview
To access Anthropic's official brand identity and style resources, use this skill.
**Keywords**: branding, corporate identity, visual identity, post-processing, styling, brand colors, typography, Anthropic brand, visual formatting, visual design
## Brand Guidelines
### Colors
**Main Colors:**
- Dark: `#141413` - Primary text and dark backgrounds
- Light: `#faf9f5` - Light backgrounds and text on dark
- Mid Gray: `#b0aea5` - Secondary elements
- Light Gray: `#e8e6dc` - Subtle backgrounds
**Accent Colors:**
- Orange: `#d97757` - Primary accent
- Blue: `#6a9bcc` - Secondary accent
- Green: `#788c5d` - Tertiary accent
### Typography
- **Headings**: Poppins (with Arial fallback)
- **Body Text**: Lora (with Georgia fallback)
- **Note**: Fonts should be pre-installed in your environment for best results
## Features
### Smart Font Application
- Applies Poppins font to headings (24pt and larger)
- Applies Lora font to body text
- Automatically falls back to Arial/Georgia if custom fonts unavailable
- Preserves readability across all systems
### Text Styling
- Headings (24pt+): Poppins font
- Body text: Lora font
- Smart color selection based on background
- Preserves text hierarchy and formatting
### Shape and Accent Colors
- Non-text shapes use accent colors
- Cycles through orange, blue, and green accents
- Maintains visual interest while staying on-brand
## Technical Details
### Font Management
- Uses system-installed Poppins and Lora fonts when available
- Provides automatic fallback to Arial (headings) and Georgia (body)
- No font installation required - works with existing system fonts
- For best results, pre-install Poppins and Lora fonts in your environment
### Color Application
- Uses RGB color values for precise brand matching
- Applied via python-pptx's RGBColor class
- Maintains color fidelity across different systems

Some files were not shown because too many files have changed in this diff Show More