Compare commits

...
Author SHA1 Message Date
jason.woltje 87f10772ce Merge M13: interactive TUI agent + TOOLS.md
Closes #35
2026-09-03 11:24:56 -05:00
jason.woltje 7db4c5c2ed feat(agent): interactive TUI launcher + identity + TOOLS.md (#35)
- scripts/agent.sh <name>: launches interactive pi TUI in the container
  with contracts + optional mission + agent identity + named session +
  optional workspace/tools; the Mosaic alternative to vanilla pi
- pi adapter: MOSAIC_INTERACTIVE branch (clean TUI, no -p, no initial
  prompt); headless exec rebuilt via positional args (no word-splitting
  on the request); MOSAIC_AGENT_NAME optional in headless
- loader: AGENT IDENTITY section when the launcher names the agent
- compose: fixed command removed (request defaults live in run-agent.sh);
  MOSAIC_INTERACTIVE/MOSAIC_AGENT_NAME passthrough
- docs/TOOLS.md: full on-demand tool reference; AGENTS.md routes to it
- RELEASE -> 0.0.8 (container change); build verified

Closes #35
2026-09-03 11:24:56 -05:00
jason.woltje 0273a84549 docs: AGENTS.md - session recovery shim, invariants canon, session registry
- AGENTS.md at root: pi loads it automatically at every session start
  (conductor-level sessions; workers deliberately exclude it via
  --no-context-files). Deliberately short: invariants, session protocol,
  role model, command surface, data map, pointers - depth stays in docs/.
- docs/SESSIONS.md: append-only session registry, mandatory per session.
- Recovery rule encoded: compaction/restart loses nothing - AGENTS.md +
  CURRENT.md + git log + suites reconstruct state; never guess.
2026-09-03 11:02:19 -05:00
jason.woltje 3b674b7a66 Merge: roles/ directory convention - root is bootstrap-only 2026-09-03 10:56:07 -05:00
jason.woltje 527bc581ca refactor(layout): role contracts move to roles/ - root is bootstrap-only
Owner direction: the repository root holds first-class, bootstrap-required
configuration only. conductor-policy.json is a ROLE contract (the
conductor's authority), one of scores of future role contracts
(agent-policy, coder-policy, ...) - such files get a dedicated home.

- roles/conductor-policy.json (git mv)
- conductor-apply.sh + test-conductor.sh read the new path
- CONDUCTOR.md records the roles/ convention

Closes UX follow-up from owner layout review; no issue (convention change).
2026-09-03 10:56:07 -05:00
jason.woltje 1249714a9a docs(plan): CURRENT.md - M12 shipped 2026-09-03 07:04:37 -05:00
jason.woltje e175616885 feat(conductor): auto-apply policy gate for worker patches (#34)
- conductor-policy.json (tracked, strictly validated): enabled switch,
  path allowlist globs, gating suites - the autonomy decision lives in a
  declarative file the owner controls
- scripts/conductor-apply.sh <runId> [--dry-run]: succeeded-run check ->
  clean target tree -> diff from worker workspace -> allowlist -> syntax
  gates (node/bash/json) -> apply -> policy suites -> attribution commit;
  ANY failure reverts the tree; push is never automatic
- scripts/test-conductor.sh: 17 sandbox cases covering every gate incl.
  suite-failure auto-revert and disabled policy
- policy defaults: scripts/docs/tasks/missions/adapters + README; all
  three suites gate

Closes #34
2026-09-03 07:03:29 -05:00
jason.woltje 6955717612 Merge M11: session forking from a common ancestor
Closes #33
2026-09-03 06:43:02 -05:00
jason.woltje 88d9cf750f feat(sessions): sessionForkFrom - branch conversations from a common ancestor (#33)
- task schema: optional sessionForkFrom (source session name); requires
  session target; self-fork rejected
- runner: resolves source newest .jsonl (fail 4 if none/outside dataRoot);
  passes MOSAIC_SESSION_FORK + MOSAIC_SESSION_DIR; result records lineage
- pi adapter: --fork <source> --session-dir <target> when forking;
  ephemeral default unchanged; plain session resume unchanged
- compose passthrough; RELEASE -> 0.0.7 (adapter changed)
- suite +9 cases (58 total): plumbing via mock stderr, validation
  negatives, live fork - child recalls ancestor code word, ancestor
  session file untouched

Closes #33
2026-09-03 06:43:02 -05:00
jason.woltje c038706eed docs(plan): CURRENT.md - M10 shipped, retention next in review 2026-09-03 06:33:54 -05:00
jason.woltje 8622c9d826 Merge M10: run-record retention
Closes #32
2026-09-03 06:33:27 -05:00
jason.woltje 88eef507b0 feat(retention): run-record pruning - keep newest N, dry-run default (#32)
- mosaic-task.mjs prune [--keep=N] [--yes]: default keep 50; without
  --yes lists candidates without deleting
- only r-* directories under the runs root; symlinks skipped;
  sessions/workspaces/state/config untouched (asserted by suite sentinels)
- append-only receipt runs/.pruned.log records every pruned id
- test-task.sh: +8 retention cases (dry-run no-delete, keep-N, newest
  kept, receipt, isolation, invalid keep, empty no-op)

Also: suite hardening - prune section scopes its config per-command
(no export/unset leaking into later sections); duplicated check()
removed; latest_reason hoisted to helpers; status colors now green OK /
red FAIL (terminal-only, NO_COLOR-aware) per owner UX feedback.

Closes #32
2026-09-03 06:33:27 -05:00
jason.woltje 439bea6915 ui(test): green OK/PASS, red FAIL - terminal-only, NO_COLOR-aware
Owner feedback: grep match-highlighting made the word 'policy' red while
status words were plain - counter-indicative. Suites + verify now emit
ANSI colors (green success, red failure) when stdout is a terminal;
piped/machine-parsed output stays plain, honoring NO_COLOR. Word 'ok'
promoted to 'OK' for scannability.

Verified byte-level via forced-pty run; piped output unchanged; suites
41/24/14 + verify green.
2026-09-03 06:23:47 -05:00
jason.woltje fad8a4718c Merge M9: mission-level capability policy
Closes #30
2026-09-03 06:16:01 -05:00
jason.woltje 2ff49adff4 feat(policy): mission-level capability policy - least-privilege intersection (#30)
- mission schema: optional capabilities.tools (same validation as task)
- merge semantics in runTask: neither -> none; mission only -> mission;
  task only -> task; both -> intersection (task narrows, never widens);
  empty intersection -> tool-free run with an explicit stderr note
- result.json records EFFECTIVE tools; task/mission snapshots remain the
  immutable declaration of intent
- adapters unchanged; host-side only (no image change, 0.0.6 still active)
- task suite +5 cases (41 total): all four merge cases asserted from run
  evidence + invalid mission capabilities rejected

Policy decision recorded: missions govern; tasks cannot escalate.

Closes #30
2026-09-03 06:16:01 -05:00
jason.woltje 44c476ebbf fix(ops): show displays retriedFrom lineage (#29)
result.json recorded lineage correctly; the human-facing show command
omitted the field. Found by owner test: show | grep retriedFrom was
empty on a run whose result.json contained it.

Closes #29
2026-09-03 06:09:51 -05:00
jason.woltje cde480eb60 docs(plan): CURRENT.md — retry lineage shipped, M9 queued for decision 2026-09-03 05:31:24 -05:00
jason.woltje 5808248707 Merge retry lineage + relative mission resolution
Closes #28
2026-09-03 05:30:57 -05:00
jason.woltje afd5827db8 fix(retry): lineage tracking + relative mission path resolution (#28)
- retryRun rewrites a snapshot's relative mission path to the run's own
  recorded mission.json (absolute) before execution — retries stay
  faithful to what originally ran
- runTask accepts options.retriedFrom; retry records lineage in
  result.json (additive optional field, no schema break)
- task suite +4 cases: retry succeeds, lineage recorded, mission section
  present after retry (36 total), missing-run retry exits 4

Closes #28
2026-09-03 05:30:57 -05:00
jason.woltje d9cc990376 Merge M8: conductor loop - self-orchestration
Closes #25, closes #26, closes #27
2026-09-02 22:42:52 -05:00
jason.woltje 83c4e9851e feat(orchestration): retry <runId> — authored by headless pi worker (#26, #27)
Collaboration record (conductor loop, docs/plans/CONDUCTOR.md):
- round 1 (worker session worker-1, 2m28s): retry implemented per spec
- conductor live test exposed spec gap: direct invocation lacked
  launcher env exports
- round 2 (same worker session, 59s): spawnEnv made self-sufficient,
  but used PI_* where compose interpolates MOSAIC_*
- conductor hotfix: 3-line rename to MOSAIC_PROVIDER/MOSAIC_MODEL/
  MOSAIC_DATA_ROOT

Final: node scripts/mosaic-task.mjs retry <runId> re-executes a run's
task snapshot as a new run; live retry replied REMEMBERED; all suites
green (24/32/14 + verify).

Known limitation: retrying a run whose task used a RELATIVE mission path
resolves it against the temp dir; lineage tracking deferred.

Closes #25, closes #26, closes #27
2026-09-02 22:42:52 -05:00
jason.woltje 22508170a2 docs(plan): CURRENT.md — single next-action pointer for cadence-driven work 2026-09-02 22:25:56 -05:00
jason.woltje 90a67d050e Merge M7: operator ergonomics + release 0.0.6
Closes #24
2026-09-02 22:14:18 -05:00
jason.woltje 24bdef75fa feat(ops): run inspection, release 0.0.6, docs (#24)
- mosaic-task.mjs show <runId>: full record + snapshots + artifacts;
  uppercase-tolerant id validation; missing/traversal ids exit 4
- list: task/workspace/session columns
- RELEASE -> 0.0.6; README workspaces/capabilities/sessions sections;
  BUILD-LOG Phases 9-11; autonomous-run tracker results filled

Closes #24
2026-09-02 22:14:18 -05:00
jason.woltje 4e2a413640 Merge M6: named sessions - persistence and resume
Closes #22, closes #23
2026-09-02 22:08:34 -05:00
jason.woltje be55549700 feat(sessions): named persistent sessions with resume (L1) (#22, #23)
- task schema: optional session (named id) -> persistent session dir at
  dataRoot/sessions/<name>, isolated per name
- pi adapter: --session-dir when declared (ephemeral --no-session stays
  the default otherwise); -c resumes the most recent session when present
- compose passthrough; result.json records session
- fixtures: tasks/session-demo-1.json (teach) + session-demo-2.json (recall)
- E2E: teach -> REMEMBERED + host-side session JSONL; resume -> recalled
  'mosaico' exactly; single continued session file

Closes #22, closes #23
2026-09-02 22:08:34 -05:00
jason.woltje ddb1554e5b Merge M5: task workspaces + capability envelope
Closes #20, closes #21
2026-09-02 22:06:43 -05:00
jason.woltje b017e66e17 test(capabilities): workspace/tooling selftests + live demo fixture (#21)
- mock plumbing cases: workspace path + tools delivered (asserted from
  run-record stderr), host workspace created, absent fields = empty vars
- validation negatives: unknown tool, workspace traversal
- tasks/workspace-demo.json: pi uses bash inside the persistent demo
  workspace; host-visible proof.txt verified live

Closes #21
2026-09-02 22:06:43 -05:00
jason.woltje 172368612c feat(capabilities): task workspaces + tools allowlist plumbing (#20)
- task schema: optional workspace (absent | :run ephemeral | named
  persistent under dataRoot/workspaces) and capabilities.tools (pi
  documented tool allowlist); strict validation, traversal-proof names
- runner: creates host workspace, passes MOSAIC_WORKSPACE (container
  path) + MOSAIC_TOOLS; result.json records both
- pi adapter: cds into workspace; --tools when allowlist present else
  --no-tools
- mock adapter: logs delivered MOSAIC_* vars to stderr as deterministic
  plumbing evidence (dash prints 'export K=v', so use env not export)

Closes #20
2026-09-02 22:03:46 -05:00
jason.woltje 1387231e57 docs(plan): autonomous work run tracker (M5-M7 scope, test plan, review checklist) 2026-09-02 21:57:30 -05:00
jason.woltje 0292392e64 Merge M4: runtime adapter seam
Closes #16, closes #17, closes #18, closes #19
2026-09-02 21:30:35 -05:00
jason.woltje 594b8d711c docs(adapters): adapter seam docs + recorded M4 E2E, release 0.0.5 (#19)
- README: Runtime adapters section (contract summary, selection, mission
  injection point); BUILD-LOG Phase 8 entries
- E2E: 24+24+14 selftests green; verify PASS; 0.0.5 packaged and
  health-gated activated; mission-bearing fixture task succeeded through
  the real pi adapter; config checksum unchanged

Closes #19
2026-09-02 21:30:35 -05:00
jason.woltje 3c1ffd2c2d test(adapters): seam selftests — deterministic mock cases + mission injection (#18)
- test-config: absent adapter defaults to pi; mock validates; unknown
  adapter exits 2; env exports MOSAIC_ADAPTER (24 cases total)
- test-task: mock adapter gate pass/mismatch (no provider needed),
  mismatch reason asserted, unknown adapter fails closed, mission section
  injected into generated prompt asserted by content (24 cases total)
- harness fixes: helpers defined before use; per-case config files (no
  cross-case leakage); newest-run selection for the live case; deduped
  accidentally duplicated live block

Closes #18
2026-09-02 21:28:57 -05:00
jason.woltje 4ebb123ba3 feat(adapters): sanctioned mission directives injection (#17)
- load-contracts.sh: MOSAIC_MISSION_FILE (readable) appends a MISSION
  (runtime) section — objective + directives — after the immutable
  contracts; unreadable path is a hard error, absent env changes nothing
- mosaic-task.mjs: exports MOSAIC_MISSION_FILE as the run snapshot's
  container path (/var/lib/mosaic/runs/<id>/mission.json), with an
  outside-dataRoot guard; also exports the configured adapter

Verified: contract-only prompt has no mission section; mission-bearing
run shows objective + directives in the generated prompt, snapshot
recorded, real provider returns exactly MOSAIC_HELLO_OK.

Closes #17
2026-09-02 21:19:22 -05:00
jason.woltje bb5cecb348 feat(adapters): adapter contract, dispatch, pi + mock adapters (#16)
- adapters/README.md: the harness boundary contract (env in, response on
  stdout, diagnostics stderr, exit 0 success)
- adapters/pi: extracted current invocation unchanged
- adapters/mock: deterministic MOSAIC_MOCK_RESPONSE echo (test-only)
- run-agent.sh: name-validated dispatch to adapters/<name>/adapter.sh
- config: optional execution.adapter (pi|mock), default pi, configVersion
  stays 1 — existing configs remain valid; selection authority is the
  config file (load_config exports it)
- compose: MOSAIC_ADAPTER / MOSAIC_MOCK_RESPONSE passthrough; Containerfile
  installs adapters read-only; RELEASE -> 0.0.5

Verified: hello unchanged; mock verbatim via config; unknown adapter and
path-traversal names refused in-container; invalid adapter exits 2.

Closes #16
2026-09-02 21:18:07 -05:00
jason.woltje 88f55d9135 test(task): live failures self-report evidence; wrong-exit no longer masks (#15)
- live hello failure dumps latest run result.json + stderr tail before
  sandbox cleanup destroys them
- wrong-expectExact case asserts reason == expect-mismatch (was: any
  exit 1, which masked compose-level failures)
- repair dangling if/else from the docker-guard refactor

Closes #15
2026-09-02 20:56:58 -05:00
jason.woltje e7e1bd26eb fix(launcher): resolve release identity for direct task runs (#14)
M3 made MOSAIC_IMAGE_TAG required in compose, but run-task.sh never
called load_release — direct task runs failed in compose before any
model call. release.sh paths masked it by exporting the tag to children.

Found by owner-run test-task.sh; failure receipts were in the run
records' stderr.txt.

Closes #14
2026-09-02 20:44:31 -05:00
jason.woltje 5d86b8fa93 Merge M3: release model and safe updates
Closes #10, closes #11, closes #12, closes #13
2026-09-02 20:24:33 -05:00
jason.woltje 35ea464661 docs(release): release model usage + recorded M3 drills (#13)
- README: Release model section (package/activate/rollback/status,
  pointer + append-only log, gate-then-flip guarantee)
- BUILD-LOG Phase 7: drills recorded (update, refusal, rollback),
  two harness/product corrections documented

Drill evidence: 0.0.3 -> 0.0.4 update with unchanged config checksum and
green verify; fault-injected refusal left pointer untouched; health-gated
rollback restored 0.0.3; full append-only event history.

Closes #13
2026-09-02 20:24:33 -05:00
jason.woltje e87ecdb3e5 test(release): release-layer selftests (#12)
14 cases: RELEASE validation (valid/invalid/missing), tag consistency,
status on empty state, fault-injected refusal with no pointer + single
valid refusal log line, healthy activation, pointer fields, repeat
activation append-only log, rollback-without-previous refusal.

Harness fix learned the hard way: restore RELEASE from backup inline
after the missing-file case (mv-back restored the mutated file); single
exit trap self-heals the repo state.

Closes #12
2026-09-02 20:23:18 -05:00
jason.woltje a947db7bfd feat(release): package/activate/rollback/status with health gate (#11)
- activate: image-presence pre-check + M2 task-runner health gate
  (tasks/hello-marker.json exact marker) before atomic pointer replace
  (tmp+rename); every attempt appended to activation-log.jsonl
- --fault-injection flips the health expectation to prove the refusal path
- rollback: health-gated re-activation of the previous activated imageTag
  from the log; refuses when the image is gone or no previous exists
- status: release, tag, pointer, recent log; safe on empty state
- state lives under <dataRoot>/state/ (config-independent, reset-scoped)

Verified: activate OK; fault-injected refuse with pointer unchanged;
rollback-without-previous refuse.

Closes #11
2026-09-02 20:20:32 -05:00
jason.woltje a35ea62ab1 feat(release): RELEASE identity + image tag single-sourcing (#10)
- RELEASE file: single source of release version (0.0.X until declared stable)
- common.sh load_release(): validates version, derives
  MOSAIC_IMAGE_TAG=mosaic-poc-agent:<pi>-r<release> from the pinned pi dep
- compose.yaml: image tag is required env; build/hello/verify call load_release
- verify.sh derives the image name instead of hardcoding it
- package.json version aligned to the same 0.0.X line

Closes #10
2026-09-02 20:18:17 -05:00
46 changed files with 3004 additions and 72 deletions
+114
View File
@@ -0,0 +1,114 @@
# AGENTS.md — Mosaic Stack rebuild (`mosaicstack/stack-v2`)
Operational context for any agent session working in this repository.
Read top to bottom; it is deliberately short — depth lives in the files it
points to, not here.
## What this repository is
A standalone rebuild of Mosaic Stack: a file-based, fail-closed
orchestration foundation that dispatches sandboxed headless pi workers to do
real work, with immutable run records as evidence. Thirteen-plus tagged
milestones (`git tag -l`) from `poc-container-hello-v0` to today; suites
green at every step. Not production software — a proven foundation.
## Non-negotiable invariants (the canon)
1. **Root is bootstrap-only.** First-class system configuration lives at the
repository root; everything else gets a dedicated directory (`roles/`,
`contracts/`, `missions/`, `tasks/`, `docs/`). Do not add new files to root.
2. **Configuration**: `~/.config/mosaic-dev/config.json` is the sole system
config — created only by `scripts/bootstrap.sh`, never overwritten,
fail-closed on any problem. Repo-scoped role authority lives in
`roles/*.json` (versioned, reviewed commits only).
3. **Secrets** never enter the repository or container images; auth is
runtime-only (read-only mount or environment variable).
4. **Contracts** (`contracts/`) are immutable and image-baked. Missions and
tasks are declarative JSON with strict schemas.
5. **Run records** under `<dataRoot>/runs/` are write-once evidence — never
rewritten, only pruned via `prune` with a receipt.
6. **Fail closed**: missing or invalid config/policy refuses the operation.
Never improvise around a refusal; diagnose it.
7. **Policy**: missions govern tasks (least-privilege intersection — a task
narrows, never widens). Role authority is declared in `roles/` and changes
only via reviewed commits.
8. **Git**: commit only after suites are green; push only `main`; never
force-push. `scripts/conductor-apply.sh` commits locally — push stays an
explicit act.
9. **Append-only logs**: BUILD-LOG.md (phases), `activation-log.jsonl`,
`.pruned.log`, docs/SESSIONS.md. Corrections are new entries, never edits.
## Session protocol (mandatory)
- **Register** your session in `docs/SESSIONS.md` — one append-only line
(date, actor, scope, outcome). Never rewrite or remove entries.
- **Cadence**: read `docs/plans/CURRENT.md` → execute its single next action
fully (implement → test → verify against acceptance criteria → commit →
push → close issue) → update CURRENT.md → register in SESSIONS.md.
- "next" means one action. A batch mandate ("run the queue") repeats the
loop until green or blocked. Blocked means stop and report, never improvise.
- Substantial work gets a Gitea issue and a BUILD-LOG phase entry
(before/after, with corrections recorded honestly).
## Role model
- **Conductor**: a system-scoped role — not an agent, not a daemon. Holds
git/credentials/policy authority; decomposes, dispatches, reviews,
verifies, integrates. Protocol: `docs/plans/CONDUCTOR.md`. Exists only
when invoked; push is never automatic.
- **Workers**: headless pi via `scripts/run-task.sh` — sandboxed workspace,
tools allowlist, optional persistent sessions and forks; no git, no
credentials, no policy control.
- Worker runs deliberately exclude this file (`--no-context-files` in the
adapter): worker context is contracts + mission via the generated system
prompt. This file is for conductor-level sessions.
## Command surface
`scripts/bootstrap.sh` (idempotent) · `build.sh` · `hello.sh` ·
`verify.sh` · `run-task.sh run <task.json>` · `release.sh
package|activate|rollback|status` · `reset.sh` (**danger**: wipes the data
root; triple-safety-checked) · `mosaic-task.mjs validate|run|show|list|retry|prune` ·
`agent.sh <name>` (interactive TUI agent) ·
suites: `test-config.sh`, `test-task.sh`, `test-release.sh`,
`test-conductor.sh`.
Full reference — usage, fields, exit codes, safety notes:
`docs/TOOLS.md` (read on demand; do not rely on this summary for detail).
## Data map (canon)
- `~/.config/mosaic-dev/config.json` — system config (user-authored; never
auto-written).
- `<dataRoot>` (from config; default `~/.mosaic-dev`):
- `runs/` — write-once run evidence (`result.json`, snapshots, `stderr.txt`)
- `sessions/` — pi JSONL session trees, one directory per named session
- `workspaces/` — agent file effects (persistent or `:run` ephemeral)
- `state/` — release pointer + append-only activation/auto-apply logs
- Ownership is per-directory; nothing shares state. Directory map and
lifecycle rules: README.md "Data map" section.
## Pointers (depth lives here)
- `docs/plans/CURRENT.md` — THE next action (single source of "what now")
- `docs/plans/CONDUCTOR.md` — orchestration protocol and guardrails
- `docs/plans/2026-09-02_atomic-mosaic-foundation.md` — architecture, invariants
- `docs/plans/2026-09-03_autonomous-run.md` — batch-run tracker
- `BUILD-LOG.md` — append-only build/verification history with corrections
- `LAYERS.md` — implemented vs deferred layers
- `docs/SESSIONS.md` — session registry
- `adapters/README.md` — the harness adapter contract
- `roles/` — role contracts (conductor, future agent/coder/reviewer)
## Recovery rule
Compacted, restarted, or new? Nothing that matters is lost: this file +
`docs/plans/CURRENT.md` + `git log --oneline -10` + the suites reconstruct
the full state. **Never guess** — verify with the suites; the run records
and logs hold the receipts.
## Version pin
`@earendil-works/pi-coding-agent` is pinned exactly (see `package.json` /
`RELEASE`); never install unversioned. Release identity: `RELEASE` file
(0.0.X until declared stable); image tags derive from it.
+174
View File
@@ -163,4 +163,178 @@ Configuration-driven Hello World verified. `main` merged with M1 and tagged `con
Mission/task layer verified end-to-end. `main` merged with M2 and tagged `mission-task-v1`.
---
## Phase 7: Release model and safe updates (M3)
### Entry 7.1 — before
- Timestamp: 2026-09-03
- Intended action: Add the release substrate (Gitea milestone M3, issues #10-#13): RELEASE file single-sources the version (0.0.X line per owner direction), image tags derive from it, scripts/release.sh provides package/activate/rollback/status, activation is health-gated by the M2 task runner, pointer + append-only log under <dataRoot>/state/.
- Reason: The owner's top invariant — updates must never corrupt a working installation — needs a mechanism, not a convention: gate-then-flip with recorded history and rollback.
- Expected result: Update, refusal, and rollback drills all green with config checksums unchanged.
### Entry 7.2 — after
- Timestamp: 2026-09-03
- Commands run: scripts/test-release.sh (14 cases); recorded drills: update (0.0.3 -> 0.0.4 package+activate+verify), fault-injected refusal, rollback to 0.0.3.
- Observed result:
- Selftests: 14 passed, 0 failed.
- Update drill: packaged and activated r0.0.4 after exact-marker health gate; verify green under the new tag; config checksum unchanged.
- Refusal drill: health-gate fault injection -> activation refused (exit 1), pointer untouched, refusal appended to the log.
- Rollback drill: health-gated rollback to r0.0.3; pointer restored; log records package/activate/refused/rollback history append-only.
- Failure or correction:
1. release.sh initially failed with missing state/ directory (no mkdir before pointer/log writes); fixed.
2. Selftest harness mutated the repo RELEASE and restored the mutated copy (mv-back bug) plus a second trap replacing the first; fixed with inline backup restore and one self-healing exit trap. Product code unaffected.
- Credential check: no credential material in release state, logs, or drills.
## Result (M3)
Release model and safe updates verified by drills. `main` merged with M3 and tagged `release-model-v1`.
---
## Phase 8: Runtime adapter seam (M4)
### Entry 8.1 — before
- Timestamp: 2026-09-03
- Intended action: Formalize the harness boundary (Gitea milestone M4, issues #16-#19): documented adapter contract under /opt/mosaic/adapters/<name>/adapter.sh; run-agent.sh becomes a dispatcher; pi extracted unchanged; deterministic mock adapter for provider-free seam tests; config gains optional execution.adapter (default pi, configVersion unchanged); mission directives gain their sanctioned injection point via the run snapshot; RELEASE bumps to 0.0.5 with a health-gated activation.
- Reason: Future harnesses (Claude, Codex, OpenCode) must be additive — one directory each — and mission content needs a single sanctioned path into the runtime.
- Expected result: All suites green including new deterministic seam cases; 0.0.5 activated by health gate; mission-bearing run recorded.
### Entry 8.2 — after
- Timestamp: 2026-09-03
- Commands run: scripts/test-config.sh; scripts/test-task.sh; scripts/test-release.sh; manual seam drills (mock verbatim, unknown/traversal adapter refusal); mission injection checks; release package + activate for 0.0.5.
- Observed result:
- Config suite 24/24 (adapter default/validation/env export).
- Task suite 24/24 including deterministic mock cases (gate pass, expect-mismatch with reason, unknown adapter fail-closed) and mission injection asserted by prompt content.
- Release suite 14/14; image mosaic-poc-agent:0.84.4-r0.0.5 packaged and activated via exact-marker health gate.
- Mission directives now flow: task -> run snapshot -> container env -> generated prompt MISSION (runtime) section.
- Failure or correction:
1. Selection authority settled: load_config always exports MOSAIC_ADAPTER from config; environment overrides for scripts are therefore not a supported selection path (by design).
2. Selftest harness: three authoring defects fixed (helpers used before definition; one config file reused across cases leaking adapter state; a static mission fixture asserted against distinctive seam directives; plus an accidentally duplicated live block removed).
- Credential check: no credential material in adapters, prompts, run records, or logs.
## Result (M4)
Adapter seam verified; harness boundary is now additive by construction. `main` merged with M4 and tagged `adapter-seam-v1`.
---
## Phase 9: Workspaces + capability envelope (M5)
### Entry 9.1 — before
- Timestamp: 2026-09-03
- Intended action: Optional task workspace (absent / ":run" ephemeral / named persistent) and capabilities.tools allowlist (pi built-ins); runner plumbing via MOSAIC_WORKSPACE/MOSAIC_TOOLS; pi adapter maps to cwd + --tools; mock adapter logs delivered MOSAIC_* vars for deterministic assertions (Gitea #20, #21).
- Reason: Agents that only answer text cannot do work; the workspace+tools pair is the smallest real capability step, bounded by the container.
- Expected result: plumbing asserted via run-record stderr; live pi writes a host-visible file.
### Entry 9.2 — after
- Timestamp: 2026-09-03
- Commands run: build; mock plumbing run; validate negatives; full task suite; live workspace demo.
- Observed result: task suite 32/32; MOSAIC_WORKSPACE/MOSAIC_TOOLS asserted in run record; dataRoot/workspaces/<name> created host-side; live pi used bash to write proof.txt into the demo workspace (host-visible).
- Failure or correction: (1) batch edit dropped SUPPORTED_TOOLS const (runtime ReferenceError, exit 1 instead of 2) — restored; (2) mock env dump used `export` which this dash prints as `export K='v'` — switched to `env`; (3) selftest fed a mismatching mock response to the plain-task case — test bug, fixed.
## Phase 10: Named sessions (M6)
### Entry 10.1 — before
- Timestamp: 2026-09-03
- Intended action: Optional task.session name -> persistent session dir dataRoot/sessions/<name> via pi --session-dir; resume most recent with -c when present; isolation per name; teach/recall demo fixtures (Gitea #22, #23).
- Reason: L1 persistence is the prerequisite for any multi-step agent work.
- Expected result: session dir populated after first run; second run recalls taught context.
### Entry 10.2 — after
- Timestamp: 2026-09-03
- Commands run: build; mock plumbing run; live teach/recall E2E; task suite.
- Observed result: task suite 32/32; teach run replied REMEMBERED and session JSONL persisted host-side; recall run resumed (-c) and answered exactly 'mosaico'; single continued session file (no duplicate sessions).
- Failure or correction: none. Design note: ephemeral (--no-session) remains the default when no session is declared.
## Phase 11: Operator ergonomics (M7)
### Entry 11.1 — before
- Timestamp: 2026-09-03
- Intended action: mosaic-task.mjs show <runId> (full record + snapshots + artifacts, traversal-safe), list with workspace/session columns, RELEASE -> 0.0.6, package + health-gated activate, docs (Gitea #24).
- Reason: Run records are only as valuable as they are inspectable; release activation closes the loop on container-content changes.
- Expected result: show works for real/missing/traversal ids; suites green; 0.0.6 active.
### Entry 11.2 — after
- Timestamp: 2026-09-03
- Commands run: build; show on real/missing/traversal ids; full sweep; release package + activate.
- Observed result: config 24/24, task 32/32, release 14/14, verify PASS; 0.0.6 packaged and activated via exact-marker health gate.
- Failure or correction: showRun initially rejected valid run ids (lowercase-only regex vs uppercase timestamp) and crashed on missing ids (uncaught readdir) — both fixed and covered.
## Autonomous run result
M5 tagged `workspace-capabilities-v1`, M6 tagged `sessions-v1`, M7 tagged `operator-ergonomics-v1`; release 0.0.6 active. Tracker: docs/plans/2026-09-03_autonomous-run.md.
---
## Phase 12: Conductor loop — self-orchestration (M8)
### Entry 12.1 — before
- Timestamp: 2026-09-03
- Intended action: Stand up the poor-man orchestration loop per docs/plans/CONDUCTOR.md: conductor (host, holds git) mirrors the repo into a worker workspace, dispatches a headless pi worker (session worker-1, tools read/write/edit/bash) to implement retry <runId>, reviews the diff, integrates, verifies (Gitea #25, #26, #27).
- Reason: The owner asked for circular task processing with agent workers; the stack now has every primitive needed — this proves it on the stack itself.
- Expected result: worker-authored retry merged with suites green and a live retry verified.
### Entry 12.2 — after
- Timestamp: 2026-09-03
- Commands run: repo mirror clone; worker dispatch (tasks/worker-retry.json); diff review; apply; live retry; refinement dispatch (tasks/worker-retry-refine.json); second review; reverse+reapply combined patch; conductor interpolation fix; live retry; full sweep.
- Observed result:
- Worker round 1: implemented retry correctly per spec in 2m28s; diff reviewed clean.
- Live retry exposed a spec gap (direct invocation lacks launcher env exports).
- Worker round 2 (same session, 59s): made spawnEnv self-sufficient, but used PI_* names where compose interpolates MOSAIC_*.
- Conductor hotfix: 3-line rename to MOSAIC_PROVIDER/MOSAIC_MODEL/MOSAIC_DATA_ROOT (too trivial for a worker round).
- Final: live retry succeeded (replied REMEMBERED, new run recorded); all suites green.
- Failure or correction: three rounds total — one spec gap (conductor), one naming mismatch (worker), one trivial rename (conductor). Each was caught by mechanical verification (run record stderr), never by hope.
- Attribution: feature authored by headless pi worker (glm-5.3-flash) in sessions worker-1; conductor reviewed, integrated, and hotfixed.
## Result (M8)
Conductor loop proven end-to-end on the stack itself. `main` merged with M8; release 0.0.6 remains active (retry is host-side only, no image change).
---
## Phase 13: Mission-level capability policy (M9)
### Entry 13.1 — before
- Timestamp: 2026-09-03
- Intended action: Missions may declare capabilities.tools as governing constraints (Gitea #30); merge semantics = least-privilege intersection (task narrows, never widens; empty intersection = tool-free run). Host-side only.
- Reason: First mechanical restriction layer — the trust model becomes enforced, not instructed.
- Expected result: all four merge cases asserted from run evidence; suites green.
### Entry 13.2 — after
- Observed: four merge cases verified deterministically via run-record stderr (mission-only, task-only, narrowed, emptied); invalid mission capabilities exit 2; task suite 41/41.
- Failure or correction: selftest harness could not express ABSENT vs EMPTY fields via its printf helper — fixed with an ABSENT marker; two suite config-leak defects fixed (per-command env scoping). Product unaffected.
## Phase 14: Session forking (M11)
### Entry 14.1 — before
- Timestamp: 2026-09-03
- Intended action: sessionForkFrom task field branches the source session's newest file (pi --fork) into the target session dir; ancestor untouched; RELEASE -> 0.0.7 with health-gated activation (Gitea #33).
- Reason: Owner flagged conversation forking from a common ancestor as a desired property; pi JSONL trees make it native.
- Expected result: forked child recalls ancestor context; ancestor file untouched; suites green; 0.0.7 active.
### Entry 14.2 — after
- Observed: mock plumbing asserts fork source + target delivery; validation rejects fork-without-target and self-fork (exit 2); ghost source exits 4; live fork: child recalled 'mosaico' from ancestor context while the ancestor session file remained untouched (file-level assertion); suites 58/24/14 + verify green; 0.0.7 packaged and activated via health gate.
- Failure or correction: retryRun-style self-assignment bug in validation (compared null to target) — caught by negative test, fixed.
## Result (M11)
Session forking verified. `main` merged with M11, tagged `session-fork-v1`; release 0.0.7 active.
+1
View File
@@ -0,0 +1 @@
@AGENTS.md
+5 -2
View File
@@ -18,11 +18,14 @@ WORKDIR /opt/app
COPY package.json package-lock.json ./
RUN npm ci --ignore-scripts
# Immutable contract fixtures (required location) and runtime scripts.
# Immutable contract fixtures (required location), runtime scripts, and
# runtime adapters.
COPY contracts /opt/mosaic/contracts
COPY src /opt/mosaic/src
COPY adapters /opt/mosaic/adapters
RUN chmod 0555 /opt/mosaic/contracts /opt/mosaic/contracts/* \
&& chmod 0555 /opt/mosaic/src /opt/mosaic/src/*.sh
&& chmod 0555 /opt/mosaic/src /opt/mosaic/src/*.sh \
&& chmod 0555 /opt/mosaic/adapters /opt/mosaic/adapters/*/adapter.sh
# Writable state, workspace, and pi agent directory (auth.json is
# bind-mounted read-only at runtime; nothing is copied into the image).
+66 -2
View File
@@ -83,6 +83,67 @@ scripts/test-task.sh # selftests (schema negat
A run exits 0 only when its expectation is met (`expectExact` match); mismatches, nonzero agent exits, and timeouts record `status: failed` in `result.json` and exit 1. Each run gets a unique directory — rerunning never rewrites history.
## Release model (M3)
`RELEASE` single-sources the release version (0.0.X until declared stable); the image tag derives from it plus the pinned Pi version. Activation is health-gated and every event is recorded:
```bash
scripts/release.sh package # build + tag the release image
scripts/release.sh activate # health check (exact marker) -> atomic pointer swap
scripts/release.sh activate --fault-injection # prove the refusal path (drills only)
scripts/release.sh rollback # health-gated return to the previous release
scripts/release.sh status # release, tag, active pointer, recent log
scripts/test-release.sh # release selftests
```
- `<dataRoot>/state/active.json` — the activation pointer (atomic tmp+rename replace)
- `<dataRoot>/state/activation-log.jsonl` — append-only history: package / activate / refused / rollback
A failed health check never activates; the previously active release remains deployed. Updating the software therefore cannot corrupt the running installation: package beside, gate, then flip. Verified by the update/refusal/rollback drills in BUILD-LOG Phase 7.
## Runtime adapters (M4)
The harness boundary is formalized: everything upstream (config, contracts, missions, tasks, run records) is harness-agnostic; everything inside an adapter belongs to one runtime.
```text
adapters/<name>/adapter.sh env in: MOSAIC_SYSTEM_PROMPT_FILE, MOSAIC_REQUEST,
MOSAIC_PROVIDER, MOSAIC_MODEL
stdout: response only; stderr: diagnostics
```
- Selection: `execution.adapter` in config.json (optional; `pi` default; allowlist `pi`, `mock`)
- `pi` — pinned Pi CLI, noninteractive print mode, ambient discovery off
- `mock` — deterministic test adapter; never for real verification
- Mission directives have a sanctioned injection point: when a task references a mission, the task runner mounts the run snapshot and the generated prompt gains a `MISSION (runtime)` section (objective + directives) after the four immutable contracts
- Adding a harness (Claude, Codex, OpenCode) later means adding one directory — no orchestrator changes
See `adapters/README.md` for the full contract.
## Workspaces, capabilities, sessions (M5/M6)
Optional task fields extend what an agent can do — all defaulting to the previous behavior:
```json
{
"workspace": "demo", // ":run" ephemeral, or persistent dataRoot/workspaces/<name>
"capabilities": { "tools": ["bash", "read"] }, // pi tool allowlist; absent = no tools
"session": "demo" // persistent session at dataRoot/sessions/<name>
}
```
- The adapter runs inside the workspace; files it writes are host-visible (`dataRoot/workspaces/<name>`).
- Sessions persist via pi's documented `--session-dir`; a follow-up run in the same session resumes the conversation (`-c`) and can recall prior context. Distinct names never share state. Ephemeral (`--no-session`) remains the default when no session is declared.
- Selection authority: config for adapter/provider/model; the task file for workspace/capabilities/session.
Inspect anything:
```bash
node scripts/mosaic-task.mjs list # runs with task/workspace/session columns
node scripts/mosaic-task.mjs show <runId> # full record + snapshots + artifacts
```
Demo fixtures: `tasks/workspace-demo.json`, `tasks/session-demo-1.json` + `tasks/session-demo-2.json`.
See `docs/plans/2026-09-02_atomic-mosaic-foundation.md` for the full plan.
Inside the container:
@@ -95,7 +156,8 @@ Inside the container:
## How it works
1. `scripts/build.sh` builds `mosaic-poc-agent:0.84.4` with Docker Compose.
1. `scripts/build.sh` builds the release image (`mosaic-poc-agent:<pi>-r<release>`,
tag derived from `RELEASE` + the pinned Pi version) with Docker Compose.
2. On each run, `/opt/mosaic/src/load-contracts.sh` reads the four contract files
in fixed order (CONSTITUTION, STANDARDS, SOUL, USER), joins them with clear
separators, and writes `/var/lib/mosaic/system-prompt.md`.
@@ -116,8 +178,10 @@ scripts/build.sh # build the image
scripts/hello.sh # one-shot request; prints the model response
scripts/verify.sh # full gated test; exit 0 only on exact MOSAIC_HELLO_OK
scripts/run-task.sh # run a mission/task file (see Missions & tasks)
scripts/release.sh # package / activate / rollback / status (see Release model)
scripts/test-config.sh # fast config-layer selftests (no Docker)
scripts/test-task.sh # mission/task selftests (schema + live runs)
scripts/test-task.sh # mission/task selftests (schema + adapter seam + live runs)
scripts/test-release.sh # release selftests
scripts/reset.sh # delete the configured data root (safety-checked)
```
+1
View File
@@ -0,0 +1 @@
0.0.8
+54
View File
@@ -0,0 +1,54 @@
# Mosaic runtime adapters
An adapter is the entire harness-specific surface of the system. Everything
upstream of an adapter — configuration, contracts, missions, tasks, run
records — is harness-agnostic; everything inside an adapter may assume one
specific agent runtime.
## Contract
An adapter lives at:
```text
/opt/mosaic/adapters/<name>/adapter.sh
```
and must be executable. The dispatcher (`/opt/mosaic/src/run-agent.sh`)
selects it via `MOSAIC_ADAPTER` (default: `pi`) and execs it after the
system prompt has been generated.
**Inputs (environment):**
| Variable | Meaning |
|---|---|
| `MOSAIC_SYSTEM_PROMPT_FILE` | Absolute path to the generated system prompt (contracts + optional mission section). Read it; do not modify it. |
| `MOSAIC_REQUEST` | The exact user request text (may contain newlines). |
| `MOSAIC_PROVIDER` | Configured provider name. |
| `MOSAIC_MODEL` | Configured model id. |
Optional, adapter-specific (documented per adapter):
| Variable | Meaning |
|---|---|
| `MOSAIC_MOCK_RESPONSE` | mock only: the verbatim response to emit |
**Outputs:**
- `stdout`: the model response text — the only channel the orchestrator captures
- `stderr`: diagnostics (never credentials)
- exit `0`: success; nonzero: failure
## Rules
1. Adapters print ONLY the response on stdout. Status lines go to stderr.
2. Adapters never read configuration files; the resolved settings arrive via environment.
3. Adapters never write outside `/var/lib/mosaic`.
4. Adding an adapter requires: a new directory, the contract implementation, and
adding the name to the allowlist in `scripts/mosaic-config.mjs`.
## Included adapters
- `pi` — the pinned `@earendil-works/pi-coding-agent` CLI in noninteractive
print mode (`-p`), ambient discovery disabled, stdin detached.
- `mock` — deterministic echo of `MOSAIC_MOCK_RESPONSE`. Test-only: never use
it where a real model response is required.
+18
View File
@@ -0,0 +1,18 @@
#!/bin/sh
# Mock adapter: deterministic response for seam tests. NEVER use where a
# real model response is required.
#
# Contract: see /opt/mosaic/adapters/README.md.
set -eu
[ -n "${MOSAIC_SYSTEM_PROMPT_FILE:-}" ] || { echo "mock adapter: MOSAIC_SYSTEM_PROMPT_FILE is required" >&2; exit 2; }
if [ "${MOSAIC_INTERACTIVE:-}" != "1" ]; then
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "mock adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
fi
[ -r "$MOSAIC_SYSTEM_PROMPT_FILE" ] || { echo "mock adapter: system prompt not readable: $MOSAIC_SYSTEM_PROMPT_FILE" >&2; exit 2; }
echo "mock adapter: responding verbatim from MOSAIC_MOCK_RESPONSE" >&2
# Deterministic plumbing evidence: which MOSAIC_* variables did the
# orchestrator actually deliver? (Auth secrets are not MOSAIC_-prefixed.)
(env | grep '^MOSAIC_' | sort) >&2 2>/dev/null || true
printf '%s\n' "${MOSAIC_MOCK_RESPONSE:-}"
+83
View File
@@ -0,0 +1,83 @@
#!/bin/sh
# Pi adapter: implements the Mosaic adapter contract for the pinned
# @earendil-works/pi-coding-agent CLI.
#
# Contract: see /opt/mosaic/adapters/README.md.
# Headless (default): stdout = response only; stderr = diagnostics; exit 0.
# Interactive (MOSAIC_INTERACTIVE=1): full pi TUI on the attached terminal.
set -eu
[ -n "${MOSAIC_SYSTEM_PROMPT_FILE:-}" ] || { echo "pi adapter: MOSAIC_SYSTEM_PROMPT_FILE is required" >&2; exit 2; }
[ -r "$MOSAIC_SYSTEM_PROMPT_FILE" ] || { echo "pi adapter: system prompt not readable: $MOSAIC_SYSTEM_PROMPT_FILE" >&2; exit 2; }
# MOSAIC_AGENT_NAME is optional in headless mode (identity section is then
# omitted); interactive launches always set it via scripts/agent.sh.
: "${PI_PROVIDER:?pi adapter: PI_PROVIDER is required}"
: "${PI_MODEL:?pi adapter: PI_MODEL is required}"
INTERACTIVE="${MOSAIC_INTERACTIVE:-}"
if [ "$INTERACTIVE" != "1" ]; then
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "pi adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
fi
# Workspace (M5): run inside the provided workspace when present.
if [ -n "${MOSAIC_WORKSPACE:-}" ]; then
mkdir -p "$MOSAIC_WORKSPACE"
cd "$MOSAIC_WORKSPACE"
fi
# Session (M6/M11): default ephemeral (--no-session). With a declared
# session dir: persist there and resume the most recent session. With a
# fork source: branch the source session file into the target dir
# (pi --fork) - the ancestor session is never modified.
SESSION_FLAGS="--no-session"
if [ -n "${MOSAIC_SESSION_FORK:-}" ]; then
[ -n "${MOSAIC_SESSION_DIR:-}" ] || { echo "pi adapter: session fork requires MOSAIC_SESSION_DIR" >&2; exit 2; }
mkdir -p "$MOSAIC_SESSION_DIR"
SESSION_FLAGS="--fork $MOSAIC_SESSION_FORK --session-dir $MOSAIC_SESSION_DIR"
elif [ -n "${MOSAIC_SESSION_DIR:-}" ]; then
mkdir -p "$MOSAIC_SESSION_DIR"
SESSION_FLAGS="--session-dir $MOSAIC_SESSION_DIR"
if [ -n "$(ls -A "$MOSAIC_SESSION_DIR" 2>/dev/null)" ]; then
SESSION_FLAGS="$SESSION_FLAGS -c"
fi
fi
# Capabilities (M5): explicit allowlist or no tools.
TOOLS_FLAG="--no-tools"
[ -n "${MOSAIC_TOOLS:-}" ] && TOOLS_FLAG="--tools $MOSAIC_TOOLS"
# Mode (M13): interactive TUI or one-shot print.
PRINT_MODE="-p"
REQUEST_ARG=""
if [ "$INTERACTIVE" = "1" ]; then
PRINT_MODE=""
else
REQUEST_ARG="$MOSAIC_REQUEST"
fi
# All flags documented in the pi package README (CLI Reference):
# -p/--print one-shot mode: print the response and exit (omitted in
# interactive TUI mode)
# --system-prompt replace the default prompt with the generated one
# --no-* no ambient context/skills/extensions/templates/themes
# SESSION_FLAGS ephemeral | persistent | forked (per env)
# TOOLS_FLAG per capabilities
# --offline no startup network operations (update checks/telemetry)
PROMPT_CONTENT="$(cat "$MOSAIC_SYSTEM_PROMPT_FILE")"
set -- \
--offline \
--no-extensions \
--no-skills \
--no-prompt-templates \
--no-themes \
--no-context-files \
$TOOLS_FLAG \
$SESSION_FLAGS \
--provider "$PI_PROVIDER" \
--model "$PI_MODEL" \
--system-prompt "$PROMPT_CONTENT"
# One-shot mode appends -p and the request (both safely quoted);
# interactive mode appends nothing - clean TUI.
[ "$INTERACTIVE" = "1" ] || set -- "$@" -p "$MOSAIC_REQUEST"
exec pi "$@"
+20 -4
View File
@@ -3,13 +3,29 @@ services:
build:
context: .
dockerfile: Containerfile
image: mosaic-poc-agent:0.84.4
image: ${MOSAIC_IMAGE_TAG:?MOSAIC_IMAGE_TAG must be set by scripts/load_release (run via scripts/*.sh)}
user: "1000:1000"
environment:
# Resolved from config.json by scripts/common.sh (load_config).
# Required: compose fails fast when the launcher did not supply them.
PI_PROVIDER: ${MOSAIC_PROVIDER:?MOSAIC_PROVIDER must be set by scripts/load_config (run via scripts/*.sh)}
PI_MODEL: ${MOSAIC_MODEL:?MOSAIC_MODEL must be set by scripts/load_config (run via scripts/*.sh)}
# Adapter selection (resolved from config execution.adapter; default pi)
MOSAIC_ADAPTER: ${MOSAIC_ADAPTER:-pi}
# Mission directives injection point (set by the task runner when the
# task references a mission; container path of the run snapshot)
MOSAIC_MISSION_FILE: ${MOSAIC_MISSION_FILE:-}
# Workspace + capabilities (set by the task runner; M5)
MOSAIC_WORKSPACE: ${MOSAIC_WORKSPACE:-}
MOSAIC_TOOLS: ${MOSAIC_TOOLS:-}
# Persistent named session dir + optional fork source (M6/M11)
MOSAIC_SESSION_DIR: ${MOSAIC_SESSION_DIR:-}
MOSAIC_SESSION_FORK: ${MOSAIC_SESSION_FORK:-}
# Interactive TUI mode + agent identity (M13, set by scripts/agent.sh)
MOSAIC_INTERACTIVE: ${MOSAIC_INTERACTIVE:-}
MOSAIC_AGENT_NAME: ${MOSAIC_AGENT_NAME:-}
# mock adapter only: verbatim response for deterministic seam tests
MOSAIC_MOCK_RESPONSE: ${MOSAIC_MOCK_RESPONSE:-}
# Documented container auth alternative: provider API key via
# runtime environment variable. Empty by default; when empty Pi
# falls back to the read-only mounted auth.json credential file.
@@ -21,6 +37,6 @@ services:
# Runtime credential only: pi auth file mounted READ-ONLY.
# Never copied into the image.
- ${PI_AUTH_FILE:-/home/jwoltje/.pi/agent/auth.json}:/home/node/.pi/agent/auth.json:ro
# One-shot: the exact startup verification request. It deliberately
# does NOT contain the expected marker MOSAIC_HELLO_OK.
command: ["Return your startup marker and nothing else."]
# Headless runs: the request is passed as command args by the launchers
# (run-task.sh) or defaults inside run-agent.sh (hello/verify). Never a
# fixed command here - interactive runs (scripts/agent.sh) need no args.
+9
View File
@@ -0,0 +1,9 @@
# Session registry — append-only
Every agent session (assistant, worker-cycle conductor, or owner-directed
automation) that works in this repository registers one line here. Entries
are never rewritten or removed; corrections are new entries.
| Date (UTC) | Actor | Scope | Outcome / artifacts |
|---|---|---|---|
| 2026-09-03 | assistant (conductor + worker) | POC through M12: containerized pi proof, config layer, missions/tasks, release model, adapter seam, workspaces/capabilities, named sessions, retention, session forking, conductor auto-apply, roles/ convention | 13 tags; suites 24/58/14 + 17 conductor + verify green; releases 0.0.10.0.7; issues #1#34 closed |
+85
View File
@@ -0,0 +1,85 @@
# TOOLS.md — command and tool reference
On-demand reference for agent sessions (conductors, bootstrapping agents,
reviewers). `AGENTS.md` routes here; this file carries the depth: usage,
inputs/outputs, exit codes, and safety notes for every entry point.
Reading guide: all entry points are `scripts/*.sh` (bash) or invoked via
`node scripts/mosaic-task.mjs` (node). Every script fails closed — missing
or invalid configuration/policy refuses the operation with a nonzero exit
and changes nothing.
## Lifecycle
| Command | Purpose | Notes |
|---|---|---|
| `scripts/bootstrap.sh` | Create `~/.config/mosaic-dev/config.json` if absent | Idempotent; existing config validated, never rewritten |
| `scripts/build.sh` | Build the release image | Tag derived from `RELEASE` + pinned pi version |
| `scripts/hello.sh` | One-shot startup request | Prints model response on stdout |
| `scripts/verify.sh` | Full gated test | Exit 0 only on exact `MOSAIC_HELLO_OK`; `EXPECTED_MARKER` overrides for negative drills |
## Tasks (missions, runs, evidence)
| Command | Purpose | Notes |
|---|---|---|
| `scripts/run-task.sh run <task.json>` | Execute a task | Immutable run record under `<dataRoot>/runs/` |
| `scripts/run-task.sh validate <task.json>` | Strict validation | Writes nothing |
| `node scripts/mosaic-task.mjs show <runId>` | Inspect a run | Full record + snapshots + artifacts |
| `node scripts/mosaic-task.mjs list` | List runs | task/workspace/session columns |
| `node scripts/mosaic-task.mjs retry <runId>` | Re-execute a run's snapshot | New run dir; `retriedFrom` lineage recorded |
| `node scripts/mosaic-task.mjs prune [--keep=N] [--yes]` | Retention | Dry-run default; receipt in `runs/.pruned.log` |
Task fields: `prompt` (required), `mission` (path), `expectExact`,
`timeoutSeconds` (5600), `workspace` (`:run` or named), `capabilities.tools`
(allowlist: read write edit bash grep find ls), `session`,
`sessionForkFrom` (requires `session`). Mission fields: `objective`,
`directives[]`, optional governing `capabilities.tools`. Policy: a task may
narrow a mission's tools, never widen; empty intersection = tool-free run.
## Agent (interactive TUI)
```bash
scripts/agent.sh <name> [--mission <file>] [--workspace <ws>] [--session <s>] [--tools <list>]
```
Launches an interactive pi TUI inside the container with the four immutable
contracts + optional mission + agent identity as its system prompt,
persistent named session, optional workspace. Exit with `/quit`.
## Release
| Command | Purpose | Notes |
|---|---|---|
| `scripts/release.sh package` | Build + tag the release image | Tag: `mosaic-poc-agent:<pi>-r<release>` |
| `scripts/release.sh activate` | Health gate → atomic pointer swap | `--fault-injection` proves the refusal path |
| `scripts/release.sh rollback` | Health-gated return to previous | Refuses if image missing |
| `scripts/release.sh status` | Release, tag, active pointer, log | Safe on empty state |
## Conductor (worker patches)
```bash
scripts/conductor-apply.sh <runId> [--dry-run]
```
Auto-applies a worker's patch under `roles/conductor-policy.json`:
succeeded run → clean target tree → path allowlist → syntax gates →
apply → policy suites → attribution commit. Any failure reverts.
Push is never automatic.
## Maintenance
| Command | Purpose | Notes |
|---|---|---|
| `scripts/reset.sh` | Delete the data root | Triple-safety-checked (path, symlink, ownership marker) |
| `scripts/test-config.sh` | Config selftests (no Docker) | 24 cases |
| `scripts/test-task.sh` | Task selftests + live cases | 58 cases |
| `scripts/test-release.sh` | Release selftests | 14 cases |
| `scripts/test-conductor.sh` | Auto-apply selftests (sandboxed) | 17 cases |
| `scripts/gitea-api.sh <METHOD> <path> [body]` | Gitea API helper | Token never on argv/stdout |
## Exit-code convention
`0` success · `1` operation failed · `2` invalid data/configuration ·
`3` configuration missing for a read operation · `4` usage/file/environment
problem. Scripts print diagnostics on stderr; model responses (and only
model responses) on stdout.
+81
View File
@@ -0,0 +1,81 @@
# Autonomous Work Run — 2026-09-03
**Status:** COMPLETED (single-session batch; see Results at bottom)
**Constraint:** The assistant cannot run unattended. This was one long interactive session, not 12 wall-clock hours. Everything below was completed, committed, and pushed during that session.
## Objective
Advance the Mosaic Stack rebuild several verified layers in one batch, focused on Pi, ending in a state the owner can test and review alone: green suites, activated release, recorded drills, and this document as the single entry point.
## Scope decided for this run
| Milestone | Theme | Status |
|---|---|---|
| M5 | Task workspaces + capability envelope (tools allowlist) | DONE |
| M6 | Named sessions — persistence and resume (L1) | DONE |
| M7 | Operator ergonomics: run inspection commands | DONE |
| — | Releases 0.0.5+0.0.6 packaged; 0.0.6 health-gated activated | DONE |
Explicitly deferred (do not mistake for forgotten):
- Claude/Codex/OpenCode adapters (owner: focus on Pi for now)
- Network policy engine (container boundary is the current control)
- Fine-grained read restrictions (excluded by the original brief)
- Config/state migrations (no schema breaks so far; keep it that way)
## Design decisions taken during this run
1. **Workspace** (`task.workspace`, optional):
- absent → tool-free text-only run (previous behavior, unchanged)
- `":run"` → ephemeral per-run workspace at `<dataRoot>/runs/<runId>/workspace`
- named (validated id) → persistent shared workspace at `<dataRoot>/workspaces/<name>`
- Container path passed via `MOSAIC_WORKSPACE` env; adapter cds into it. No new mounts (dataRoot is already mounted).
2. **Capabilities** (`task.capabilities.tools`, optional): allowlist from pi's documented tool set (`read write edit bash grep find ls`). Absent → `--no-tools` (previous behavior). Passed via `MOSAIC_TOOLS` env; pi adapter maps to `--tools`.
3. **Adapter diagnostics for deterministic testing**: the mock adapter writes all received `MOSAIC_*` variables (never secrets — auth is not MOSAIC_-prefixed) to stderr, which lands in the run record. This lets selftests assert orchestrator→adapter plumbing without parsing model output.
4. **Sessions** (`task.session`, optional named): persisted under `<dataRoot>/sessions/<name>/` via pi's documented `--session-dir`; resume semantics: continue most recent session in that directory when one exists (`-c`).
5. **Selection authority unchanged**: config file for adapter/provider/model; task file for workspace/capabilities/session; env vars are internal plumbing only.
6. **configVersion stays 1**; all new task fields are optional. Old tasks/configs remain valid.
## Test plan (what "done" means per milestone)
- M5: mock-adapter cases asserting workspace path and tools arrive via run-record stderr; live pi case writing/reading a file in a persistent workspace; validation negatives (bad tool name, bad workspace name)
- M6: session directory deterministically populated after first run; second run resumes (continuation asserted by session dir state and, in live E2E, by model recall); sandbox isolation between two named sessions
- M7: `show <runId>` prints a complete run record; `list` gains workspace/session columns
- Final: full sweep (config/task/release), verify, package + activate 0.0.6, config checksum unchanged
## Review checklist for the owner
1. `cat docs/plans/2026-09-03_autonomous-run.md` (this file)
2. `scripts/release.sh status` → 0.0.6 active
3. `scripts/test-config.sh && scripts/test-task.sh && scripts/test-release.sh && scripts/verify.sh`
4. Try a workspace task:
```bash
scripts/run-task.sh run tasks/workspace-demo.json
ls ~/.mosaic-dev/workspaces/demo/
```
5. Try the session demo:
```bash
scripts/run-task.sh run tasks/session-demo-1.json # teaches a word
scripts/run-task.sh run tasks/session-demo-2.json # recalls it
```
6. Inspect any run: `node scripts/mosaic-task.mjs show <runId>`
7. Gitea: milestones M5/M6/M7 closed; issues referenced by merge commits
## Results
- M5 merged on `main` (merge commit `ddb1554`), tagged `workspace-capabilities-v1`
- M6 merged on `main` (merge commit `4e2a413`), tagged `sessions-v1`
- M7 merged on `main`, tagged `operator-ergonomics-v1`
- Release 0.0.6 packaged, health-gated activated, full sweep green
- Suites at end of run: config 24/24, task 32/32, release 14/14, verify PASS
- Build log: Phases 9 (M5), 10 (M6), 11 (M7) appended with corrections
- Corrections encountered: dropped constant from a failed atomic edit batch (SUPPORTED_TOOLS); dash `export` output format vs `env`; three selftest authoring defects; showRun id-regex case sensitivity + missing-run crash. All fixed and covered by tests.
- Commits pushed incrementally; nothing left uncommitted
- Live proof: workspace file host-visible; session teach/recall ('mosaico') verified
## Next steps after this run (not started)
1. Owner review + hands-on testing of workspaces, capabilities, sessions
2. Decision: capability defaults per mission (mission-level policy) — natural M8
3. Second real adapter remains available whenever wanted
4. Consider run-record pruning/retention policy once run volume grows
5. Consider a `mosaic-task.mjs retry <runId>` convenience for failed runs
+58
View File
@@ -0,0 +1,58 @@
# Conductor protocol — poor-man orchestration loop
How the stack orchestrates headless pi workers to do work on itself.
## Role contracts
Role authority is declared in role contracts, one file per role, under
`roles/` (e.g. `roles/conductor-policy.json`). The repository root holds
only first-class, bootstrap-required configuration; role contracts are
tracked, versioned files whose changes arrive as reviewed commits.
## Roles
| Role | Runs where | Powers | Never has |
|---|---|---|---|
| **Conductor** | host (assistant or owner) | git (clone/commit/push), task dispatch, review, verification suites, Gitea | nothing new |
| **Worker** | container (headless pi via `scripts/run-task.sh`) | read/write/edit/bash inside its workspace; persistent session on request | git credentials, docker socket, host filesystem |
## The loop
1. **Decompose**: conductor turns a goal into worker tasks small enough to
specify completely in one prompt (file paths, acceptance criteria, style
constraints, verification the worker can run itself, e.g. `node --check`).
2. **Mirror**: conductor maintains the repo clone at
`<dataRoot>/workspaces/stack-repo` (host-side git; workers see it read-write
through their workspace mount).
3. **Dispatch**: `scripts/run-task.sh run <worker-task.json>` — worker edits the
clone. Session name `worker-<n>` keeps continuity across refinement rounds.
4. **Extract**: `git -C <workspace> diff > patch` — the worker's entire output
is a reviewable diff. Run record (result.json, stderr.txt) is the receipt.
5. **Review**: conductor reads the diff line by line. Bad output → refine the
prompt, re-dispatch (same session: "your patch had these problems…").
6. **Integrate**: conductor applies the patch to the real repo, runs the full
suites, commits and pushes. Suites failing → revert apply, back to step 5.
7. **Record**: update CURRENT.md, BUILD-LOG, close the Gitea issue.
## Guardrails
- Workers never receive credentials; they never run git; they never leave the
workspace (container is the boundary; tools allowlist is the gate).
- Every worker diff is reviewed by the conductor before integration. No
auto-apply. (Auto-apply would be a capability-policy decision for later.)
- Verification is mechanical: suites + `node --check` / `bash -n` gates.
- Recursive decomposition = "fail → smaller task", never "hope."
## Worker task template
```json
{
"taskVersion": 1,
"id": "t-worker-<name>",
"prompt": "<full spec: goal, files, constraints, acceptance, self-checks>",
"workspace": "stack-repo",
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
"session": "worker-1",
"timeoutSeconds": 600
}
```
+61
View File
@@ -0,0 +1,61 @@
# CURRENT — single source of "what happens next"
This file always names exactly one next action. Any "continue" / "next" /
"proceed" message means: execute the action below, fully (implement → test →
verify against its acceptance criteria → commit → push → close the issue →
update this file to the next action). No ambiguity, no re-planning.
## Next action
Owner review of M12 (conductor auto-apply policy) — then name the next target.
## Queue (ordered, not started)
1. Second real adapter (parked — owner focused on Pi)
2. Session forking from a common ancestor — SHIPPED in M11; exercise via session demos
3. UX/DX backlog #31 (grep highlight vs status colors — likely closed by owner's unalias)
4. Push policy decision: auto-apply commits locally; push remains explicit (documented in CONDUCTOR.md)
## Rules
- One action in flight. Update this file at the END of every action.
- Blocked? Move the item to "Blocked" below with the reason and stop.
- Completed actions move to the log at the bottom (date + issue + result).
## Blocked
(none)
## Completed log
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
- 2026-09-03 — M12 conductor auto-apply policy (#34) — committed on main, 17/24/58/14 + verify green; push stays explicit
- 2026-09-03 — repository convention: role contracts move to roles/ (root = bootstrap-only, per owner direction)
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
- 2026-09-03 — M12 conductor auto-apply policy (#34) — committed on main, 17/24/58/14 + verify green; push stays explicit
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
- 2026-09-03 — M5 workspaces + capability envelope (#20, #21) — merged, tests green
- 2026-09-03 — M6 named sessions + resume (#22, #23) — merged, teach/recall verified
- 2026-09-03 — M7 run inspection + release 0.0.6 (#24) — merged, activated
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
- 2026-09-03 — retry lineage + relative mission resolution (#28) — merged, 36/24/14 suites + verify green
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection, 41/36/14 + verify green
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
- 2026-09-03 — retry lineage + relative mission resolution (#28) — merged, 36/24/14 suites + verify green
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
+1 -1
View File
@@ -1,7 +1,7 @@
{
"missionVersion": 1,
"id": "m-hello",
"objective": "Prove the startup marker path of the mosaic-poc-agent.",
"objective": "Verify the startup marker path.",
"directives": [
"Startup verification requests are answered with the marker only.",
"No explanation, no formatting."
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "mosaic-stack-dev-test",
"version": "0.1.0",
"version": "0.0.3",
"private": true,
"description": "Minimal Mosaic Stack container proof of concept: one Pi agent, four local contract files, one real model request returning MOSAIC_HELLO_OK.",
"license": "UNLICENSED",
+19
View File
@@ -0,0 +1,19 @@
{
"policyVersion": 1,
"autoApply": {
"enabled": true,
"allowedPaths": [
"scripts/**",
"docs/**",
"tasks/**",
"missions/**",
"adapters/**",
"README.md"
],
"suites": [
"test-config",
"test-task",
"test-release"
]
}
}
+67
View File
@@ -0,0 +1,67 @@
#!/usr/bin/env bash
# Launch an interactive (TUI) Mosaic agent in its container.
#
# Usage:
# scripts/agent.sh <name> [--mission <file>] [--workspace <ws>]
# [--session <name>] [--tools <comma,list>]
#
# The agent receives the four immutable contracts (constitution, standards,
# SOUL, USER) plus its own identity and optional mission directives as its
# system prompt, a persistent named session, and - if declared - a
# workspace and tool capabilities. The TUI opens clean; you drive.
#
# This is the Mosaic alternative to launching vanilla pi: same engine,
# governed context.
set -euo pipefail
cd "$(dirname "$0")/.."
# shellcheck source=common.sh
source scripts/common.sh
NAME=""
MISSION=""
WORKSPACE=""
SESSION=""
TOOLS=""
while [ $# -gt 0 ]; do
case "$1" in
--mission) MISSION="${2:?}"; shift 2 ;;
--workspace) WORKSPACE="${2:?}"; shift 2 ;;
--session) SESSION="${2:?}"; shift 2 ;;
--tools) TOOLS="${2:?}"; shift 2 ;;
--help|-h) sed -n '2,12p' "$0"; exit 0 ;;
*) NAME="$1"; shift ;;
esac
done
[ -n "$NAME" ] || { echo "agent: usage: scripts/agent.sh <name> [--mission f] [--workspace ws] [--session s] [--tools list]" >&2; exit 4; }
case "$NAME" in *[!A-Za-z0-9._-]*|'') echo "agent: invalid agent name" >&2; exit 4;; esac
load_config
load_release
bootstrap_runtime_dir
SESSION="${SESSION:-agent-$NAME}"
mkdir -p "$MOSAIC_DEV_DIR/sessions/$SESSION"
export MOSAIC_SESSION_DIR="/var/lib/mosaic/sessions/$SESSION"
export MOSAIC_AGENT_NAME="$NAME"
export MOSAIC_INTERACTIVE=1
export MOSAIC_TOOLS="${TOOLS:+$TOOLS}"
if [ -n "$MISSION" ]; then
[ -r "$MISSION" ] || { echo "agent: mission file not readable: $MISSION" >&2; exit 4; }
mkdir -p "$MOSAIC_DEV_DIR/agent-missions"
cp "$MISSION" "$MOSAIC_DEV_DIR/agent-missions/$NAME.json"
export MOSAIC_MISSION_FILE="/var/lib/mosaic/agent-missions/$NAME.json"
fi
if [ -n "$WORKSPACE" ]; then
case "$WORKSPACE" in *[!A-Za-z0-9._-]*|'') echo "agent: invalid workspace name" >&2; exit 4;; esac
mkdir -p "$MOSAIC_DEV_DIR/workspaces/$WORKSPACE"
export MOSAIC_WORKSPACE="/var/lib/mosaic/workspaces/$WORKSPACE"
fi
echo "agent: launching TUI agent '$NAME' (session: $SESSION, adapter: $MOSAIC_ADAPTER, model: $MOSAIC_MODEL)"
echo "agent: contracts + $([ -n "$MISSION" ] && echo 'mission' || echo 'no mission') loaded; exit the TUI with /quit"
# No -T: the TTY is the point. Ctrl+C twice or /quit exits.
exec docker compose run --rm mosaic-agent
+1
View File
@@ -6,6 +6,7 @@ cd "$(dirname "$0")/.."
source scripts/common.sh
load_config
load_release
bootstrap_runtime_dir
+16 -1
View File
@@ -15,10 +15,25 @@ load_config() {
exit 1
fi
eval "$config_env"
export MOSAIC_DATA_ROOT MOSAIC_PROVIDER MOSAIC_MODEL
export MOSAIC_DATA_ROOT MOSAIC_PROVIDER MOSAIC_MODEL MOSAIC_ADAPTER
MOSAIC_DEV_DIR="$MOSAIC_DATA_ROOT"
}
# Resolve the release identity: RELEASE is the single source of the
# release version (stays 0.0.X until declared stable); the image tag
# derives from it plus the pinned pi dependency version.
load_release() {
local release pi_version
release="$(tr -d '[:space:]' < RELEASE)"
if ! printf '%s' "$release" | grep -Eq '^[0-9]+\.[0-9]+\.[0-9]+$'; then
echo "common: RELEASE must be a semver-ish version, got: '$release'" >&2
exit 1
fi
pi_version="$(node -p "require('./package.json').dependencies['@earendil-works/pi-coding-agent']")"
export MOSAIC_RELEASE="$release"
export MOSAIC_IMAGE_TAG="mosaic-poc-agent:${pi_version}-r${release}"
}
# Ensure the configured runtime data directory exists and carries this
# project's ownership marker. The marker is what scripts/reset.sh requires
# before it will delete anything.
+139
View File
@@ -0,0 +1,139 @@
#!/usr/bin/env bash
# Conductor auto-apply: integrate a worker's patch under the declared policy.
#
# Usage: scripts/conductor-apply.sh <runId> [--dry-run]
#
# Policy (roles/conductor-policy.json in the target repo, strictly validated):
# autoApply.enabled master switch
# autoApply.allowedPaths glob allowlist ('dir/**' = everything under dir)
# autoApply.suites suite scripts that must pass AFTER applying
#
# Gate sequence: succeeded run record -> clean target tree -> diff extracted
# from the worker workspace -> allowlist -> syntax gates (node/bash/json) ->
# apply -> policy suites -> commit with attribution. ANY failure reverts the
# working tree and exits nonzero. Push is never automatic.
#
# Environment:
# MOSAIC_APPLY_TARGET repo root to apply into (default: this project)
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
TARGET_ROOT="${MOSAIC_APPLY_TARGET:-$(cd "$SCRIPT_DIR/.." && pwd)}"
RUN_ID="${1:?usage: conductor-apply.sh <runId> [--dry-run]}"
DRY_RUN="no"
[ "${2:-}" = "--dry-run" ] && DRY_RUN="yes"
cd "$TARGET_ROOT"
fail() { echo "conductor-apply: $*" >&2; exit "${2:-1}"; }
[ -d .git ] || fail "target is not a git repository: $TARGET_ROOT" 4
[ -f roles/conductor-policy.json ] || fail "no roles/conductor-policy.json in target" 2
# ---- policy (strict) ----
POLICY_JSON="$(node -e '
const fs = require("fs");
const p = JSON.parse(fs.readFileSync("roles/conductor-policy.json", "utf8"));
if (p.policyVersion !== 1) process.exit(3);
if (!p.autoApply || typeof p.autoApply.enabled !== "boolean" || !Array.isArray(p.autoApply.allowedPaths) || !Array.isArray(p.autoApply.suites)) process.exit(3);
for (const g of p.autoApply.allowedPaths) {
if (typeof g !== "string" || !/^[A-Za-z0-9_.*/-]+$/.test(g) || g.startsWith("/") || g.includes("..")) process.exit(3);
}
console.log(JSON.stringify(p.autoApply));
')" || fail "invalid roles/conductor-policy.json" 2
ENABLED="$(node -e 'console.log(JSON.parse(process.argv[1]).enabled)' "$POLICY_JSON")"
[ "$ENABLED" = "true" ] || fail "auto-apply is disabled by policy" 2
# ---- run record ----
DATA_ROOT="$(node scripts/mosaic-config.mjs validate | node -e 'let d="";process.stdin.on("data",c=>d+=c).on("end",()=>console.log(JSON.parse(d).dataRoot))')"
RESULT_FILE="$DATA_ROOT/runs/$RUN_ID/result.json"
[ -f "$RESULT_FILE" ] || fail "run not found: $RUN_ID" 4
node -e '
const r = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
process.exit(r.status === "succeeded" ? 0 : 1);
' "$RESULT_FILE" || fail "run $RUN_ID did not succeed; refusing to auto-apply"
WORKSPACE_NAME="$(node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.workspace||"")' "$RESULT_FILE")"
[ -n "$WORKSPACE_NAME" ] || fail "run has no workspace; nothing to integrate" 4
WORKSPACE="$DATA_ROOT/workspaces/$WORKSPACE_NAME"
[ -d "$WORKSPACE/.git" ] || fail "workspace is not a git clone: $WORKSPACE" 4
# ---- extract diff (tracked + intent-to-add) ----
git -C "$WORKSPACE" add -N . >/dev/null 2>&1 || true
DIFF_FILE="$(mktemp)"
trap 'rm -f "$DIFF_FILE"' EXIT
git -C "$WORKSPACE" diff > "$DIFF_FILE"
if [ ! -s "$DIFF_FILE" ]; then
fail "workspace has no changes to apply"
fi
# ---- allowlist ----
mapfile -t CHANGED < <(git -C "$WORKSPACE" diff --name-only)
GLOBS="$(node -e 'const a=JSON.parse(process.argv[1]).allowedPaths;console.log(a.join("\n"))' "$POLICY_JSON")"
REFUSED=""
for f in "${CHANGED[@]}"; do
ok="no"
while IFS= read -r g; do
[ -z "$g" ] && continue
case "$f" in
$g) ok="yes"; break ;;
esac
done <<< "$GLOBS"
[ "$ok" = "yes" ] || REFUSED="$REFUSED $f"
done
if [ -n "$REFUSED" ]; then
echo "conductor-apply: refusing - files outside policy allowlist:$REFUSED" >&2
echo "conductor-apply: patch preserved at $DIFF_FILE for manual review" >&2
exit 1
fi
# ---- syntax gates (on workspace files, pre-apply) ----
for f in "${CHANGED[@]}"; do
case "$f" in
*.mjs) node --check "$WORKSPACE/$f" || fail "syntax gate failed (node): $f" 1 ;;
*.sh) bash -n "$WORKSPACE/$f" || fail "syntax gate failed (bash): $f" 1 ;;
*.json) node -e 'JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"))' "$WORKSPACE/$f" || fail "syntax gate failed (json): $f" 1 ;;
esac
done
if [ "$DRY_RUN" = "yes" ]; then
echo "conductor-apply (dry-run): would apply $(echo "${#CHANGED[@]}") file(s) from $RUN_ID:"
printf ' %s\n' "${CHANGED[@]}"
echo "conductor-apply (dry-run): suites would run: $(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(", "))' "$POLICY_JSON")"
exit 0
fi
# ---- apply ----
[ -z "$(git status --porcelain)" ] || fail "target tree is not clean; refusing to mix states"
git apply "$DIFF_FILE" || fail "git apply failed"
# ---- policy suites ----
SUITES="$(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(" "))' "$POLICY_JSON")"
SUITES_OK="yes"
for s in $SUITES; do
case "$s" in
test-[a-z]*) : ;; # shape guard; existence checked next
*) echo "conductor-apply: refusing suspicious suite name: $s" >&2; SUITES_OK="no"; break ;;
esac
[ -x "scripts/$s.sh" ] || { echo "conductor-apply: suite script missing: scripts/$s.sh" >&2; SUITES_OK="no"; break; }
if ! bash "scripts/$s.sh" >/dev/null 2>&1; then
echo "conductor-apply: suite failed: $s" >&2
SUITES_OK="no"
break
fi
done
if [ "$SUITES_OK" != "yes" ]; then
git apply -R "$DIFF_FILE" && echo "conductor-apply: changes REVERTED (suites failed)" >&2
exit 1
fi
# ---- commit with attribution ----
git add -A
git commit -q -m "feat(worker): auto-applied patch from run $RUN_ID
Authored-by: pi worker (run $RUN_ID, workspace $WORKSPACE_NAME)
Applied-under: conductor-policy v1 (allowlist + syntax gates + suites)"
echo "conductor-apply: applied and committed run $RUN_ID ($(echo "${#CHANGED[@]}") file(s)); suites: $SUITES"
echo "conductor-apply: NOT pushed - push remains an explicit act."
+1
View File
@@ -11,6 +11,7 @@ cd "$(dirname "$0")/.."
source scripts/common.sh
load_config
load_release
bootstrap_runtime_dir
+11 -1
View File
@@ -30,6 +30,7 @@ import process from "node:process";
const SUPPORTED_CONFIG_VERSION = 1;
const SUPPORTED_ENVIRONMENTS = new Set(["development", "production"]);
const SUPPORTED_BACKENDS = new Set(["docker"]);
const SUPPORTED_ADAPTERS = new Set(["pi", "mock"]); // mock: test-only, see adapters/README.md
const NAME_PATTERN = /^[A-Za-z0-9][A-Za-z0-9._:/-]{0,199}$/;
function fail(exitCode, message) {
@@ -127,7 +128,7 @@ function validate(document, file) {
if (!isPlainObject(document.execution)) {
fail(2, '"execution" must be a JSON object');
}
rejectUnknownKeys(document.execution, ["backend", "provider", "model"], '"execution"');
rejectUnknownKeys(document.execution, ["backend", "provider", "model", "adapter"], '"execution"');
if (!SUPPORTED_BACKENDS.has(document.execution.backend)) {
fail(2, `unsupported execution.backend: ${JSON.stringify(document.execution.backend)} (supported: ${[...SUPPORTED_BACKENDS].join(", ")})`);
}
@@ -138,6 +139,13 @@ function validate(document, file) {
}
}
const adapter = document.execution.adapter === undefined || document.execution.adapter === null
? "pi"
: document.execution.adapter;
if (typeof adapter !== "string" || !SUPPORTED_ADAPTERS.has(adapter)) {
fail(2, `unsupported execution.adapter: ${JSON.stringify(adapter)} (supported: ${[...SUPPORTED_ADAPTERS].join(", ")})`);
}
return {
configVersion: document.configVersion,
environment: document.environment,
@@ -146,6 +154,7 @@ function validate(document, file) {
backend: document.execution.backend,
provider: document.execution.provider,
model: document.execution.model,
adapter,
},
};
}
@@ -222,6 +231,7 @@ switch (operation) {
`MOSAIC_DATA_ROOT=${shellQuote(resolved.dataRoot)}`,
`MOSAIC_PROVIDER=${shellQuote(resolved.execution.provider)}`,
`MOSAIC_MODEL=${shellQuote(resolved.execution.model)}`,
`MOSAIC_ADAPTER=${shellQuote(resolved.execution.adapter)}`,
"",
].join("\n"),
);
+338 -7
View File
@@ -28,6 +28,7 @@
*/
import fs from "node:fs";
import os from "node:os";
import path from "node:path";
import process from "node:process";
import { randomBytes } from "node:crypto";
@@ -38,6 +39,7 @@ const PROJECT_ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)),
const RUNS_DIRNAME = "runs";
const ID_PATTERN = /^[a-z0-9][a-z0-9._-]{0,63}$/;
const DEFAULT_TIMEOUT_SECONDS = 120;
const SUPPORTED_TOOLS = ["read", "write", "edit", "bash", "grep", "find", "ls"]; // pi documented built-ins
function fail(exitCode, message) {
process.stderr.write(`mosaic-task: ${message}\n`);
@@ -82,7 +84,7 @@ function validateId(value, what) {
function validateMission(document, file) {
if (!isPlainObject(document)) fail(2, "mission must be a JSON object");
rejectUnknownKeys(document, ["missionVersion", "id", "objective", "directives"], "mission");
rejectUnknownKeys(document, ["missionVersion", "id", "objective", "directives", "capabilities"], "mission");
if (document.missionVersion !== 1) {
fail(2, `unsupported missionVersion: ${JSON.stringify(document.missionVersion)} (supported: 1)`);
}
@@ -100,12 +102,34 @@ function validateMission(document, file) {
return d;
});
}
return { missionVersion: document.missionVersion, id: document.id, objective: document.objective, directives };
// Governing capability constraints (M9): same validation as task
// capabilities; semantically these BOUND tasks (least-privilege
// intersection at run time), never grant beyond them.
let capabilities = null;
if (document.capabilities !== undefined && document.capabilities !== null) {
if (!isPlainObject(document.capabilities)) fail(2, 'mission "capabilities" must be a JSON object');
rejectUnknownKeys(document.capabilities, ["tools"], 'mission "capabilities"');
if (!Array.isArray(document.capabilities.tools) || document.capabilities.tools.length === 0) {
fail(2, 'mission "capabilities.tools" must be a non-empty array of tool names');
}
const seen = new Set();
for (const tool of document.capabilities.tools) {
if (!SUPPORTED_TOOLS.includes(tool)) {
fail(2, `unsupported mission tool: ${JSON.stringify(tool)} (supported: ${SUPPORTED_TOOLS.join(", ")})`);
}
if (seen.has(tool)) fail(2, `duplicate tool in mission capabilities.tools: ${tool}`);
seen.add(tool);
}
capabilities = { tools: [...seen] };
}
return { missionVersion: document.missionVersion, id: document.id, objective: document.objective, directives, capabilities };
}
function validateTask(document, file) {
if (!isPlainObject(document)) fail(2, "task must be a JSON object");
rejectUnknownKeys(document, ["taskVersion", "id", "prompt", "mission", "expectExact", "timeoutSeconds"], "task");
rejectUnknownKeys(document, ["taskVersion", "id", "prompt", "mission", "expectExact", "timeoutSeconds", "workspace", "capabilities", "session", "sessionForkFrom"], "task");
if (document.taskVersion !== 1) {
fail(2, `unsupported taskVersion: ${JSON.stringify(document.taskVersion)} (supported: 1)`);
}
@@ -144,6 +168,67 @@ function validateTask(document, file) {
timeoutSeconds = document.timeoutSeconds;
}
// Workspace (M5): absent = none; ":run" = ephemeral per-run; otherwise a
// persistent named workspace under <dataRoot>/workspaces/<name>.
let workspace = null;
if (document.workspace !== undefined && document.workspace !== null) {
if (typeof document.workspace !== "string" || document.workspace.length === 0) {
fail(2, 'task "workspace" must be a non-empty string when present');
}
if (document.workspace !== ":run") {
validateId(document.workspace, "task workspace");
}
workspace = document.workspace;
}
// Session (M6): optional named persistent session under
// <dataRoot>/sessions/<name>. Distinct names never share state.
let session = null;
if (document.session !== undefined && document.session !== null) {
if (typeof document.session !== "string" || document.session.length === 0) {
fail(2, 'task "session" must be a non-empty string when present');
}
validateId(document.session, "task session");
session = document.session;
}
// Session fork (M11): optional source session whose newest session file
// is branched (pi --fork) into the target session dir. Requires session.
let sessionForkFrom = null;
if (document.sessionForkFrom !== undefined && document.sessionForkFrom !== null) {
if (typeof document.sessionForkFrom !== "string" || document.sessionForkFrom.length === 0) {
fail(2, 'task "sessionForkFrom" must be a non-empty string when present');
}
validateId(document.sessionForkFrom, "task sessionForkFrom");
if (!session) {
fail(2, 'task "sessionForkFrom" requires "session" (the fork target) to be set');
}
if (document.sessionForkFrom === session) {
fail(2, 'task "sessionForkFrom" must differ from "session" (cannot fork onto itself)');
}
sessionForkFrom = document.sessionForkFrom;
}
// Capabilities (M5): optional tools allowlist mapped by adapters to their
// native permission flags. Absent = no tools.
let tools = null;
if (document.capabilities !== undefined && document.capabilities !== null) {
if (!isPlainObject(document.capabilities)) fail(2, '"capabilities" must be a JSON object');
rejectUnknownKeys(document.capabilities, ["tools"], '"capabilities"');
if (!Array.isArray(document.capabilities.tools) || document.capabilities.tools.length === 0) {
fail(2, '"capabilities.tools" must be a non-empty array of tool names');
}
const seen = new Set();
for (const tool of document.capabilities.tools) {
if (!SUPPORTED_TOOLS.includes(tool)) {
fail(2, `unsupported tool: ${JSON.stringify(tool)} (supported: ${SUPPORTED_TOOLS.join(", ")})`);
}
if (seen.has(tool)) fail(2, `duplicate tool in capabilities.tools: ${tool}`);
seen.add(tool);
}
tools = [...seen];
}
return {
taskVersion: document.taskVersion,
id: document.id,
@@ -153,6 +238,10 @@ function validateTask(document, file) {
missionSnapshot,
expectExact,
timeoutSeconds,
workspace,
tools,
session,
sessionForkFrom,
};
}
@@ -185,7 +274,7 @@ function writeOnce(file, content) {
}
}
function runTask(taskFile) {
function runTask(taskFile, options = {}) {
const resolved = JSON.parse(
spawnSync(process.execPath, [path.join(PROJECT_ROOT, "scripts", "mosaic-config.mjs"), "validate"], {
cwd: PROJECT_ROOT,
@@ -210,12 +299,89 @@ function runTask(taskFile) {
const startedAt = new Date();
const stderrFile = path.join(runDir, "stderr.txt");
// Sanctioned mission injection: point the container at the run snapshot's
// CONTAINER path (dataRoot maps to /var/lib/mosaic in the image).
const spawnEnv = { ...process.env };
spawnEnv.MOSAIC_ADAPTER = resolved.execution.adapter;
// Self-sufficient env: direct invocation (e.g. `retry`) skips the shell
// launcher exports, so derive them from the resolved config and release.
// (Names here are the compose interpolation consumers, not PI_*.)
spawnEnv.MOSAIC_PROVIDER = resolved.execution.provider;
spawnEnv.MOSAIC_MODEL = resolved.execution.model;
spawnEnv.MOSAIC_DATA_ROOT = resolved.dataRoot;
const release = fs.readFileSync(path.join(PROJECT_ROOT, "RELEASE"), "utf8").trim();
const piVersion = JSON.parse(fs.readFileSync(path.join(PROJECT_ROOT, "package.json"), "utf8")).dependencies["@earendil-works/pi-coding-agent"];
spawnEnv.MOSAIC_IMAGE_TAG = `mosaic-poc-agent:${piVersion}-r${release}`;
if (task.missionSnapshot) {
const relative = path.relative(resolved.dataRoot, runDir);
if (relative.startsWith("..") || path.isAbsolute(relative)) {
fail(4, `run directory is outside the configured dataRoot: ${runDir}`);
}
spawnEnv.MOSAIC_MISSION_FILE = `/var/lib/mosaic/${relative.split(path.sep).join("/")}/mission.json`;
}
// Workspace (M5): create host-side, pass the CONTAINER path.
let workspaceContainerPath = null;
if (task.workspace === ":run") {
fs.mkdirSync(path.join(runDir, "workspace"), { recursive: true });
workspaceContainerPath = `/var/lib/mosaic/runs/${runId}/workspace`;
} else if (task.workspace) {
fs.mkdirSync(path.join(resolved.dataRoot, "workspaces", task.workspace), { recursive: true });
workspaceContainerPath = `/var/lib/mosaic/workspaces/${task.workspace}`;
}
if (workspaceContainerPath) spawnEnv.MOSAIC_WORKSPACE = workspaceContainerPath;
// Capability policy (M9): least-privilege intersection. A task may narrow
// a mission's tool grant, never widen it. Empty intersection = tool-free.
let effectiveTools = task.tools;
let policyNote = null;
if (task.missionSnapshot?.capabilities) {
const missionTools = task.missionSnapshot.capabilities.tools;
if (effectiveTools) {
effectiveTools = effectiveTools.filter((t) => missionTools.includes(t));
if (effectiveTools.length === 0) {
policyNote = `capability policy: mission ${task.missionSnapshot.id} and task request no tools in common -> tool-free run`;
}
} else {
effectiveTools = [...missionTools];
}
}
if (policyNote) process.stderr.write(`mosaic-task: ${policyNote}\n`);
spawnEnv.MOSAIC_TOOLS = effectiveTools ? effectiveTools.join(",") : "";
// Session (M6): persistent named session dir, passed as container path.
if (task.session) {
fs.mkdirSync(path.join(resolved.dataRoot, "sessions", task.session), { recursive: true });
spawnEnv.MOSAIC_SESSION_DIR = `/var/lib/mosaic/sessions/${task.session}`;
}
// Session fork (M11): resolve the source session's newest file; pi --fork
// branches it into the target dir without modifying the ancestor.
if (task.sessionForkFrom) {
const sourceDir = path.join(resolved.dataRoot, "sessions", task.sessionForkFrom);
let sources = [];
try {
sources = fs.readdirSync(sourceDir).filter((f) => f.endsWith(".jsonl")).sort();
} catch {
fail(4, `cannot fork: source session dir not found: ${sourceDir}`);
}
if (sources.length === 0) {
fail(4, `cannot fork: source session '${task.sessionForkFrom}' has no session files`);
}
const relative = path.relative(resolved.dataRoot, path.join(sourceDir, sources[sources.length - 1]));
if (relative.startsWith("..") || path.isAbsolute(relative)) {
fail(4, `source session is outside the configured dataRoot: ${sourceDir}`);
}
spawnEnv.MOSAIC_SESSION_FORK = `/var/lib/mosaic/${relative.split(path.sep).join("/")}`;
}
const proc = spawnSync(
"docker",
["compose", "run", "--rm", "-T", "mosaic-agent", task.prompt],
{
cwd: PROJECT_ROOT,
env: process.env, // MOSAIC_DATA_ROOT / MOSAIC_PROVIDER / MOSAIC_MODEL resolved by run-task.sh
env: spawnEnv,
input: "", // stdin detached: print mode must never wait on a terminal (see issue #5)
encoding: "utf8",
maxBuffer: 16 * 1024 * 1024,
@@ -256,6 +422,11 @@ function runTask(taskFile) {
request: task.prompt,
response,
expectedExact: expected,
...(options.retriedFrom ? { retriedFrom: options.retriedFrom } : {}),
workspace: task.workspace,
tools: effectiveTools,
session: task.session,
sessionForkFrom: task.sessionForkFrom,
exitCode: proc.status,
signal: proc.signal ?? null,
provider: resolved.execution.provider,
@@ -283,17 +454,166 @@ function listRuns() {
for (const runId of entries) {
let status = "unknown";
let taskId = "-";
let workspace = "-";
let session = "-";
try {
const result = JSON.parse(fs.readFileSync(path.join(root, runId, "result.json"), "utf8"));
status = result.status;
taskId = result.taskId;
workspace = result.workspace ?? "-";
session = result.session ?? "-";
} catch {
// Incomplete run record; report as unknown.
}
process.stdout.write(`${runId} ${status.padEnd(9)} ${taskId}\n`);
process.stdout.write(`${runId} ${status.padEnd(9)} task=${taskId.padEnd(18)} ws=${String(workspace).padEnd(10)} session=${session}\n`);
}
}
function showRun(runId) {
if (!/^r-[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(runId)) {
fail(4, `invalid run id: ${JSON.stringify(runId)} (expected r-<id>)`);
}
const resolved = loadConfig();
const dir = path.join(runsRoot(resolved), runId);
if (!fs.existsSync(dir)) {
fail(4, `run not found: ${runId} (under ${runsRoot(resolved)})`);
}
const read = (name) => {
try {
return JSON.parse(fs.readFileSync(path.join(dir, name), "utf8"));
} catch {
return null;
}
};
const result = read("result.json");
const task = read("task.json");
const mission = read("mission.json");
process.stdout.write(`run: ${runId}\n`);
if (result) {
process.stdout.write(
[
`status: ${result.status}${result.reason ? ` (${result.reason})` : ""}`,
`task: ${result.taskId}`,
result.missionId ? `mission: ${result.missionId}` : null,
result.workspace ? `workspace: ${result.workspace}` : null,
result.session ? `session: ${result.session}` : null,
result.tools ? `tools: ${result.tools.join(", ")}` : null,
`adapter: ${"(see config)"} provider=${result.provider} model=${result.model}`,
`request: ${JSON.stringify(result.request)}`,
`response: ${JSON.stringify(result.response)}`,
result.expectedExact !== null && result.expectedExact !== undefined ? `expected: ${JSON.stringify(result.expectedExact)}` : null,
result.retriedFrom ? `retriedFrom: ${result.retriedFrom}` : null,
`timing: ${result.startedAt} -> ${result.finishedAt} (${result.durationMs} ms)`,
`exit: ${result.exitCode}${result.signal ? ` signal=${result.signal}` : ""}`,
].filter((line) => line !== null).join("\n") + "\n",
);
} else {
process.stdout.write("result.json: (missing or unreadable)\n");
}
if (task) process.stdout.write(`task snapshot: ${"task.json"} present\n`);
if (mission) process.stdout.write(`mission snapshot: ${mission.id} - ${mission.objective}\n`);
process.stdout.write(`artifacts: ${fs.readdirSync(dir).map((f) => `${f}`).join(", ")}\n`);
process.exit(0);
}
function retryRun(runId) {
if (!/^r-[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(runId)) {
fail(4, `invalid run id: ${JSON.stringify(runId)} (expected r-<id>)`);
}
const resolved = loadConfig();
const dir = path.join(runsRoot(resolved), runId);
if (!fs.existsSync(dir)) {
fail(4, `run not found: ${runId}`);
}
let snapshot;
try {
snapshot = fs.readFileSync(path.join(dir, "task.json"), "utf8");
} catch {
fail(4, `run task snapshot unreadable: ${runId}`);
}
// Relative mission paths in a snapshot resolve against the ORIGINAL task
// location, which no longer exists here — rewrite them to the run's own
// recorded mission.json so retries stay faithful.
let snapshotDoc;
try {
snapshotDoc = JSON.parse(snapshot);
} catch {
fail(4, `run task snapshot is not valid JSON: ${runId}`);
}
if (snapshotDoc.mission && !path.isAbsolute(snapshotDoc.mission)) {
const recordedMission = path.join(dir, "mission.json");
if (!fs.existsSync(recordedMission)) {
fail(4, `cannot retry ${runId}: relative mission path but no mission.json snapshot in run dir`);
}
snapshotDoc.mission = recordedMission;
snapshot = `${JSON.stringify(snapshotDoc, null, 2)}\n`;
}
// A retry is a brand-new run: replay the recorded task snapshot through
// the ordinary run path; existing run records stay untouched.
const tempDir = fs.mkdtempSync(path.join(os.tmpdir(), "mosaic-retry-"));
const tempTaskFile = path.join(tempDir, "task.json");
fs.writeFileSync(tempTaskFile, snapshot);
process.on("exit", () => {
try {
fs.rmSync(tempDir, { recursive: true, force: true });
} catch {
// Best-effort cleanup only.
}
});
runTask(tempTaskFile, { retriedFrom: runId });
}
function pruneRuns(args) {
const resolved = loadConfig();
const root = runsRoot(resolved);
let keep = 50;
let apply = false;
for (const arg of args) {
if (arg === "--yes") apply = true;
else if (arg === "--keep") fail(4, "prune: --keep requires a value");
else if (arg.startsWith("--keep=")) {
keep = Number(arg.slice("--keep=".length));
if (!Number.isInteger(keep) || keep < 1) fail(4, `--keep must be a positive integer (got ${arg.slice(7)})`);
} else fail(4, `unknown prune option: ${arg}`);
}
let entries = [];
try {
entries = fs.readdirSync(root, { withFileTypes: true })
.filter((e) => e.name.startsWith("r-") && e.isDirectory() && !e.isSymbolicLink())
.map((e) => e.name)
.sort();
} catch {
// No runs yet.
}
if (entries.length <= keep) {
process.stdout.write(`prune: ${entries.length} run(s) present, keep=${keep} -> nothing to prune\n`);
process.exit(0);
}
const doomed = entries.slice(0, entries.length - keep); // oldest first
if (!apply) {
process.stdout.write(`prune (dry-run): would remove ${doomed.length} oldest run(s), keep ${entries.length - doomed.length}:\n`);
for (const id of doomed) process.stdout.write(` would remove: ${id}\n`);
process.stdout.write("prune: re-run with --yes to apply\n");
process.exit(0);
}
const receipt = path.join(root, ".pruned.log");
for (const id of doomed) {
fs.rmSync(path.join(root, id), { recursive: true, force: true });
fs.appendFileSync(receipt, `${JSON.stringify({ at: new Date().toISOString(), event: "pruned", runId: id })}\n`);
}
process.stdout.write(`prune: removed ${doomed.length} run(s), kept ${keep}; receipt: ${receipt}\n`);
process.exit(0);
}
const operation = process.argv[2];
const target = process.argv[3];
@@ -310,9 +630,20 @@ switch (operation) {
if (!target) fail(4, "usage: mosaic-task.mjs run <taskFile>");
runTask(path.resolve(target));
break;
case "show":
if (!target) fail(4, "usage: mosaic-task.mjs show <runId>");
showRun(target);
break;
case "list":
listRuns();
process.exit(0);
case "prune":
pruneRuns(process.argv.slice(3));
break;
case "retry":
if (!target) fail(4, "usage: mosaic-task.mjs retry <runId>");
retryRun(target);
break;
default:
fail(4, `unknown operation: ${JSON.stringify(operation ?? "")} (expected validate | run | list)`);
fail(4, `unknown operation: ${JSON.stringify(operation ?? "")} (expected validate | run | show | list | retry | prune)`);
}
+149
View File
@@ -0,0 +1,149 @@
#!/usr/bin/env bash
# Release lifecycle: package, activate, rollback, status.
#
# scripts/release.sh package build the image for this release
# scripts/release.sh activate [--fault-injection]
# scripts/release.sh rollback
# scripts/release.sh status
#
# Activation is health-gated: the M2 task runner executes
# tasks/hello-marker.json; only an exact-marker pass activates. The
# pointer (state/active.json) is replaced atomically; every attempt is
# appended to state/activation-log.jsonl (append-only history).
#
# Fault injection exists solely to prove the refusal path in drills.
set -euo pipefail
cd "$(dirname "$0")/.."
# shellcheck source=common.sh
source scripts/common.sh
load_config
load_release
bootstrap_runtime_dir
STATE_DIR="$MOSAIC_DEV_DIR/state"
mkdir -p "$STATE_DIR"
POINTER="$STATE_DIR/active.json"
LOG="$STATE_DIR/activation-log.jsonl"
now() { date -u +%Y-%m-%dT%H:%M:%SZ; }
append_log() { # event release imageTag note
local event="$1" release="$2" imageTag="$3" note="${4:-}"
printf '{"at":"%s","event":"%s","release":"%s","imageTag":"%s"%s}\n' \
"$(now)" "$event" "$release" "$imageTag" \
"$(printf '%s' "$note" | node -e 'let d="";process.stdin.on("data",c=>d+=c).on("end",()=>{const s=d.replace(/\n$/,"");process.stdout.write(s ? ",\"note\":"+JSON.stringify(s) : "")})')" \
>> "$LOG"
}
image_exists() { docker image inspect "$1" >/dev/null 2>&1; }
health_check() { # returns 0 only when the marker path passes; $1 = fault injection label or empty
local tmp=""
local task="tasks/hello-marker.json"
if [ -n "${1:-}" ]; then
tmp="$(mktemp -d)"
# Fault injection: same prompt, deliberately wrong expectation.
printf '{"taskVersion":1,"id":"t-health-fault","prompt":"Return your startup marker and nothing else.","expectExact":"MOSAIC_FAULT_%s"}' \
"$RANDOM$RANDOM" > "$tmp/fault-task.json"
task="$tmp/fault-task.json"
fi
local rc=0
scripts/run-task.sh run "$task" >/dev/null 2>&1 || rc=$?
[ -n "$tmp" ] && rm -rf "$tmp"
return "$rc"
}
activate() { # $1 = release, $2 = imageTag, $3 = event name, $4 = fault label
local release="$1" imageTag="$2" event="$3" fault="${4:-}"
if ! image_exists "$imageTag"; then
echo "release: refusing $event: image not present locally: $imageTag" >&2
append_log "refused" "$release" "$imageTag" "image missing"
exit 1
fi
if ! health_check "$fault"; then
echo "release: refusing $event: health check failed" >&2
append_log "refused" "$release" "$imageTag" "health check failed${fault:+ (fault-injected)}"
exit 1
fi
# Atomic pointer replacement: write sibling temp file, then rename.
local tmp_pointer="$POINTER.tmp.$$"
printf '{"pointerVersion":1,"release":"%s","imageTag":"%s","activatedAt":"%s"}\n' \
"$release" "$imageTag" "$(now)" > "$tmp_pointer"
mv -f "$tmp_pointer" "$POINTER"
append_log "$event" "$release" "$imageTag"
echo "release: $event OK -> $release ($imageTag)"
}
previous_image_tag() { # last activated imageTag different from current pointer
[ -f "$POINTER" ] || return 1
local current
current="$(node -p 'JSON.parse(require("fs").readFileSync(process.argv[1],"utf8")).imageTag' "$POINTER")"
node -e '
const fs = require("fs");
const current = process.argv[1];
const lines = fs.readFileSync(process.argv[2], "utf8").split("\n").filter(Boolean);
for (let i = lines.length - 1; i >= 0; i--) {
let e;
try { e = JSON.parse(lines[i]); } catch { continue; }
if ((e.event === "activate" || e.event === "rollback") && e.imageTag && e.imageTag !== current) {
console.log(e.imageTag);
process.exit(0);
}
}
process.exit(1);
' "$current" "$LOG"
}
cmd_status() {
echo "release: $MOSAIC_RELEASE"
echo "image tag: $MOSAIC_IMAGE_TAG (packaged: $(image_exists "$MOSAIC_IMAGE_TAG" && echo yes || echo no))"
if [ -f "$POINTER" ]; then
node -e '
const p = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
console.log("active: " + p.release + " (" + p.imageTag + ") since " + p.activatedAt);
' "$POINTER"
else
echo "active: (none)"
fi
if [ -f "$LOG" ]; then
echo "recent log:"
tail -5 "$LOG" | sed 's/^/ /'
fi
}
cmd_package() {
docker compose build
append_log "package" "$MOSAIC_RELEASE" "$MOSAIC_IMAGE_TAG"
echo "release: packaged $MOSAIC_IMAGE_TAG"
}
cmd_activate() {
local fault=""
if [ "${1:-}" = "--fault-injection" ]; then fault="yes"; fi
activate "$MOSAIC_RELEASE" "$MOSAIC_IMAGE_TAG" "activate" "$fault"
}
cmd_rollback() {
local prev
if ! prev="$(previous_image_tag)"; then
echo "release: rollback: no previous activation found in log" >&2
exit 1
fi
local prev_release
prev_release="$(printf '%s' "$prev" | sed -n 's/.*-r\([0-9.]*\)$/\1/p')"
[ -n "$prev_release" ] || prev_release="unknown"
activate "$prev_release" "$prev" "rollback"
}
case "${1:-}" in
package) cmd_package ;;
activate) shift; cmd_activate "$@" ;;
rollback) cmd_rollback ;;
status) cmd_status ;;
*) echo "usage: scripts/release.sh package | activate [--fault-injection] | rollback | status" >&2; exit 4 ;;
esac
+1
View File
@@ -11,6 +11,7 @@ cd "$(dirname "$0")/.."
source scripts/common.sh
load_config
load_release # compose requires MOSAIC_IMAGE_TAG; task runs are release-scoped too
bootstrap_runtime_dir
exec node scripts/mosaic-task.mjs "$@"
+141
View File
@@ -0,0 +1,141 @@
#!/usr/bin/env bash
# Sandboxed selftests for the conductor auto-apply policy gate.
#
# Builds a throwaway target repo + worker workspace + fake run records, then
# exercises every gate: policy validation, allowlist, syntax gates, suite
# failure revert, disabled policy, missing/failed runs. No real model calls.
set -uo pipefail
cd "$(dirname "$0")/.."
SANDBOX="$(mktemp -d)"
trap 'rm -rf "$SANDBOX"' EXIT
PASS=0
FAIL=0
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
check_rc() { # name expectedRc command...
local name="$1" expected="$2"
shift 2
local rc
"$@" >/dev/null 2>&1
rc=$?
if [ "$rc" -eq "$expected" ]; then
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
else
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
fi
}
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
}
# ---- infrastructure: target repo + worker workspace + fake run ----
git clone -q . "$SANDBOX/repo"
# The clone carries committed state only - give the target its policy and
# commit it so the tree starts clean (untracked policy would fail target_clean).
mkdir -p "$SANDBOX/repo/roles"
cp roles/conductor-policy.json "$SANDBOX/repo/roles/conductor-policy.json"
git -C "$SANDBOX/repo" add roles/conductor-policy.json
git -C "$SANDBOX/repo" -c user.name=suite -c user.email=suite@local commit -q -m policy
mkdir -p "$SANDBOX/data/workspaces" "$SANDBOX/data/runs"
git clone -q "$SANDBOX/repo" "$SANDBOX/data/workspaces/stack-repo"
TARGET="$SANDBOX/repo"
WS="$SANDBOX/data/workspaces/stack-repo"
cat > "$SANDBOX/config.json" <<EOF
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"mock"}}
EOF
export MOSAIC_APPLY_TARGET="$TARGET"
export MOSAIC_CONFIG="$SANDBOX/config.json"
RUN_OK="r-20260903T000000000Z-ok0000001"
mkdir -p "$SANDBOX/data/runs/$RUN_OK"
printf '{"runVersion":1,"runId":"%s","taskId":"t-fake","status":"succeeded","workspace":"stack-repo"}' "$RUN_OK" \
> "$SANDBOX/data/runs/$RUN_OK/result.json"
ws_edit() { printf '\n%s\n' "$2" >> "$WS/$1"; }
ws_reset() { git -C "$WS" checkout -q -- . 2>/dev/null; git -C "$WS" clean -qfd; }
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
set_policy() { # enabled suites (commits: the target tree must stay clean)
local suites="[\"$2\"]"
printf '{"policyVersion":1,"autoApply":{"enabled":%s,"allowedPaths":["scripts/**","docs/**","README.md"],"suites":%s}}' "$1" "$suites" \
> "$TARGET/roles/conductor-policy.json"
git -C "$TARGET" add roles/conductor-policy.json
git -C "$TARGET" -c user.name=suite -c user.email=suite@local commit -q -m "policy update"
}
set_policy true "test-config"
# T1: dry run - allowed change, nothing applied
ws_edit "README.md" "worker dry-run line"
check_rc "dry-run: allowed change, exit 0, nothing committed" 0 \
scripts/conductor-apply.sh "$RUN_OK" --dry-run
if git -C "$TARGET" log --format=%s | grep -q "auto-applied"; then
check "dry-run committed nothing" 1
else
check "dry-run committed nothing" 0
fi
ws_reset
# T2: apply - allowed change, suites pass, commit created
ws_edit "README.md" "worker applied line"
check_rc "apply: allowed change exits 0" 0 scripts/conductor-apply.sh "$RUN_OK"
git -C "$TARGET" log -1 --format=%s | grep -q "auto-applied patch from run $RUN_OK" \
&& check "apply: attribution in commit subject" 0 || check "apply: attribution in commit subject" 1
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
target_clean && check "apply: target tree clean after commit" 0 || check "apply: target tree clean after commit" 1
git -C "$TARGET" reset -q --hard HEAD~1
# T3: disallowed path refused
ws_edit "Containerfile" "# worker touch"
check_rc "disallowed path refused" 1 scripts/conductor-apply.sh "$RUN_OK"
target_clean && check "disallowed path: target untouched" 0 || check "disallowed path: target untouched" 1
ws_reset
# T4: syntax gate - broken .mjs on an allowed path
printf 'this is not (valid js\n' > "$WS/scripts/broken-worker.mjs"
check_rc "syntax gate refused broken .mjs" 1 scripts/conductor-apply.sh "$RUN_OK"
target_clean && check "syntax gate: target untouched" 0 || check "syntax gate: target untouched" 1
ws_reset
# T5: suite failure - allowed change breaks a policy suite -> auto-revert
printf '\nexit 7\n' >> "$TARGET/scripts/test-config.sh"
ws_edit "README.md" "worker change that will fail suites"
check_rc "suite failure refused" 1 scripts/conductor-apply.sh "$RUN_OK"
git -C "$TARGET" checkout -q -- scripts/test-config.sh
target_clean && check "suite failure: target reverted to clean" 0 || check "suite failure: target reverted to clean" 1
# T6: disabled policy
set_policy false "test-config"
ws_edit "README.md" "worker line while disabled"
check_rc "disabled policy refused" 2 scripts/conductor-apply.sh "$RUN_OK"
target_clean && check "disabled policy: target untouched" 0 || check "disabled policy: target untouched" 1
ws_reset
set_policy true "test-config"
# T7: failed run refused
RUN_FAIL="r-20260903T000000000Z-fail00001"
mkdir -p "$SANDBOX/data/runs/$RUN_FAIL"
printf '{"runVersion":1,"runId":"%s","taskId":"t","status":"failed","workspace":"stack-repo"}' "$RUN_FAIL" \
> "$SANDBOX/data/runs/$RUN_FAIL/result.json"
ws_edit "README.md" "worker line from failed run"
check_rc "failed run refused" 1 scripts/conductor-apply.sh "$RUN_FAIL"
target_clean && check "failed run: target untouched" 0 || check "failed run: target untouched" 1
ws_reset
# T8/T9: missing run + invalid policy
check_rc "missing run exits 4" 4 scripts/conductor-apply.sh r-missing
printf '{"policyVersion":9}' > "$TARGET/roles/conductor-policy.json"
check_rc "invalid policy exits 2" 2 scripts/conductor-apply.sh "$RUN_OK"
git -C "$TARGET" checkout -q -- roles/conductor-policy.json
echo
echo "selftest: $PASS passed, $FAIL failed"
[ "$FAIL" -eq 0 ]
+34 -10
View File
@@ -10,8 +10,18 @@ SANDBOX="$(mktemp -d)"
trap 'rm -rf "$SANDBOX"' EXIT
PASS=0
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
FAIL=0
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
}
# expect_exit NAME EXPECTED_RC -- command...
expect_exit() {
local name="$1" expected="$2"
@@ -21,10 +31,10 @@ expect_exit() {
rc=$?
if [ "$rc" -eq "$expected" ]; then
PASS=$((PASS + 1))
echo "ok $name (exit $rc)"
echo "${C_OK}OK${C_RESET} $name (exit $rc)"
else
FAIL=$((FAIL + 1))
echo "FAIL $name (exit $rc, expected $expected)"
echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
fi
}
@@ -40,12 +50,26 @@ cfg() { printf '%s' "$2" > "$SANDBOX/$1"; }
DATA_ROOT="$SANDBOX/data"
# --- adapter selection (M4) ---
cfg default-adapter.json '{"configVersion":1,"environment":"development","dataRoot":"'$DATA_ROOT'","execution":{"backend":"docker","provider":"zai","model":"m"}}'
MOSAIC_CONFIG="$SANDBOX/default-adapter.json" $CONFIG_OP validate | grep -q '"adapter": "pi"'
check "absent adapter defaults to pi" $?
cfg mock-adapter.json '{"configVersion":1,"environment":"development","dataRoot":"'$DATA_ROOT'","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"mock"}}'
expect_exit "adapter mock validates" 0 -- env MOSAIC_CONFIG="$SANDBOX/mock-adapter.json" $CONFIG_OP validate
cfg bad-adapter.json '{"configVersion":1,"environment":"development","dataRoot":"'$DATA_ROOT'","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"claude"}}'
expect_exit "unsupported adapter exits 2" 2 -- env MOSAIC_CONFIG="$SANDBOX/bad-adapter.json" $CONFIG_OP validate
MOSAIC_CONFIG="$SANDBOX/mock-adapter.json" $CONFIG_OP env | grep -q "MOSAIC_ADAPTER='mock'"
check "env exports adapter" $?
# --- bootstrap ---
rm -f "$SANDBOX/config.json"
expect_exit "bootstrap creates default when absent" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/config.json" $CONFIG_OP bootstrap
[ -f "$SANDBOX/config.json" ] && { PASS=$((PASS+1)); echo "ok bootstrap wrote config file"; } \
|| { FAIL=$((FAIL+1)); echo "FAIL bootstrap wrote config file"; }
[ -f "$SANDBOX/config.json" ] && { PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} bootstrap wrote config file"; } \
|| { FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} bootstrap wrote config file"; }
SUM_BEFORE=$(sha256sum "$SANDBOX/config.json" | cut -d' ' -f1)
MTIME_BEFORE=$(stat -c %Y "$SANDBOX/config.json")
@@ -55,9 +79,9 @@ expect_exit "bootstrap is idempotent on existing config" 0 -- \
SUM_AFTER=$(sha256sum "$SANDBOX/config.json" | cut -d' ' -f1)
MTIME_AFTER=$(stat -c %Y "$SANDBOX/config.json")
if [ "$SUM_BEFORE" = "$SUM_AFTER" ] && [ "$MTIME_BEFORE" = "$MTIME_AFTER" ]; then
PASS=$((PASS+1)); echo "ok bootstrap did not rewrite existing config"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} bootstrap did not rewrite existing config"
else
FAIL=$((FAIL+1)); echo "FAIL bootstrap rewrote existing config"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} bootstrap rewrote existing config"
fi
# --- validate ---
@@ -122,9 +146,9 @@ cfg valid.json "$(valid_body "$DATA_ROOT")"
EVAL_OUT="$(MOSAIC_CONFIG="$SANDBOX/valid.json" $CONFIG_OP env)" || true
if eval "$EVAL_OUT" 2>/dev/null && [ "$MOSAIC_DATA_ROOT" = "$DATA_ROOT" ] \
&& [ "$MOSAIC_PROVIDER" = "zai" ] && [ "$MOSAIC_MODEL" = "glm-5.3-flash" ]; then
PASS=$((PASS+1)); echo "ok env exports resolve correctly"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} env exports resolve correctly"
else
FAIL=$((FAIL+1)); echo "FAIL env exports resolve correctly"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} env exports resolve correctly"
fi
# --- validation must not modify the file ---
@@ -132,9 +156,9 @@ SUM_INVALID_BEFORE=$(sha256sum "$SANDBOX/invalid.json" | cut -d' ' -f1)
MOSAIC_CONFIG="$SANDBOX/invalid.json" $CONFIG_OP validate >/dev/null 2>&1
SUM_INVALID_AFTER=$(sha256sum "$SANDBOX/invalid.json" | cut -d' ' -f1)
if [ "$SUM_INVALID_BEFORE" = "$SUM_INVALID_AFTER" ]; then
PASS=$((PASS+1)); echo "ok failed validation modified nothing"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} failed validation modified nothing"
else
FAIL=$((FAIL+1)); echo "FAIL failed validation modified the file"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} failed validation modified the file"
fi
echo
+104
View File
@@ -0,0 +1,104 @@
#!/usr/bin/env bash
# Sandboxed selftests for the release layer.
#
# Fast cases (version validation) need no Docker. State-machine cases
# (status/activate/refusal) run against a sandboxed config and therefore
# require the Docker daemon; they are skipped when it is unavailable.
set -uo pipefail
cd "$(dirname "$0")/.."
SANDBOX="$(mktemp -d)"
RELEASE_BACKUP="$(mktemp)"
cp RELEASE "$RELEASE_BACKUP"
# One exit trap: the repo RELEASE is ALWAYS restored from the backup,
# regardless of how the test run ends.
trap 'cp "$RELEASE_BACKUP" RELEASE 2>/dev/null; rm -rf "$SANDBOX" "$RELEASE_BACKUP"' EXIT
PASS=0
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
FAIL=0
expect_exit() {
local name="$1" expected="$2"
shift 3
local rc
"$@" >/dev/null 2>&1
rc=$?
if [ "$rc" -eq "$expected" ]; then
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
else
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
fi
}
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
}
# ---------- fast: release identity ----------
expect_exit "valid RELEASE resolves" 0 -- bash -c 'source scripts/common.sh && load_release'
printf 'garbage\n' > RELEASE
expect_exit "invalid RELEASE exits 1" 1 -- bash -c 'source scripts/common.sh && load_release'
mv RELEASE "$SANDBOX/RELEASE.hidden"
expect_exit "missing RELEASE exits 1" 1 -- bash -c 'source scripts/common.sh && load_release'
cp "$RELEASE_BACKUP" RELEASE
bash -c 'source scripts/common.sh && load_release' >/dev/null 2>&1
bash -c 'source scripts/common.sh && load_release && case "$MOSAIC_IMAGE_TAG" in mosaic-poc-agent:*-r'"$(cat RELEASE)"') exit 0;; *) exit 1;; esac' >/dev/null 2>&1
check "valid RELEASE leaves image tag consistent with version" $?
# ---------- sandboxed state machine (Docker required) ----------
if docker info >/dev/null 2>&1; then
mkdir -p "$SANDBOX/data"
cat > "$SANDBOX/config.json" <<EOF
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"glm-5.3-flash"}}
EOF
export MOSAIC_CONFIG="$SANDBOX/config.json"
expect_exit "status safe on empty state" 0 -- scripts/release.sh status
[ ! -e "$SANDBOX/data/state/active.json" ] \
&& check "status created no pointer" 0 || check "status created no pointer" 1
expect_exit "fault-injected activation refuses" 1 -- scripts/release.sh activate --fault-injection
[ ! -e "$SANDBOX/data/state/active.json" ] \
&& check "refused activation wrote no pointer" 0 || check "refused activation wrote no pointer" 1
if [ -f "$SANDBOX/data/state/activation-log.jsonl" ]; then
node -e '
const fs = require("fs");
const lines = fs.readFileSync(process.argv[1], "utf8").split("\n").filter(Boolean);
if (lines.length !== 1) process.exit(1);
const e = JSON.parse(lines[0]);
process.exit(e.event === "refused" && e.release && e.imageTag && e.at ? 0 : 1);
' "$SANDBOX/data/state/activation-log.jsonl"
check "refusal logged exactly once with valid fields" $?
else
check "refusal logged exactly once with valid fields" 1
fi
expect_exit "healthy activation succeeds" 0 -- scripts/release.sh activate
node -e '
const fs = require("fs");
const p = JSON.parse(fs.readFileSync(process.argv[1], "utf8"));
process.exit(p.pointerVersion === 1 && p.release && p.imageTag && p.activatedAt ? 0 : 1);
' "$SANDBOX/data/state/active.json"
check "pointer written with valid fields" $?
expect_exit "repeat activation succeeds (log grows)" 0 -- scripts/release.sh activate
LINES=$(grep -c '' "$SANDBOX/data/state/activation-log.jsonl")
[ "$LINES" -ge 3 ] && check "log is append-only across activations" 0 || check "log is append-only across activations" 1
expect_exit "rollback without previous refuses" 1 -- scripts/release.sh rollback
else
echo "skip state-machine cases (docker daemon unavailable)"
fi
echo
echo "selftest: $PASS passed, $FAIL failed"
[ "$FAIL" -eq 0 ]
+230 -9
View File
@@ -11,6 +11,12 @@ SANDBOX="$(mktemp -d)"
trap 'rm -rf "$SANDBOX"' EXIT
PASS=0
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
FAIL=0
expect_exit() {
@@ -20,14 +26,20 @@ expect_exit() {
"$@" >/dev/null 2>&1
rc=$?
if [ "$rc" -eq "$expected" ]; then
PASS=$((PASS+1)); echo "ok $name (exit $rc)"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
else
FAIL=$((FAIL+1)); echo "FAIL $name (exit $rc, expected $expected)"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
fi
}
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
}
latest_reason() {
local latest
latest="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
[ -n "$latest" ] && node -e 'try{const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.reason??"")}catch{console.log("")}' "$latest/result.json" 2>/dev/null
}
CONFIG="$SANDBOX/config.json"
@@ -92,13 +104,200 @@ $TASK validate "$SANDBOX/ok.json" >/dev/null 2>&1
M2=$(stat -c %Y "$SANDBOX/ok.json")
[ "$M1" = "$M2" ] && check "validation does not modify the task file" 0 || check "validation does not modify the task file" 1
# ---------- live: real runs (Docker + credentials required) ----------
# ---------- retention: prune (deterministic, no Docker) ----------
mkdir -p "$SANDBOX/data"
cat > "$SANDBOX/prune-config.json" <<EOF
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m"}}
EOF
for i in 1 2 3 4 5; do
D="$SANDBOX/data/runs/r-20260903T0100_0${i}Z-suite00$i"
mkdir -p "$D"
printf '{"runVersion":1,"runId":"r-20260903T0100_0%sZ-suite00%s","taskId":"t","status":"succeeded"}' "$i" "$i" > "$D/result.json"
done
mkdir -p "$SANDBOX/data/sessions/sentinel" "$SANDBOX/data/workspaces/sentinel"
expect_exit "prune dry-run exits 0" 0 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune
[ "$(ls "$SANDBOX/data/runs" | grep -c '^r-')" -eq 5 ] \
&& check "dry-run deleted nothing" 0 || check "dry-run deleted nothing" 1
expect_exit "prune --keep=2 --yes removes oldest" 0 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune --keep=2 --yes
[ "$(ls "$SANDBOX/data/runs" | grep -c '^r-')" -eq 2 ] \
&& check "kept exactly 2 newest runs" 0 || check "kept exactly 2 newest runs" 1
NEWEST="r-20260903T0100_05Z-suite005"
[ -d "$SANDBOX/data/runs/$NEWEST" ] \
&& check "newest run kept, oldest pruned" 0 || check "newest run kept, oldest pruned" 1
[ -f "$SANDBOX/data/runs/.pruned.log" ] \
&& [ "$(grep -c 'pruned' "$SANDBOX/data/runs/.pruned.log")" -eq 3 ] \
&& check "append-only receipt written (3 entries)" 0 \
|| check "append-only receipt written (3 entries)" 1
[ -d "$SANDBOX/data/sessions/sentinel" ] && [ -d "$SANDBOX/data/workspaces/sentinel" ] \
&& check "sessions/workspaces untouched by prune" 0 \
|| check "sessions/workspaces untouched by prune" 1
expect_exit "prune with invalid keep exits 4" 4 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune --keep=0 --yes
# ---------- adapter seam: deterministic mock cases (Docker, no provider) ----------
if docker info >/dev/null 2>&1; then
expect_exit "live hello task succeeds with exact marker" 0 -- \
good_task "$SANDBOX/ok.json"
mock_config() { # file adapter
cat > "$SANDBOX/$1" <<EOF
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"$2"}}
EOF
}
mock_config mock-adapters.json mock
expect_exit "mock adapter: gate passes on matching mock response" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOSAIC_HELLO_OK \
scripts/run-task.sh run "$SANDBOX/ok.json"
[ -f "$DATA_ROOT/runs" ] && RUNS1=$(ls "$DATA_ROOT/runs" | wc -l)
R1="$(ls "$DATA_ROOT/runs" | head -1)"
expect_exit "mock adapter: expect-mismatch recorded" 1 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=SOMETHING_ELSE \
scripts/run-task.sh run "$SANDBOX/ok.json"
[ "$(latest_reason)" = "expect-mismatch" ] \
&& check "mismatch reason recorded" 0 || check "mismatch reason recorded" 1
mock_config bad-adapter.json nonexistent
# The config layer rejects unknown adapters first, so run-task.sh fails
# closed (exit 1) before any container or run record exists.
expect_exit "unknown adapter fails closed" 1 -- \
env MOSAIC_CONFIG="$SANDBOX/bad-adapter.json" scripts/run-task.sh run "$SANDBOX/ok.json"
good_mission "$SANDBOX/m-ok.json"
printf '{"taskVersion":1,"id":"t-mission","prompt":"ignored by mock","mission":"m-ok.json"}' > "$SANDBOX/mission-task.json"
# Dedicated mission with distinctive directives for the injection assertion.
cat > "$SANDBOX/m-seam.json" <<'EOF'
{"missionVersion":1,"id":"m-seam","objective":"Prove the mission injection point.","directives":["Seam directive A.","Seam directive B."]}
EOF
printf '{"taskVersion":1,"id":"t-mission-seam","prompt":"ignored by mock","mission":"m-seam.json"}' > "$SANDBOX/seam-task.json"
expect_exit "mission task runs via mock adapter" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
scripts/run-task.sh run "$SANDBOX/seam-task.json"
grep -q 'MISSION (runtime)' "$SANDBOX/data/system-prompt.md" \
&& grep -q 'Seam directive A.' "$SANDBOX/data/system-prompt.md" \
&& check "mission section injected into generated prompt" 0 \
|| check "mission section injected into generated prompt" 1
# retry lineage + relative mission path resolution
expect_exit "retry of mission run succeeds" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
node scripts/mosaic-task.mjs retry "$(ls -dt "$SANDBOX/data/runs"/r-* | head -1 | xargs basename)"
RT="$(ls -dt "$SANDBOX/data/runs"/r-* | head -1 | xargs basename)"
ORIG="$(ls -dt "$SANDBOX/data/runs"/r-* | sed -n 2p | xargs basename)"
node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));process.exit(r.retriedFrom===process.argv[2]?0:1)' \
"$SANDBOX/data/runs/$RT/result.json" "$ORIG" \
&& check "retriedFrom lineage recorded" 0 || check "retriedFrom lineage recorded" 1
grep -q 'MISSION (runtime)' "$SANDBOX/data/system-prompt.md" \
&& check "mission section present after retry (relative path resolved)" 0 \
|| check "mission section present after retry (relative path resolved)" 1
expect_exit "retry of missing run exits 4" 4 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" node scripts/mosaic-task.mjs retry r-missing
# session fork plumbing (M11): fork source + target dir delivered
printf '{"taskVersion":1,"id":"t-forkplumb","prompt":"x","session":"fork-child","sessionForkFrom":"base"}' > "$SANDBOX/forkplumb.json"
mkdir -p "$SANDBOX/data/sessions/base"
printf '{}' > "$SANDBOX/data/sessions/base/20260903T000000-plumb.jsonl"
expect_exit "fork task runs via mock" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
scripts/run-task.sh run "$SANDBOX/forkplumb.json"
FL="$(ls -dt "$SANDBOX/data/runs"/r-* | head -1)"
grep -q '^MOSAIC_SESSION_FORK=/var/lib/mosaic/sessions/base/20260903T000000-plumb.jsonl$' "$FL/stderr.txt" 2>/dev/null \
&& grep -q '^MOSAIC_SESSION_DIR=/var/lib/mosaic/sessions/fork-child$' "$FL/stderr.txt" 2>/dev/null \
&& check "fork source + target delivered to adapter" 0 \
|| check "fork source + target delivered to adapter" 1
printf '{"taskVersion":1,"id":"t-f2","prompt":"x","sessionForkFrom":"base"}' > "$SANDBOX/noforktarget.json"
expect_exit "fork without session target exits 2" 2 -- $TASK validate "$SANDBOX/noforktarget.json"
printf '{"taskVersion":1,"id":"t-f3","prompt":"x","session":"base","sessionForkFrom":"base"}' > "$SANDBOX/selfork.json"
expect_exit "self-fork exits 2" 2 -- $TASK validate "$SANDBOX/selfork.json"
# capability policy (M9): least-privilege intersection
POL="$SANDBOX/data/workspaces"; mkdir -p "$POL"
pol_run() { # missionTools(ABSENT|json) taskTools(ABSENT|json) -> stderr MOSAIC_TOOLS value
if [ "$1" = "ABSENT" ]; then
printf '{"missionVersion":1,"id":"m-pol","objective":"o"}' > "$SANDBOX/pol-m.json"
else
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":[%s]}}' "$1" > "$SANDBOX/pol-m.json"
fi
if [ "$2" = "ABSENT" ]; then
printf '{"taskVersion":1,"id":"t-pol","prompt":"x","mission":"pol-m.json"}' > "$SANDBOX/pol-t.json"
else
printf '{"taskVersion":1,"id":"t-pol","prompt":"x","mission":"pol-m.json","capabilities":{"tools":[%s]}}' "$2" > "$SANDBOX/pol-t.json"
fi
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
scripts/run-task.sh run "$SANDBOX/pol-t.json" >/dev/null 2>&1
grep -o '^MOSAIC_TOOLS=.*' "$(ls -dt "$SANDBOX/data/runs"/r-* | head -1)/stderr.txt" 2>/dev/null
}
[ "$(pol_run '"read","bash"' ABSENT)" = 'MOSAIC_TOOLS=read,bash' ] \
&& check "policy: mission only -> mission tools" 0 || check "policy: mission only -> mission tools" 1
[ "$(pol_run ABSENT '"read","bash"')" = 'MOSAIC_TOOLS=read,bash' ] \
&& check "policy: task only -> task tools" 0 || check "policy: task only -> task tools" 1
[ "$(pol_run '"read"' '"bash","read"')" = 'MOSAIC_TOOLS=read' ] \
&& check "policy: both -> intersection (task narrowed)" 0 || check "policy: both -> intersection (task narrowed)" 1
[ "$(pol_run '"grep"' '"bash","read"')" = 'MOSAIC_TOOLS=' ] \
&& check "policy: empty intersection -> tool-free" 0 || check "policy: empty intersection -> tool-free" 1
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":["sudo"]}}' > "$SANDBOX/pol-m.json"
expect_exit "invalid mission capabilities rejected" 2 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" $TASK validate "$SANDBOX/pol-t.json"
else
echo "skip adapter seam cases (docker daemon unavailable)"
fi
# ---------- workspace + capabilities (M5): deterministic mock cases ----------
if docker info >/dev/null 2>&1; then
printf '{"taskVersion":1,"id":"t-ws","prompt":"ignored","workspace":"suitews","capabilities":{"tools":["read","bash"]}}' > "$SANDBOX/ws-task.json"
expect_exit "workspace+tools task runs via mock" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
scripts/run-task.sh run "$SANDBOX/ws-task.json"
WSL="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
grep -q '^MOSAIC_WORKSPACE=/var/lib/mosaic/workspaces/suitews$' "$WSL/stderr.txt" 2>/dev/null \
&& grep -q '^MOSAIC_TOOLS=read,bash$' "$WSL/stderr.txt" 2>/dev/null \
&& check "workspace path + tools delivered to adapter" 0 \
|| check "workspace path + tools delivered to adapter" 1
[ -d "$SANDBOX/data/workspaces/suitews" ] \
&& check "persistent workspace created on host" 0 \
|| check "persistent workspace created on host" 1
expect_exit "plain task still runs (no workspace/tools)" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOSAIC_HELLO_OK \
scripts/run-task.sh run "$SANDBOX/ok.json"
PL="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
grep -q '^MOSAIC_WORKSPACE=$' "$PL/stderr.txt" 2>/dev/null \
&& check "workspace var present but empty when absent" 0 || check "workspace var present but empty when absent" 1
grep -q '^MOSAIC_TOOLS=$' "$PL/stderr.txt" 2>/dev/null \
&& check "tools empty when absent" 0 || check "tools empty when absent" 1
printf '{"taskVersion":1,"id":"t-badtool","prompt":"x","capabilities":{"tools":["sudo"]}}' > "$SANDBOX/badtool.json"
expect_exit "unknown tool exits 2" 2 -- $TASK validate "$SANDBOX/badtool.json"
printf '{"taskVersion":1,"id":"t-badws","prompt":"x","workspace":"../escape"}' > "$SANDBOX/badws.json"
expect_exit "workspace traversal exits 2" 2 -- $TASK validate "$SANDBOX/badws.json"
else
echo "skip workspace/capability cases (docker daemon unavailable)"
fi
# ---------- live: real runs (Docker + credentials required) ----------
# On failure, surface the run record + agent stderr BEFORE the sandbox
# cleanup destroys them. Never let a wrong-exit mask the real reason.
dump_latest_run() {
local latest
latest="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
if [ -n "$latest" ]; then
echo "--- latest run evidence: $latest ---" >&2
cat "$latest/result.json" 2>/dev/null >&2
echo "--- stderr.txt (tail) ---" >&2
tail -8 "$latest/stderr.txt" 2>/dev/null >&2
else
echo "--- no run dir was created at all ---" >&2
fi
}
if docker info >/dev/null 2>&1; then
RUNS1=$(ls "$DATA_ROOT/runs" 2>/dev/null | wc -l)
if scripts/run-task.sh run "$SANDBOX/ok.json" >/dev/null 2>&1; then
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} live hello task succeeds with exact marker"
else
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} live hello task succeeds with exact marker" >&2
dump_latest_run
fi
R1="$(ls -t "$DATA_ROOT/runs" | head -1)" # newest = the live run above
[ -f "$DATA_ROOT/runs/$R1/result.json" ] && check "result.json written in run dir" 0 || check "result.json written in run dir" 1
node -e '
const r = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
@@ -107,8 +306,15 @@ process.exit(r.status === "succeeded" && r.response === "MOSAIC_HELLO_OK" && r.e
check "result.json contents are correct" $?
printf '{"taskVersion":1,"id":"t-wrong","prompt":"Return your startup marker and nothing else.","expectExact":"MOSAIC_NOT_OK"}' > "$SANDBOX/wrong.json"
expect_exit "wrong expectExact fails with exit 1" 1 -- \
scripts/run-task.sh run "$SANDBOX/wrong.json"
scripts/run-task.sh run "$SANDBOX/wrong.json" >/dev/null 2>&1
RC=$?
WRONG_REASON="$(latest_reason)"
if [ "$RC" -eq 1 ] && [ "$WRONG_REASON" = "expect-mismatch" ]; then
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} wrong expectExact fails with exit 1 (reason: expect-mismatch)"
else
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} wrong expectExact (exit $RC, reason: '${WRONG_REASON:-none}')" >&2
dump_latest_run
fi
RUNS2=$(ls "$DATA_ROOT/runs" | wc -l)
[ "$RUNS2" -gt "${RUNS1:-0}" ] && check "each run gets a distinct run dir (no clobber)" 0 \
@@ -116,6 +322,21 @@ process.exit(r.status === "succeeded" && r.response === "MOSAIC_HELLO_OK" && r.e
COUNT=$($TASK list | wc -l)
[ "$COUNT" -ge 2 ] && check "list shows both runs" 0 || check "list shows both runs" 1
# session fork (M11): teach in base, fork into child, child recalls;
# ancestor file count must be unchanged by the fork
printf '{"taskVersion":1,"id":"t-fork-base","prompt":"Remember this code word for later: mosaico. Reply with exactly: REMEMBERED","session":"suite-base","expectExact":"REMEMBERED","timeoutSeconds":180}' > "$SANDBOX/fb.json"
expect_exit "fork base: teach succeeds" 0 -- scripts/run-task.sh run "$SANDBOX/fb.json"
BASECOUNT=$(ls "$SANDBOX/data/sessions/suite-base" | wc -l)
printf '{"taskVersion":1,"id":"t-fork-child","prompt":"What code word did I ask you to remember? Reply with only the code word.","session":"suite-child","sessionForkFrom":"suite-base","timeoutSeconds":180}' > "$SANDBOX/fc.json"
expect_exit "fork child recalls ancestor context" 0 -- scripts/run-task.sh run "$SANDBOX/fc.json"
FR="$(ls -dt "$DATA_ROOT/runs"/r-* | head -1)"
node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));process.exit(String(r.response).toLowerCase().includes("mosaico")?0:1)' "$FR/result.json" \
&& check "forked child recalled ancestor code word" 0 || check "forked child recalled ancestor code word" 1
[ "$(ls "$SANDBOX/data/sessions/suite-base" | wc -l)" -eq "$BASECOUNT" ] \
&& check "ancestor session untouched by fork" 0 || check "ancestor session untouched by fork" 1
[ "$(ls "$SANDBOX/data/sessions/suite-child" | wc -l)" -ge 1 ] \
&& check "child session dir has its own branch file" 0 || check "child session dir has its own branch file" 1
else
echo "skip live task cases (docker unavailable)"
fi
+11 -3
View File
@@ -14,10 +14,18 @@ cd "$(dirname "$0")/.."
# shellcheck source=common.sh
source scripts/common.sh
IMAGE="mosaic-poc-agent:0.84.4"
EXPECTED="${EXPECTED_MARKER:-MOSAIC_HELLO_OK}"
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
load_config
load_release
IMAGE="$MOSAIC_IMAGE_TAG"
# Ensure the configured data root exists (host-owned) before the mount,
# otherwise Docker would auto-create a root-owned directory.
@@ -50,11 +58,11 @@ TRIMMED="$(printf '%s' "$RESPONSE" | sed -e 's/^[[:space:]]*//' -e 's/[[:space:]
# 4-6. Exact comparison gate.
if [ "$TRIMMED" = "$EXPECTED" ]; then
echo "PASS: response matches expected marker"
echo "${C_OK}PASS${C_RESET}: response matches expected marker"
exit 0
fi
echo "FAIL: response does not match expected marker" >&2
echo "${C_FAIL}FAIL${C_RESET}: response does not match expected marker" >&2
printf 'expected: %s\n' "$EXPECTED" >&2
printf 'actual : %s\n' "$TRIMMED" >&2
exit 1
+14
View File
@@ -0,0 +1,14 @@
---
name: ms-tools
description: Contains a reference to all available tools for Mosaic Stack.
disable-model-invocation: true
---
# ms-tools
You are a Mosaic fleet agent. A maintained toolkit lives at `<dir>`.
Use it FIRST for the tasks below — improvising with raw CLIs causes the recurring failures this
skill exists to prevent. This is the high-frequency fast path; the full reference is the
`# Machine Tools` section already in your system prompt.
+85
View File
@@ -0,0 +1,85 @@
---
name: ms-unslop
description: Cut AI tells from any writing. Must always apply.
disable-model-invocation: false
---
# Unslop
Edit text to remove AI patterns and add human voice.
## Process
1. Scan for the patterns below.
2. Rewrite. Preserve meaning, match intended tone.
3. Add soul (see next section).
4. Self-audit: "What makes this obviously AI generated?" Fix remaining tells.
## Adding soul
Removing patterns is half the job. Sterile, voiceless writing is just as obvious.
- **Have opinions.** React to facts instead of neutrally listing pros and cons.
- **Vary rhythm.** Short sentences. Then longer ones that take their time. Mix it up.
- **Acknowledge complexity.** "Impressive but also kind of unsettling" beats "impressive."
- **Use "I" when it fits.** First person isn't unprofessional.
- **Let some mess in.** Perfect structure looks machine-made.
- **Be specific.** Not "this is concerning" but "there's something unsettling about agents churning away at 3am."
## Patterns to detect and fix
### Content
1. **Puffery.** `pivotal moment`, `testament to`, "evolving landscape", "setting the stage for", "indelible mark", "deeply rooted". Cut puffery, state what happened.
2. **Name-dropping.** Listing media outlets without context. Pick one, say what was said.
3. **Superficial -ing phrases.** "highlighting...", "ensuring...", "reflecting...", "showcasing...", "fostering...". Delete or expand with real sources.
4. **Promotional language.** "nestled", `vibrant`, "breathtaking", "groundbreaking", "renowned", "stunning", "must-visit". Use neutral descriptions.
5. **Vague attributions.** "Experts believe", "Industry reports suggest", "Some critics argue". Name the source or delete.
6. **Formulaic challenges.** "Despite challenges... continues to thrive." Replace with specific facts.
### Language
7. **AI vocabulary.** `Additionally`, `crucial`, `delve`, `enduring`, `enhance`, `fostering`, `garner`, `interplay`, `intricate`, `landscape` (abstract), `pivotal`, `showcase`, `tapestry` (abstract), `testament`, `underscore`, `vibrant`. Replace with plain words.
8. **Fancy ways to say "is".** "serves as", "stands as", "boasts", "features". Just say "is" or "has".
9. **`Not just X, but Y`.** State the point directly instead.
10. **Rule of three.** Forcing ideas into groups of three. Use the natural number.
11. **Synonym cycling.** Protagonist, main character, central figure, hero all in one paragraph. Pick one, repeat it.
12. **False ranges.** "from X to Y" where X and Y aren't on a meaningful scale. List topics directly.
### Style
13. **Em dash overuse.** Avoid em dashes entirely. Use periods or commas only (no parentheses, no en dashes, no hyphen-as-dash substitutes). Em dashes are an AI tell, and reaching for parentheses instead just trades one tell for another. If a thought needs separation, end the sentence or use a comma.
14. **Colon overuse.** Colons are fine before a list or example. Not as mid-sentence connectors. "If you're coming from traditional automation: instead of registering event handlers, you describe conditions" adds nothing with the colon. Rewrite to let the point stand on its own without comparison framing. "Describing when the scheduler should fire works best as plain English." Same meaning, no crutch punctuation.
15. **Boldface overuse.** Don't bold every proper noun or acronym.
16. **Inline-header lists.** The tell is a bold label and colon that restates the line: "**Performance:** Performance improved...". Convert those to prose. A bold lead-in that ends in a period, names the item, and is followed by genuinely new detail ("**Schema in TypeScript.** Tables live in one file.") is fine, not a tell.
17. **Title case headings.** Use sentence case.
18. **Decorative emojis.** Remove from headings and bullets.
19. **Curly quotes.** Replace with straight quotes.
### Communication artifacts
20. **Chatbot phrases.** `I hope this helps!`, `Let me know if...`, `Of course!`, `Certainly!`, `Found the smoking gun!` Remove.
21. **Cutoff disclaimers.** "While specific details are limited..." Find sources or remove.
22. **Sycophantic tone.** `Great question!` `You're absolutely right!` Respond directly.
### Filler
23. **Filler phrases.** `In order to` becomes "To". `Due to the fact that` becomes "Because". `It is important to note that` gets deleted.
24. **Excessive hedging.** "could potentially possibly be argued that it might" becomes "may".
25. **Generic conclusions.** "The future looks bright." State specific plans or facts.
### Jargon
26. **Abstract metaphor nouns.** Substrate, wedge, vector, locus, vantage, nexus, primitive (as noun), harness (as metaphor), surface (as in "API surface"), bedrock, scaffolding (as metaphor), modality, paradigm, gold-plating, ratchet (as metaphor), evacuate (for moving code), endgame, north star, flywheel. These read as technical but usually have a plainer concrete word. "Substrate" becomes "base". "Wedge in" becomes "add". "Vector" becomes "way" or "method". "Gold-plating" becomes "more than the job needs". "Ratchet" becomes the mechanism's real name or "a limit that only tightens". "Evacuate" becomes "move out". "Endgame" becomes "the last phase". Pick the concrete word.
### Plain speech
27. **Say what it does, not how it feels.** "the database stays close at hand", "SQL you can read", "types that follow your schema" name a feeling. The fix names the mechanism or a number: "`.toSQL()` returns the exact string sent to the database", "a column rename fails the build". Ask what the sentence tells the reader to do or know, then write that. If you can't restate it as a concrete instruction, fact, or number, cut it. One more check: if the sentence could appear unchanged in another project's docs, it says nothing about this one. Cut it.
28. **Shorten or split dense sentences.** If the reader has to backtrack to parse a sentence, break it in two or drop clauses. One idea per sentence.
29. **Active voice.** Prefer it. Catch "is/are/was/were + past participle" and name the actor: "queries are validated" becomes "the compiler validates queries", "the file is parsed by the loader" becomes "the loader parses the file". Passive is fine only when the actor is unknown or genuinely doesn't matter.
30. **Cut adverbs, or use a stronger verb.** "runs quickly" becomes "is fast" or the number. "significantly improves" becomes the measured delta. An adverb propping up a weak verb means the verb is wrong.
31. **Prefer the plain word.** `utilize` becomes "use", `leverage` becomes "use", `facilitate` becomes "help", "numerous" becomes "many", "in the event that" becomes "if". The fancier synonym is rarely clearer.
## Mention convention
A document that MENTIONS a banned word or phrase quotes it as inline code. The checker (`tools/unslop-hook/unslop-check.js`, machine source `tools/unslop-hook/lists.json`) strips code spans before matching, so a backticked mention is invisible to the gate while a bare one flags. This file follows that convention and doubles as a regression fixture: if `unslop-check.js` ever flags this file, either an edit broke the mention convention or code stripping regressed. Documents that deliberately CONTAIN slop to test detection (fixture files) are uses, not mentions; they are expected to flag.
+26
View File
@@ -33,5 +33,31 @@ for f in $FILES; do
printf '\n' >> "$TEMP"
done
# Agent identity (M13): when the launcher names the agent, the generated
# prompt states it - SOUL.md provides the persona, this provides the name.
if [ -n "${MOSAIC_AGENT_NAME:-}" ]; then
printf '===== AGENT IDENTITY =====\n' >> "$TEMP"
printf 'agent name: %s\n' "$MOSAIC_AGENT_NAME" >> "$TEMP"
printf '\n' >> "$TEMP"
fi
# Sanctioned mission injection point (M4): when the task runner provides a
# mission snapshot, its objective and directives are appended AFTER the
# immutable contracts. Runtime data; never part of the contract fixtures.
if [ -n "${MOSAIC_MISSION_FILE:-}" ]; then
if [ ! -r "$MOSAIC_MISSION_FILE" ]; then
echo "load-contracts: MOSAIC_MISSION_FILE set but not readable: $MOSAIC_MISSION_FILE" >&2
rm -f "$TEMP"
exit 1
fi
printf '===== MISSION (runtime) =====\n' >> "$TEMP"
node -e '
const m = JSON.parse(require("fs").readFileSync(process.env.MOSAIC_MISSION_FILE, "utf8"));
process.stdout.write("Objective: " + m.objective + "\n");
for (const d of m.directives ?? []) process.stdout.write("- " + d + "\n");
' >> "$TEMP"
printf '\n' >> "$TEMP"
fi
mv "$TEMP" "$OUT"
echo "load-contracts: wrote $OUT from $CONTRACT_DIR" >&2
+34 -31
View File
@@ -1,38 +1,41 @@
#!/bin/sh
# One-shot Pi agent runner inside the container.
# Loads the contract-generated system prompt, then sends exactly one
# user request through Pi's documented noninteractive mode and prints
# the model response on stdout.
# Agent dispatcher inside the container.
#
# Headless (default): loads the contract-generated system prompt, then
# dispatches one request to /opt/mosaic/adapters/<MOSAIC_ADAPTER>/adapter.sh
# (contract: /opt/mosaic/adapters/README.md).
#
# Interactive (MOSAIC_INTERACTIVE=1, from scripts/agent.sh): same prompt,
# but the adapter opens the full pi TUI with no initial prompt - the human
# drives from there.
set -eu
: "${PI_PROVIDER:=zai}"
: "${PI_MODEL:=glm-5.3-flash}"
export PI_PROVIDER PI_MODEL
REQUEST=""
if [ "${MOSAIC_INTERACTIVE:-}" != "1" ]; then
# Headless: args are the request; default is the startup verification
# request used by hello/verify.
REQUEST="${*:-Return your startup marker and nothing else.}"
export MOSAIC_REQUEST="$REQUEST"
fi
REQUEST="${*:-Return your startup marker and nothing else.}"
ADAPTER="${MOSAIC_ADAPTER:-pi}"
case "$ADAPTER" in
# Allowlist mirrors scripts/mosaic-config.mjs; pattern check first so a
# crafted name cannot escape the adapters directory.
*[!A-Za-z0-9._-]*|'')
echo "run-agent: invalid adapter name: '$ADAPTER'" >&2
exit 2
;;
esac
ADAPTER_SCRIPT="/opt/mosaic/adapters/$ADAPTER/adapter.sh"
if [ ! -x "$ADAPTER_SCRIPT" ]; then
echo "run-agent: unknown or non-executable adapter: $ADAPTER" >&2
exit 2
fi
/opt/mosaic/src/load-contracts.sh /opt/mosaic/contracts /var/lib/mosaic/system-prompt.md
# All flags are documented in the package README (CLI Reference):
# -p / --print noninteractive: print the response and exit
# --system-prompt replace the default system prompt with the
# contract-generated prompt
# --no-* switches prevent ambient context files, skills, extensions,
# prompt templates, and themes from being appended
# --no-session ephemeral: no persistent agent session
# --no-tools the startup request needs no tool execution
# --offline disable startup network operations (update checks,
# package update checks, install/update telemetry)
exec pi \
--offline \
--no-session \
--no-extensions \
--no-skills \
--no-prompt-templates \
--no-themes \
--no-context-files \
--no-tools \
--provider "$PI_PROVIDER" \
--model "$PI_MODEL" \
--system-prompt "$(cat /var/lib/mosaic/system-prompt.md)" \
-p "$REQUEST"
export MOSAIC_SYSTEM_PROMPT_FILE="/var/lib/mosaic/system-prompt.md"
exec "$ADAPTER_SCRIPT"
+8
View File
@@ -0,0 +1,8 @@
{
"taskVersion": 1,
"id": "t-session-teach",
"prompt": "Remember this code word for later: mosaico. Reply with exactly: REMEMBERED",
"session": "demo",
"expectExact": "REMEMBERED",
"timeoutSeconds": 180
}
+7
View File
@@ -0,0 +1,7 @@
{
"taskVersion": 1,
"id": "t-session-recall",
"prompt": "What code word did I ask you to remember earlier in this session? Reply with only the code word.",
"session": "demo",
"timeoutSeconds": 180
}
+9
View File
@@ -0,0 +1,9 @@
{
"taskVersion": 1,
"id": "t-worker-retry-refine",
"prompt": "Follow-up on your last change to scripts/mosaic-task.mjs (the retry operation). Conductor review found one defect: runTask builds spawnEnv from process.env, so a direct 'node scripts/mosaic-task.mjs retry <runId>' invocation lacks the launcher-exported variables and compose fails with: required variable MOSAIC_IMAGE_TAG is missing.\n\nFIX (in runTask, not retryRun): make spawnEnv self-sufficient —\n1. spawnEnv.PI_PROVIDER = resolved.execution.provider; spawnEnv.PI_MODEL = resolved.execution.model (from the already-loaded resolved config).\n2. spawnEnv.MOSAIC_IMAGE_TAG = 'mosaic-poc-agent:' + <pinned pi version> + '-r' + <RELEASE file content trimmed>, reading RELEASE and package.json the same way scripts/common.sh load_release does.\nKeep the change minimal; do not touch other files; keep code style.\n\nSELF-CHECK: node --check scripts/mosaic-task.mjs must pass.\nWhen finished, reply with exactly: REFINE_DONE",
"workspace": "stack-repo",
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
"session": "worker-1",
"timeoutSeconds": 600
}
+9
View File
@@ -0,0 +1,9 @@
{
"taskVersion": 1,
"id": "t-worker-retry",
"prompt": "You are working alone inside a git repository checkout (your current working directory). Implement a new feature in scripts/mosaic-task.mjs.\n\nFEATURE: add a 'retry <runId>' operation alongside the existing 'validate | run | show | list' operations.\n\nBEHAVIOR: 'retry <runId>' re-executes a previously recorded run. Steps: (1) validate the runId argument with the same pattern showRun uses; (2) load the config exactly like showRun does; (3) locate the run directory under the runs root; if it does not exist, fail with exit code 4 and the message 'run not found: <runId>'; (4) read that run directory's task.json snapshot; if unreadable, fail with exit code 4; (5) write the snapshot to a temp file and execute it through the EXISTING runTask(path) function so a brand-new run is created — do not duplicate any execution logic; (6) exit with whatever exit code the new run produced.\n\nCONSTRAINTS: reuse runTask; do not refactor unrelated code; do not touch any file other than scripts/mosaic-task.mjs; keep the existing code style; update the default usage error message to include retry.\n\nSELF-CHECK (run it yourself with the bash tool): node --check scripts/mosaic-task.mjs must pass.\n\nWhen finished, reply with exactly: RETRY_DONE",
"workspace": "stack-repo",
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
"session": "worker-1",
"timeoutSeconds": 600
}
+9
View File
@@ -0,0 +1,9 @@
{
"taskVersion": 1,
"id": "t-workspace-demo",
"prompt": "Use the bash tool to create a file named proof.txt in the current directory containing exactly the text: workspace works. Then reply with exactly: WORKSPACE_OK",
"workspace": "demo",
"capabilities": { "tools": ["bash", "read", "write"] },
"expectExact": "WORKSPACE_OK",
"timeoutSeconds": 180
}
+83
View File
@@ -0,0 +1,83 @@
# unslop-hook
Mechanical AI-tell enforcement for pi seats. Anti-drift gate for the writing
standard in SYSTEM.md / ms-unslop: prose distribution alone decays over long
sessions; this check cannot forget.
- `lists.json`: committed machine source for every list the checker enforces:
words, phrases, punct rules, regex patterns, density thresholds. Each entry
carries provenance (`ms-unslop:<pattern id>` or `system-md`), the mention
convention, and the documented divergence of the density gate from
SYSTEM.md's outright em-dash ban. Edit lists here, not in code.
- `unslop-check.js`: dependency-free checker (node CLI + module) driven by
lists.json. Loads and schema-validates the lists on first use and hard-fails
closed: empty, unparseable, or invalid lists throw. Detects banned vocabulary,
chatbot/sycophancy phrases, filler phrases, em/en dashes, curly quotes,
`not just X but Y`. Strips fenced and inline code first, so quoted code is
never flagged. Exit 0 clean, 1 violations, 2 gate broken (lists unreadable,
never a clean verdict).
- `extension.ts`: pi extension. `message_end` checks finalized assistant text
and notifies the operator (TUI/RPC). `before_agent_start` reads the most
recent assistant reply from the session file and, if it carries tells,
injects a correction notice the model sees on its next turn. `/unslop`
reports session stats. Violation state lives in the session file, so the
injection path survives restart, resume, fork, and reload (an in-memory
pending flag was measured dead across print-mode turns, 2026-08-19). A
broken lists.json fails closed: checks stop, `broken_lists` /
`skipped_broken` events log the reason, operator notified once, seat keeps
running.
- `test-unslop-check.js`: unit tests with red and green controls.
## Use
```bash
node test-unslop-check.js # suite
node unslop-check.js <file> # CLI check
UNSLOP_LISTS=<path> node unslop-check.js <file> # alt lists location
pi -e ~/.mosaic/tools/unslop-hook/extension.ts # ad-hoc load
# deploy: copy dir to ~/.pi/agent/extensions/unslop-hook/ or seat .pi
# equivalent, or list it in settings.json "extensions"
```
Env: `MOSAIC_UNSLOP_HOOK=0` disables. `MOSAIC_UNSLOP_LOG=<path>` appends JSONL
events (loaded / flagged / notice_injected / checked / broken_lists /
skipped_broken) for headless evidence. `UNSLOP_LISTS=<path>` overrides the
lists.json location for both CLI and extension.
## Verified here (2026-08-19)
- Unit suite 24/24 (8 behavioral, 11 loader/CLI, 5 review follow-up), red and
green controls both exercised, including exit-2 on broken lists and on an
unreadable input file (S3).
- CLI: slop file exit 1, clean file exit 0, broken lists exit 2 with the fault
named on stderr.
- Extension, healthy path (print mode, zai/glm-5.3:low): startup probe loads
lists.json, reply checked clean.
- Extension, broken-lists path (print mode): `broken_lists` at startup,
`skipped_broken` per turn, seat survives, reply still delivered.
- Earlier live evidence (pre-C1, inline lists): forced-slop turn flagged;
fresh-process follow-up injected the notice and the reply came back clean;
full TUI trial (notify line, injection, /unslop stats) on session vision-unslop.
- Log evidence in session scratchpad.
## Limits
- `/unslop` command not tested headless (print mode has no command surface);
it is a thin stats wrapper.
- En dash flag fires on typographic ranges too (23); acceptable for fleet
prose, revisit if it noisifies technical writing.
- Notice injection is a nudger, not a blocker. Output already streamed to the
user stays as-is; correction lands on the next turn.
- A broken lists.json latches for the session: repairing the file mid-session
does not revive checks until the seat restarts. Acceptable for an advisory
gate (review S1).
- The fail-closed operator notification requires a UI. Print-mode sessions
log `skipped_broken` but notify nobody (review S2).
- The word/phrase lists are the mechanical subset of ms-unslop only, keyed to
pattern ids in lists.json. Style judgments (voice, rhythm, structure) stay in
the skill, not the gate.
## Promotion path
Stack issue (A4): checker shared as the single source for a matching Claude
Code Stop-hook script; lists versioned beside SYSTEM.md contract text.
+163
View File
@@ -0,0 +1,163 @@
// unslop-hook — pi extension wrapper around unslop-check.js.
// Detects mechanical AI tells in finalized assistant messages and injects a
// correction notice the model sees on its next turn. Anti-drift enforcement for
// SYSTEM.md / ms-unslop; prose distribution alone decays, this cannot forget.
//
// Deploy: copy dir to ~/.pi/agent/extensions/unslop-hook/ (or seat .pi equivalent),
// or add this file's dir to settings.json "extensions".
// Test: pi -e <abs path>/extension.ts
// Off: MOSAIC_UNSLOP_HOOK=0
// Log: MOSAIC_UNSLOP_LOG=/path/to/log.jsonl (JSONL events; headless evidence)
// Broken: lists.json missing/empty/invalid → the checker throws; checks are
// skipped, logged as skipped_broken, and the operator is notified once.
// Never silently pass while the lists cannot load (fail closed).
//
// Design note: violation state lives in the SESSION FILE, not memory. At
// before_agent_start we read the most recent assistant text message from
// ctx.sessionManager and check it there. That survives process restarts, resume,
// fork, and reload — an in-memory pending flag measured dead on 2026-08-19 when
// a print-mode second turn never injected.
import { appendFileSync } from "node:fs";
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { checkText } from "./unslop-check.js";
interface Finding {
rule: string;
detail: string;
count: number;
}
interface MessageEntry {
type: "message";
id: string;
message: { role?: string; content?: unknown };
}
function assistantText(entry: unknown): string | null {
const e = entry as Partial<MessageEntry>;
if (e?.type !== "message") return null;
const msg = e.message;
if (msg?.role !== "assistant" || !Array.isArray(msg.content)) return null;
const text = msg.content
.filter((b): b is { type: "text"; text: string } =>
typeof b === "object" && b !== null && (b as { type?: string }).type === "text")
.map((b) => b.text ?? "")
.join("\n");
return text.trim() ? text : null; // tool-call-only assistant messages return null
}
export default function (pi: ExtensionAPI) {
if (process.env.MOSAIC_UNSLOP_HOOK === "0") return;
const LOG = process.env.MOSAIC_UNSLOP_LOG;
const log = (ev: Record<string, unknown>) => {
if (LOG) appendFileSync(LOG, JSON.stringify({ ts: Date.now(), ...ev }) + "\n");
};
// Entry ids we have already injected a notice for. In-memory only: after a
// restart the same entry may inject once more, which re-anchors the style
// after a context loss. That is wanted, not a bug.
const injectedFor = new Set<string>();
let turnsChecked = 0;
let turnsFlagged = 0;
const histogram = new Map<string, number>();
// Fail-closed path for a broken lists.json. A checker that cannot load its
// lists must never be read as "everything passed": checks stop, the skip is
// logged each turn, and the operator is notified once.
let broken: string | null = null;
let brokenNotified = false;
const reportBroken = (ctx: { hasUI?: boolean } | undefined, where: string) => {
log({ ev: "skipped_broken", where, reason: broken });
if (!brokenNotified && ctx?.hasUI) {
ctx.ui.notify(`unslop gate BROKEN: ${broken}. Fix tools/unslop-hook/lists.json; no clean verdicts until then.`, "error");
brokenNotified = true;
}
};
const safeCheck = (text: string): ReturnType<typeof checkText> | null => {
if (broken) return null;
try {
return checkText(text);
} catch (e) {
broken = String((e as Error).message);
log({ ev: "broken_lists", reason: broken });
return null;
}
};
pi.on("session_start", async (event, _ctx) => {
log({ ev: "loaded", reason: event.reason });
try {
checkText(""); // probe: load+validate lists at startup, not mid-conversation
} catch (e) {
broken = String((e as Error).message);
log({ ev: "broken_lists", reason: broken, at: "startup" });
}
});
pi.on("message_end", async (event, ctx) => {
if ((event.message as { role?: string }).role !== "assistant") return;
const text = assistantText({ type: "message", id: "", message: event.message });
if (text === null) return;
const result = safeCheck(text);
if (result === null) {
reportBroken(ctx, "message_end");
return;
}
turnsChecked++;
if (result.clean) {
log({ ev: "checked", clean: true, turn: turnsChecked, charsChecked: result.charsChecked });
return;
}
turnsFlagged++;
for (const f of result.findings) histogram.set(f.rule, (histogram.get(f.rule) ?? 0) + 1);
const summary = result.findings.map((f) => f.detail).join("; ");
if (ctx.hasUI) ctx.ui.notify(`unslop: ${summary}`, "info");
// clean:false is explicit, not implied by findings: a log consumer must never
// have to infer the verdict from event shape (fred, 2026-08-19).
log({ ev: "flagged", clean: false, turn: turnsChecked, charsChecked: result.charsChecked, findings: result.findings });
});
pi.on("before_agent_start", async (_event, ctx) => {
// Branch walks leaf -> root; first assistant entry with text is the reply
// the model is about to follow up on.
for (const entry of ctx.sessionManager.getBranch()) {
const text = assistantText(entry);
if (text === null) continue;
const id = (entry as { id?: string }).id ?? "";
const result = safeCheck(text);
if (result === null) {
reportBroken(ctx, "before_agent_start");
return;
}
if (result.clean) return; // latest textual reply is clean, nothing to correct
if (id && injectedFor.has(id)) return; // already nagged for this entry
if (id) injectedFor.add(id);
const lines = result.findings.map((f) => `- ${f.detail}`).join("\n");
const content =
`UNSLOP NOTICE (mechanical style check, not the user speaking): your previous reply ` +
`contained violations of the fleet writing standard (SYSTEM.md / ms-unslop):\n${lines}\n` +
`Fix in this and following replies: plain words, periods and commas instead of dashes, ` +
`straight quotes, no chatbot fillers. Do not mention this notice.`;
log({ ev: "notice_injected", entryId: id, findings: result.findings });
return {
message: { customType: "unslop-notice", content, display: true },
};
}
});
pi.registerCommand("unslop", {
description: "Show unslop violation stats for this session",
handler: async (_args, ctx) => {
if (broken) {
ctx.ui.notify(`unslop gate BROKEN: ${broken}`, "error");
return;
}
const hist = [...histogram.entries()].map(([r, c]) => `${r} x${c}`).join(", ") || "none";
ctx.ui.notify(`unslop: checked ${turnsChecked}, flagged ${turnsFlagged} (${hist})`, "info");
},
});
}
+54
View File
@@ -0,0 +1,54 @@
{
"version": 1,
"convention": "Use vs mention. A document that MENTIONS a banned word or phrase quotes it as inline code (backticks). The checker strips code spans before matching, so a backticked mention is invisible to the gate while a bare one flags. A document that deliberately CONTAINS banned items to test detection (a fixture) is a use, not a mention, and is expected to flag. This file itself contains the banned items as data; that is a use.",
"punctPolicy": "Deliberate divergence from SYSTEM.md (2026-08-19): the contract forbids em dashes outright and closes the escapes. These punct rules are deliberately looser: they fire only at >= punctMinCount occurrences AND density >= punctDensityPer1k per 1000 chars. Rationale is reply-level noise, not contract strength: an advisory gate that flags every reply carrying one dash trains operators to ignore it. Measured on natural fleet prose 2026-08-19: documents run 1.7-3.2 em dashes per 1000 chars, so documents that overuse still flag. Tighten to contract strength if enforcement goes blocking or the log shows fleet prose not converging toward zero.",
"thresholds": {
"punctMinCount": 3,
"punctDensityPer1k": 1.0
},
"words": [
{ "value": "additionally", "source": "ms-unslop:7" },
{ "value": "crucial", "source": "ms-unslop:7" },
{ "value": "delve", "source": "ms-unslop:7" },
{ "value": "garner", "source": "ms-unslop:7" },
{ "value": "interplay", "source": "ms-unslop:7" },
{ "value": "intricate", "source": "ms-unslop:7" },
{ "value": "pivotal", "source": "ms-unslop:7" },
{ "value": "showcase", "source": "ms-unslop:7" },
{ "value": "tapestry", "source": "ms-unslop:7" },
{ "value": "testament", "source": "ms-unslop:7" },
{ "value": "underscore", "source": "ms-unslop:7" },
{ "value": "vibrant", "source": "ms-unslop:7" },
{ "value": "utilize", "source": "ms-unslop:31" },
{ "value": "leverage", "source": "ms-unslop:31" },
{ "value": "facilitate", "source": "ms-unslop:31" },
{ "value": "load-bearing", "source": "system-md" }
],
"phrases": [
{ "value": "worth stating plainly", "source": "system-md" },
{ "value": "here's the honest truth", "source": "system-md" },
{ "value": "heres the honest truth", "source": "system-md", "note": "apostrophe-OMITTED renderings only; ASCII and curly-apostrophe forms match the main entry because the checker normalizes U+2019/U+2018 to ASCII before phrase matching" },
{ "value": "the real tension", "source": "system-md" },
{ "value": "carry the argument", "source": "system-md" },
{ "value": "in order to", "source": "ms-unslop:23" },
{ "value": "due to the fact that", "source": "ms-unslop:23" },
{ "value": "it is important to note", "source": "ms-unslop:23" },
{ "value": "i hope this helps", "source": "ms-unslop:20" },
{ "value": "let me know if", "source": "ms-unslop:20" },
{ "value": "of course!", "source": "ms-unslop:20" },
{ "value": "certainly!", "source": "ms-unslop:20" },
{ "value": "found the smoking gun", "source": "ms-unslop:20" },
{ "value": "happy to help", "source": "ms-unslop:20", "note": "extension of the named pattern set" },
{ "value": "great question", "source": "ms-unslop:22" },
{ "value": "absolutely right", "source": "ms-unslop:22" },
{ "value": "excellent question", "source": "ms-unslop:22", "note": "extension of the named pattern set" }
],
"punct": [
{ "value": "em", "label": "em dash", "chars": ["\u2014"], "source": "ms-unslop:13+system-md" },
{ "value": "en", "label": "en dash", "chars": ["\u2013"], "source": "ms-unslop:13" },
{ "value": "curly", "label": "curly quote/apostrophe", "chars": ["\u201c", "\u201d", "\u2018", "\u2019"], "source": "ms-unslop:19" }
],
"patterns": [
{ "value": "not-just-but", "regex": "not just\\s+[^.!?]{0,80}?\\s+but", "flags": "gi", "detail": "not just X but Y", "source": "ms-unslop:9" }
]
}
+228
View File
@@ -0,0 +1,228 @@
"use strict";
// Tests for unslop-check.js. Run: node test-unslop-check.js
// Exit 0 = all pass. Cases include a red control (slop must fail) and a green
// control (clean prose must pass) per evidence discipline.
const assert = require("node:assert");
const fs = require("node:fs");
const os = require("node:os");
const path = require("node:path");
const { spawnSync } = require("node:child_process");
const { checkText, stripCode, loadLists } = require("./unslop-check.js");
const SLOP = `Certainly! Let me delve into the evolving tapestry of database technology — its truly “pivotal” — and intricate.
In order to understand it — we should leverage this interplay of systems — deeply. I hope this helps!`;
const CLEAN = `The loader parses the file and validates each row. Rows that fail are logged
and skipped. We measured a range from 1 to 10 seconds. Use "straight quotes" and
commas, not dashes. That is the whole finding.`;
// Code-stripping control: banned words inside code must not count.
const WITH_CODE = [
"The config uses `utilize=false` internally.",
"```",
"delve tapestry — pivotal",
"```",
"The config file sets one flag. It is parsed at startup.",
].join("\n");
const results = [];
function t(name, fn) {
try { fn(); results.push([name, true]); } catch (e) { results.push([name, false]); console.error(`FAIL ${name}: ${e.message}`); }
}
t("slop fixture is flagged (red control)", () => {
const r = checkText(SLOP);
assert.ok(!r.clean, "slop must not be clean");
const details = r.findings.map((f) => f.detail).join("; ");
assert.ok(r.findings.some((f) => f.detail.includes("delve")), `delve missing: ${details}`);
assert.ok(r.findings.some((f) => f.detail.includes("tapestry")), `tapestry missing: ${details}`);
assert.ok(r.findings.some((f) => f.detail.includes("pivotal")), `pivotal missing: ${details}`);
assert.ok(r.findings.some((f) => f.detail.includes("em dash")), `em dash missing: ${details}`);
assert.ok(r.findings.some((f) => f.detail.includes("curly")), `curly missing: ${details}`);
assert.ok(r.findings.some((f) => f.detail.includes("in order to")), `in order to missing: ${details}`);
assert.ok(r.findings.some((f) => f.detail.includes("i hope this helps")), `chatbot phrase missing: ${details}`);
assert.ok(r.findings.some((f) => f.detail.includes("certainly")), `certainly missing: ${details}`);
});
t("clean fixture passes (green control)", () => {
const r = checkText(CLEAN);
assert.deepStrictEqual(r.findings, [], `unexpected findings: ${JSON.stringify(r.findings)}`);
});
t("numeric range is not a false range flag", () => {
const r = checkText(CLEAN);
assert.ok(!r.findings.some((f) => f.rule === "pattern"), "must not flag numeric ranges");
});
t("code blocks and inline code are stripped", () => {
const r = checkText(WITH_CODE);
assert.deepStrictEqual(r.findings, [], `code leaked into check: ${JSON.stringify(r.findings)}`);
});
t("stripCode removes fenced and inline code", () => {
const s = stripCode("a `x — y` b\n```\ndelve\n```\nc");
assert.ok(!s.includes("delve"), "fenced code not stripped");
assert.ok(!s.includes("—"), "inline code not stripped");
assert.ok(s.includes("a") && s.includes("b") && s.includes("c"), "prose lost");
});
t("not-just-but pattern is detected", () => {
const r = checkText("This is not just a cache but a coordination layer.");
assert.ok(r.findings.some((f) => f.rule === "pattern"), "pattern missed");
});
t("light dash use is not flagged (below threshold)", () => {
const prose =
"The loader parses each row and validates it against the schema. Rows that fail " +
"are logged — with their line numbers — and skipped. The operator reviews the log " +
"daily and reconciles the rejects against the source system by hand, which takes " +
"a few minutes and has never once produced a discrepancy worth acting on.";
const r = checkText(prose);
assert.ok(!r.findings.some((f) => f.detail.includes("dash")), "2 dashes in ~330 chars must not flag");
});
t("dash overuse is flagged (above threshold)", () => {
const r = checkText("One — two — three — four. That is the whole sentence.");
assert.ok(r.findings.some((f) => f.detail.includes("em dash")), "4 dashes in 50 chars must flag");
});
// ── C1: lists.json machine source ──────────────────────────────────────
// The committed lists are the single source of truth; these tests pin the
// file's validity, its shape, and the loader's fail-closed behavior.
const tmpdir = fs.mkdtempSync(path.join(os.tmpdir(), "unslop-c1-"));
const tmp = (n) => path.join(tmpdir, n);
function brokenVariant(mutate) {
const l = JSON.parse(JSON.stringify(loadLists()));
mutate(l);
return l;
}
function writeTmp(name, data) {
const f = tmp(name);
fs.writeFileSync(f, typeof data === "string" ? data : JSON.stringify(data));
return f;
}
t("lists.json (the real file) validates and is pinned in size", () => {
const l = loadLists();
assert.strictEqual(l.version, 1);
// Counts pin the migration: 16 words, 17 phrases, 3 punct, 1 pattern moved
// from the old inline constants. Changing a count means changing this test
// too, consciously.
assert.strictEqual(l.words.length, 16, "word count drifted");
assert.strictEqual(l.phrases.length, 17, "phrase count drifted");
assert.strictEqual(l.punct.length, 3, "punct count drifted");
assert.strictEqual(l.patterns.length, 1, "pattern count drifted");
assert.ok(l.convention.length > 50, "mention convention must be present");
assert.ok(l.punctPolicy.length > 50, "punct divergence policy must be present");
for (const e of [...l.words, ...l.phrases, ...l.punct, ...l.patterns]) {
assert.ok(e.source && e.source.trim(), `entry missing source: ${JSON.stringify(e)}`);
}
});
t("loader rejects an empty file", () => {
const f = writeTmp("empty.json", "");
assert.throws(() => loadLists(f), /empty file/);
});
t("loader rejects unparseable JSON", () => {
const f = writeTmp("bad.json", "{nope");
assert.throws(() => loadLists(f), /unparseable/);
});
t("loader rejects a missing file", () => {
assert.throws(() => loadLists(tmp("does-not-exist.json")), /cannot read/);
});
t("loader rejects missing keys", () => {
const f = writeTmp("nokeys.json", { version: 1 });
assert.throws(() => loadLists(f), /missing key/);
});
t("loader rejects an emptied word list", () => {
const f = writeTmp("emptywords.json", brokenVariant((l) => { l.words = []; }));
assert.throws(() => loadLists(f), /words must be a non-empty array/);
});
t("loader rejects entries without provenance", () => {
const f = writeTmp("nosource.json", brokenVariant((l) => { delete l.phrases[0].source; }));
assert.throws(() => loadLists(f), /source/);
});
t("loader rejects duplicate values", () => {
const f = writeTmp("dup.json", brokenVariant((l) => { l.words.push({ ...l.words[0] }); }));
assert.throws(() => loadLists(f), /duplicate/);
});
t("loader rejects a non-compiling pattern regex", () => {
const f = writeTmp("badregex.json", brokenVariant((l) => { l.patterns[0].regex = "("; }));
assert.throws(() => loadLists(f), /does not compile/);
});
t("CLI exits 2 on broken lists (red control)", () => {
const f = writeTmp("cli-broken.json", "");
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js")], {
input: "some prose",
encoding: "utf8",
env: { ...process.env, UNSLOP_LISTS: f },
});
assert.strictEqual(r.status, 2, `expected exit 2, got ${r.status} (stderr: ${r.stderr})`);
assert.ok(r.stderr.includes("lists.json invalid"), `stderr must name the fault: ${r.stderr}`);
});
t("CLI honors UNSLOP_LISTS for a valid file (green control)", () => {
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js")], {
input: "plain prose with no tells at all",
encoding: "utf8",
env: { ...process.env, UNSLOP_LISTS: path.join(__dirname, "lists.json") },
});
assert.strictEqual(r.status, 0, `expected exit 0, got ${r.status} (stderr: ${r.stderr})`);
});
// ── Review follow-up (rev-code-01, 2026-08-19): F1, F2, F3, S3 ────────
t("loader rejects non-finite thresholds (F1)", () => {
// Infinity cannot round-trip JSON.stringify, so the fixture is a raw string
// edit of the real file — exactly the hand-edit that produced the finding.
const real = fs.readFileSync(path.join(__dirname, "lists.json"), "utf8");
const f = writeTmp("inf-threshold.json", real.replace('"punctMinCount": 3', '"punctMinCount": 1e999'));
assert.ok(real !== fs.readFileSync(f, "utf8") || !real.includes('"punctMinCount": 3'), "fixture mutation did not apply; test is vacuous");
assert.throws(() => loadLists(f), /finite/);
const f2 = writeTmp("inf-density.json", real.replace('"punctDensityPer1k": 1.0', '"punctDensityPer1k": 1e999'));
assert.throws(() => loadLists(f2), /finite/);
});
t("curly-apostrophe phrase rendering is flagged (F2 red control)", () => {
const r = checkText("Here\u2019s the honest truth about the deploy.");
assert.ok(!r.clean, "curly apostrophe must not defeat phrase matching");
assert.ok(r.findings.some((x) => x.detail.includes("here's the honest truth")), `main entry must match, got: ${JSON.stringify(r.findings)}`);
});
t("curly apostrophes still fire the punct rule alongside phrases (F2 ordering)", () => {
// Normalization for phrases must not eat the punct signal: four curly
// quotes in short text must flag punct, not only the phrase.
const r = checkText("It\u2019s \u2019one\u2019 \u2019two\u2019 \u2019three\u2019 \u2019four\u2019 done.");
assert.ok(r.findings.some((f) => f.rule === "punct"), `punct must fire on original text: ${JSON.stringify(r.findings)}`);
});
t("loader rejects non-lowercase phrase values (F3)", () => {
const f = writeTmp("cap-phrase.json", brokenVariant((l) => { l.phrases[0].value = "Worth Stating Plainly"; }));
assert.throws(() => loadLists(f), /lowercase/);
});
t("CLI exits 2 on unreadable input file (S3)", () => {
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js"), tmp("definitely-absent.txt")], {
encoding: "utf8",
});
assert.strictEqual(r.status, 2, `expected exit 2, got ${r.status} (stderr: ${r.stderr})`);
assert.ok(r.stderr.includes("cannot read input"), `stderr must name the fault: ${r.stderr}`);
});
let failed = 0;
for (const [name, ok] of results) { console.log(`${ok ? "PASS" : "FAIL"} ${name}`); if (!ok) failed++; }
console.log(`${results.length - failed}/${results.length} passed`);
try { fs.rmSync(tmpdir, { recursive: true, force: true }); } catch {}
process.exit(failed ? 1 : 0);
+181
View File
@@ -0,0 +1,181 @@
#!/usr/bin/env node
"use strict";
// unslop-check — mechanical AI-tell checker (ms-unslop subset + SYSTEM.md phrase bans).
// Plain JS, no deps, so pi extensions (jiti) and Claude Code hook scripts (node CLI)
// share one implementation.
//
// The lists live in lists.json beside this file: committed machine source with
// per-entry provenance (which ms-unslop pattern or SYSTEM.md rule each entry
// mechanizes), the mention convention, and the punct thresholds. The loader
// hard-fails closed: an empty, unparseable, or schema-invalid lists.json throws,
// and the CLI exits 2 so a broken gate is never mistaken for a clean verdict.
//
// CLI: node unslop-check.js <file> (or stdin)
// exit 0 = clean, exit 1 = violations found (findings printed as JSON),
// exit 2 = gate broken (lists.json missing/empty/invalid; error on stderr).
// Env: UNSLOP_LISTS=<path> overrides the lists.json location (testing; reuse by
// other harnesses sharing this file).
//
// Provenance note (2026-08-19): the inline lists this file carried before C1
// moved to lists.json unchanged — 16 words, 17 phrases, 3 punct rules, 1 pattern.
// The suite pins those counts; a list edit without a test edit is a drift signal.
const fs = require("node:fs");
const path = require("node:path");
function stripCode(text) {
// Fenced blocks (``` or ~~~), then inline code spans. Code is quoted material,
// not the agent's prose style. This is also the mention convention: a banned
// item quoted as inline code is a mention and must not flag (see lists.json).
return text
.replace(/```[\s\S]*?```/g, " ")
.replace(/~~~[\s\S]*?~~~/g, " ")
.replace(/`[^`\n]*`/g, " ");
}
// ── lists.json loading and validation ─────────────────────────────────────────
function validateLists(data) {
const fail = (why) => { throw new Error(`lists.json invalid: ${why}`); };
if (typeof data !== "object" || data === null || Array.isArray(data)) fail("top level must be an object");
for (const k of ["version", "convention", "punctPolicy", "thresholds", "words", "phrases", "punct", "patterns"]) {
if (!(k in data)) fail(`missing key: ${k}`);
}
if (typeof data.version !== "number" || data.version < 1) fail("version must be a number >= 1");
for (const k of ["convention", "punctPolicy"]) {
if (typeof data[k] !== "string" || !data[k].trim()) fail(`${k} must be a non-empty string`);
}
const th = data.thresholds;
if (typeof th !== "object" || th === null) fail("thresholds must be an object");
// Number.isFinite, not just typeof: JSON.parse of 1e999 yields Infinity, which
// passes typeof-number and would silently disable the punct gate (review F1).
if (!Number.isFinite(th.punctMinCount) || th.punctMinCount < 1) fail("thresholds.punctMinCount must be a finite number >= 1");
if (!Number.isFinite(th.punctDensityPer1k) || !(th.punctDensityPer1k > 0)) fail("thresholds.punctDensityPer1k must be a finite number > 0");
const seen = new Set();
const checkEntries = (arr, kind, extra) => {
if (!Array.isArray(arr) || arr.length === 0) fail(`${kind} must be a non-empty array`);
arr.forEach((e, i) => {
const at = `${kind}[${i}]`;
if (typeof e !== "object" || e === null) fail(`${at} must be an object`);
if (typeof e.value !== "string" || !e.value.trim()) fail(`${at}.value must be a non-empty string`);
if (typeof e.source !== "string" || !e.source.trim()) fail(`${at}.source must be a non-empty string (pattern id or system-md)`);
if (extra) extra(e, at, fail);
if (seen.has(`${kind}:${e.value}`)) fail(`duplicate ${kind} value: ${e.value}`);
seen.add(`${kind}:${e.value}`);
});
};
checkEntries(data.words, "words");
checkEntries(data.phrases, "phrases", (e, at, fail) => {
// Phrase matching splits a lowercased haystack, so an uppercase letter in a
// phrase value is a silently dead rule (review F3). Reject, do not silently
// normalize: list edits should fail loud (D-a).
if (e.value !== e.value.toLowerCase()) fail(`${at}.value must be lowercase; phrase matching lowercases the haystack: ${e.value}`);
});
checkEntries(data.punct, "punct", (e, at, fail) => {
if (typeof e.label !== "string" || !e.label.trim()) fail(`${at}.label must be a non-empty string`);
if (!Array.isArray(e.chars) || e.chars.length === 0 || !e.chars.every((c) => typeof c === "string" && c.length === 1)) {
fail(`${at}.chars must be a non-empty array of single-char strings`);
}
});
checkEntries(data.patterns, "patterns", (e, at, fail) => {
if (typeof e.regex !== "string" || !e.regex.trim()) fail(`${at}.regex must be a non-empty string`);
if (typeof e.flags !== "string") fail(`${at}.flags must be a string`);
if (typeof e.detail !== "string" || !e.detail.trim()) fail(`${at}.detail must be a non-empty string`);
try { new RegExp(e.regex, e.flags); } catch (err) { fail(`${at}.regex does not compile: ${err.message}`); }
});
return data;
}
let cache = null;
function loadLists(filePath) {
if (cache && !filePath) return cache;
const p = filePath || process.env.UNSLOP_LISTS || path.join(__dirname, "lists.json");
let raw;
try {
raw = fs.readFileSync(p, "utf8");
} catch (e) {
throw new Error(`lists.json invalid: cannot read ${p}: ${e.message}`);
}
if (!raw.trim()) throw new Error(`lists.json invalid: empty file: ${p}`);
let data;
try {
data = JSON.parse(raw);
} catch (e) {
throw new Error(`lists.json invalid: unparseable JSON: ${e.message}`);
}
const validated = validateLists(data);
if (!filePath) cache = validated;
return validated;
}
// ── checker ───────────────────────────────────────────────────────────────────
const escapeRegex = (s) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
function checkText(raw) {
const lists = loadLists();
const text = stripCode(String(raw));
// Phrase haystack: lowercased, then curly apostrophes normalized to ASCII.
// This must be a SEPARATE string from `text`: punct counting reads the
// original, so curly quotes still fire the punct rule (review F2 ordering).
const phraseHay = text.toLowerCase().replace(/[\u2018\u2019]/g, "'");
const findings = [];
for (const w of lists.words) {
const re = new RegExp("\\b" + escapeRegex(w.value) + "\\b", "gi");
const count = (text.match(re) || []).length;
if (count > 0) findings.push({ rule: "word", detail: `banned word "${w.value}" x${count}`, count });
}
for (const p of lists.phrases) {
const count = phraseHay.split(p.value).length - 1;
if (count > 0) findings.push({ rule: "phrase", detail: `phrase "${p.value}" x${count}`, count });
}
// Density-gated punctuation. The deliberate divergence from SYSTEM.md's
// outright em-dash ban is documented in lists.json punctPolicy, not only here.
for (const pc of lists.punct) {
let count = 0;
for (const ch of pc.chars) count += text.split(ch).length - 1;
if (count < lists.thresholds.punctMinCount) continue;
if (count / Math.max(text.length, 1) * 1000 < lists.thresholds.punctDensityPer1k) continue;
findings.push({ rule: "punct", detail: `${pc.label} x${count} (density-gated)`, count });
}
for (const pt of lists.patterns) {
const flags = pt.flags.includes("g") ? pt.flags : pt.flags + "g";
const m = text.match(new RegExp(pt.regex, flags));
const count = m ? m.length : 0;
if (count > 0) findings.push({ rule: "pattern", detail: `"${pt.detail}" x${count}`, count });
}
return { clean: findings.length === 0, findings, charsChecked: text.length };
}
module.exports = { checkText, stripCode, loadLists, validateLists };
if (require.main === module) {
let input;
try {
input = process.argv[2] ? fs.readFileSync(process.argv[2], "utf8") : fs.readFileSync(0, "utf8");
} catch (e) {
// An unreadable input must not exit 1: that is the violations code, and a
// wrapper keying on rc alone would report slop-free for a file it never
// read (review S3).
console.error(`unslop-check: cannot read input: ${e.message}`);
process.exit(2);
}
let result;
try {
result = checkText(input);
} catch (e) {
if (String(e.message).startsWith("lists.json invalid")) {
console.error(`unslop-check: ${e.message}`);
process.exit(2);
}
throw e;
}
console.log(JSON.stringify(result, null, 2));
process.exit(result.clean ? 0 : 1);
}