Compare commits
42
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
87f10772ce | ||
|
|
7db4c5c2ed | ||
|
|
0273a84549 | ||
|
|
3b674b7a66 | ||
|
|
527bc581ca | ||
|
|
1249714a9a | ||
|
|
e175616885 | ||
|
|
6955717612 | ||
|
|
88d9cf750f | ||
|
|
c038706eed | ||
|
|
8622c9d826 | ||
|
|
88eef507b0 | ||
|
|
439bea6915 | ||
|
|
fad8a4718c | ||
|
|
2ff49adff4 | ||
|
|
44c476ebbf | ||
|
|
cde480eb60 | ||
|
|
5808248707 | ||
|
|
afd5827db8 | ||
|
|
d9cc990376 | ||
|
|
83c4e9851e | ||
|
|
22508170a2 | ||
|
|
90a67d050e | ||
|
|
24bdef75fa | ||
|
|
4e2a413640 | ||
|
|
be55549700 | ||
|
|
ddb1554e5b | ||
|
|
b017e66e17 | ||
|
|
172368612c | ||
|
|
1387231e57 | ||
|
|
0292392e64 | ||
|
|
594b8d711c | ||
|
|
3c1ffd2c2d | ||
|
|
4ebb123ba3 | ||
|
|
bb5cecb348 | ||
|
|
88f55d9135 | ||
|
|
e7e1bd26eb | ||
|
|
5d86b8fa93 | ||
|
|
35ea464661 | ||
|
|
e87ecdb3e5 | ||
|
|
a947db7bfd | ||
|
|
a35ea62ab1 |
@@ -0,0 +1,114 @@
|
||||
# AGENTS.md — Mosaic Stack rebuild (`mosaicstack/stack-v2`)
|
||||
|
||||
Operational context for any agent session working in this repository.
|
||||
Read top to bottom; it is deliberately short — depth lives in the files it
|
||||
points to, not here.
|
||||
|
||||
## What this repository is
|
||||
|
||||
A standalone rebuild of Mosaic Stack: a file-based, fail-closed
|
||||
orchestration foundation that dispatches sandboxed headless pi workers to do
|
||||
real work, with immutable run records as evidence. Thirteen-plus tagged
|
||||
milestones (`git tag -l`) from `poc-container-hello-v0` to today; suites
|
||||
green at every step. Not production software — a proven foundation.
|
||||
|
||||
## Non-negotiable invariants (the canon)
|
||||
|
||||
1. **Root is bootstrap-only.** First-class system configuration lives at the
|
||||
repository root; everything else gets a dedicated directory (`roles/`,
|
||||
`contracts/`, `missions/`, `tasks/`, `docs/`). Do not add new files to root.
|
||||
2. **Configuration**: `~/.config/mosaic-dev/config.json` is the sole system
|
||||
config — created only by `scripts/bootstrap.sh`, never overwritten,
|
||||
fail-closed on any problem. Repo-scoped role authority lives in
|
||||
`roles/*.json` (versioned, reviewed commits only).
|
||||
3. **Secrets** never enter the repository or container images; auth is
|
||||
runtime-only (read-only mount or environment variable).
|
||||
4. **Contracts** (`contracts/`) are immutable and image-baked. Missions and
|
||||
tasks are declarative JSON with strict schemas.
|
||||
5. **Run records** under `<dataRoot>/runs/` are write-once evidence — never
|
||||
rewritten, only pruned via `prune` with a receipt.
|
||||
6. **Fail closed**: missing or invalid config/policy refuses the operation.
|
||||
Never improvise around a refusal; diagnose it.
|
||||
7. **Policy**: missions govern tasks (least-privilege intersection — a task
|
||||
narrows, never widens). Role authority is declared in `roles/` and changes
|
||||
only via reviewed commits.
|
||||
8. **Git**: commit only after suites are green; push only `main`; never
|
||||
force-push. `scripts/conductor-apply.sh` commits locally — push stays an
|
||||
explicit act.
|
||||
9. **Append-only logs**: BUILD-LOG.md (phases), `activation-log.jsonl`,
|
||||
`.pruned.log`, docs/SESSIONS.md. Corrections are new entries, never edits.
|
||||
|
||||
## Session protocol (mandatory)
|
||||
|
||||
- **Register** your session in `docs/SESSIONS.md` — one append-only line
|
||||
(date, actor, scope, outcome). Never rewrite or remove entries.
|
||||
- **Cadence**: read `docs/plans/CURRENT.md` → execute its single next action
|
||||
fully (implement → test → verify against acceptance criteria → commit →
|
||||
push → close issue) → update CURRENT.md → register in SESSIONS.md.
|
||||
- "next" means one action. A batch mandate ("run the queue") repeats the
|
||||
loop until green or blocked. Blocked means stop and report, never improvise.
|
||||
- Substantial work gets a Gitea issue and a BUILD-LOG phase entry
|
||||
(before/after, with corrections recorded honestly).
|
||||
|
||||
## Role model
|
||||
|
||||
- **Conductor**: a system-scoped role — not an agent, not a daemon. Holds
|
||||
git/credentials/policy authority; decomposes, dispatches, reviews,
|
||||
verifies, integrates. Protocol: `docs/plans/CONDUCTOR.md`. Exists only
|
||||
when invoked; push is never automatic.
|
||||
- **Workers**: headless pi via `scripts/run-task.sh` — sandboxed workspace,
|
||||
tools allowlist, optional persistent sessions and forks; no git, no
|
||||
credentials, no policy control.
|
||||
- Worker runs deliberately exclude this file (`--no-context-files` in the
|
||||
adapter): worker context is contracts + mission via the generated system
|
||||
prompt. This file is for conductor-level sessions.
|
||||
|
||||
## Command surface
|
||||
|
||||
`scripts/bootstrap.sh` (idempotent) · `build.sh` · `hello.sh` ·
|
||||
`verify.sh` · `run-task.sh run <task.json>` · `release.sh
|
||||
package|activate|rollback|status` · `reset.sh` (**danger**: wipes the data
|
||||
root; triple-safety-checked) · `mosaic-task.mjs validate|run|show|list|retry|prune` ·
|
||||
`agent.sh <name>` (interactive TUI agent) ·
|
||||
suites: `test-config.sh`, `test-task.sh`, `test-release.sh`,
|
||||
`test-conductor.sh`.
|
||||
|
||||
Full reference — usage, fields, exit codes, safety notes:
|
||||
`docs/TOOLS.md` (read on demand; do not rely on this summary for detail).
|
||||
|
||||
## Data map (canon)
|
||||
|
||||
- `~/.config/mosaic-dev/config.json` — system config (user-authored; never
|
||||
auto-written).
|
||||
- `<dataRoot>` (from config; default `~/.mosaic-dev`):
|
||||
- `runs/` — write-once run evidence (`result.json`, snapshots, `stderr.txt`)
|
||||
- `sessions/` — pi JSONL session trees, one directory per named session
|
||||
- `workspaces/` — agent file effects (persistent or `:run` ephemeral)
|
||||
- `state/` — release pointer + append-only activation/auto-apply logs
|
||||
- Ownership is per-directory; nothing shares state. Directory map and
|
||||
lifecycle rules: README.md "Data map" section.
|
||||
|
||||
## Pointers (depth lives here)
|
||||
|
||||
- `docs/plans/CURRENT.md` — THE next action (single source of "what now")
|
||||
- `docs/plans/CONDUCTOR.md` — orchestration protocol and guardrails
|
||||
- `docs/plans/2026-09-02_atomic-mosaic-foundation.md` — architecture, invariants
|
||||
- `docs/plans/2026-09-03_autonomous-run.md` — batch-run tracker
|
||||
- `BUILD-LOG.md` — append-only build/verification history with corrections
|
||||
- `LAYERS.md` — implemented vs deferred layers
|
||||
- `docs/SESSIONS.md` — session registry
|
||||
- `adapters/README.md` — the harness adapter contract
|
||||
- `roles/` — role contracts (conductor, future agent/coder/reviewer)
|
||||
|
||||
## Recovery rule
|
||||
|
||||
Compacted, restarted, or new? Nothing that matters is lost: this file +
|
||||
`docs/plans/CURRENT.md` + `git log --oneline -10` + the suites reconstruct
|
||||
the full state. **Never guess** — verify with the suites; the run records
|
||||
and logs hold the receipts.
|
||||
|
||||
## Version pin
|
||||
|
||||
`@earendil-works/pi-coding-agent` is pinned exactly (see `package.json` /
|
||||
`RELEASE`); never install unversioned. Release identity: `RELEASE` file
|
||||
(0.0.X until declared stable); image tags derive from it.
|
||||
+174
@@ -163,4 +163,178 @@ Configuration-driven Hello World verified. `main` merged with M1 and tagged `con
|
||||
|
||||
Mission/task layer verified end-to-end. `main` merged with M2 and tagged `mission-task-v1`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7: Release model and safe updates (M3)
|
||||
|
||||
### Entry 7.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Add the release substrate (Gitea milestone M3, issues #10-#13): RELEASE file single-sources the version (0.0.X line per owner direction), image tags derive from it, scripts/release.sh provides package/activate/rollback/status, activation is health-gated by the M2 task runner, pointer + append-only log under <dataRoot>/state/.
|
||||
- Reason: The owner's top invariant — updates must never corrupt a working installation — needs a mechanism, not a convention: gate-then-flip with recorded history and rollback.
|
||||
- Expected result: Update, refusal, and rollback drills all green with config checksums unchanged.
|
||||
|
||||
### Entry 7.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: scripts/test-release.sh (14 cases); recorded drills: update (0.0.3 -> 0.0.4 package+activate+verify), fault-injected refusal, rollback to 0.0.3.
|
||||
- Observed result:
|
||||
- Selftests: 14 passed, 0 failed.
|
||||
- Update drill: packaged and activated r0.0.4 after exact-marker health gate; verify green under the new tag; config checksum unchanged.
|
||||
- Refusal drill: health-gate fault injection -> activation refused (exit 1), pointer untouched, refusal appended to the log.
|
||||
- Rollback drill: health-gated rollback to r0.0.3; pointer restored; log records package/activate/refused/rollback history append-only.
|
||||
- Failure or correction:
|
||||
1. release.sh initially failed with missing state/ directory (no mkdir before pointer/log writes); fixed.
|
||||
2. Selftest harness mutated the repo RELEASE and restored the mutated copy (mv-back bug) plus a second trap replacing the first; fixed with inline backup restore and one self-healing exit trap. Product code unaffected.
|
||||
- Credential check: no credential material in release state, logs, or drills.
|
||||
|
||||
## Result (M3)
|
||||
|
||||
Release model and safe updates verified by drills. `main` merged with M3 and tagged `release-model-v1`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 8: Runtime adapter seam (M4)
|
||||
|
||||
### Entry 8.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Formalize the harness boundary (Gitea milestone M4, issues #16-#19): documented adapter contract under /opt/mosaic/adapters/<name>/adapter.sh; run-agent.sh becomes a dispatcher; pi extracted unchanged; deterministic mock adapter for provider-free seam tests; config gains optional execution.adapter (default pi, configVersion unchanged); mission directives gain their sanctioned injection point via the run snapshot; RELEASE bumps to 0.0.5 with a health-gated activation.
|
||||
- Reason: Future harnesses (Claude, Codex, OpenCode) must be additive — one directory each — and mission content needs a single sanctioned path into the runtime.
|
||||
- Expected result: All suites green including new deterministic seam cases; 0.0.5 activated by health gate; mission-bearing run recorded.
|
||||
|
||||
### Entry 8.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: scripts/test-config.sh; scripts/test-task.sh; scripts/test-release.sh; manual seam drills (mock verbatim, unknown/traversal adapter refusal); mission injection checks; release package + activate for 0.0.5.
|
||||
- Observed result:
|
||||
- Config suite 24/24 (adapter default/validation/env export).
|
||||
- Task suite 24/24 including deterministic mock cases (gate pass, expect-mismatch with reason, unknown adapter fail-closed) and mission injection asserted by prompt content.
|
||||
- Release suite 14/14; image mosaic-poc-agent:0.84.4-r0.0.5 packaged and activated via exact-marker health gate.
|
||||
- Mission directives now flow: task -> run snapshot -> container env -> generated prompt MISSION (runtime) section.
|
||||
- Failure or correction:
|
||||
1. Selection authority settled: load_config always exports MOSAIC_ADAPTER from config; environment overrides for scripts are therefore not a supported selection path (by design).
|
||||
2. Selftest harness: three authoring defects fixed (helpers used before definition; one config file reused across cases leaking adapter state; a static mission fixture asserted against distinctive seam directives; plus an accidentally duplicated live block removed).
|
||||
- Credential check: no credential material in adapters, prompts, run records, or logs.
|
||||
|
||||
## Result (M4)
|
||||
|
||||
Adapter seam verified; harness boundary is now additive by construction. `main` merged with M4 and tagged `adapter-seam-v1`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 9: Workspaces + capability envelope (M5)
|
||||
|
||||
### Entry 9.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Optional task workspace (absent / ":run" ephemeral / named persistent) and capabilities.tools allowlist (pi built-ins); runner plumbing via MOSAIC_WORKSPACE/MOSAIC_TOOLS; pi adapter maps to cwd + --tools; mock adapter logs delivered MOSAIC_* vars for deterministic assertions (Gitea #20, #21).
|
||||
- Reason: Agents that only answer text cannot do work; the workspace+tools pair is the smallest real capability step, bounded by the container.
|
||||
- Expected result: plumbing asserted via run-record stderr; live pi writes a host-visible file.
|
||||
|
||||
### Entry 9.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: build; mock plumbing run; validate negatives; full task suite; live workspace demo.
|
||||
- Observed result: task suite 32/32; MOSAIC_WORKSPACE/MOSAIC_TOOLS asserted in run record; dataRoot/workspaces/<name> created host-side; live pi used bash to write proof.txt into the demo workspace (host-visible).
|
||||
- Failure or correction: (1) batch edit dropped SUPPORTED_TOOLS const (runtime ReferenceError, exit 1 instead of 2) — restored; (2) mock env dump used `export` which this dash prints as `export K='v'` — switched to `env`; (3) selftest fed a mismatching mock response to the plain-task case — test bug, fixed.
|
||||
|
||||
## Phase 10: Named sessions (M6)
|
||||
|
||||
### Entry 10.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Optional task.session name -> persistent session dir dataRoot/sessions/<name> via pi --session-dir; resume most recent with -c when present; isolation per name; teach/recall demo fixtures (Gitea #22, #23).
|
||||
- Reason: L1 persistence is the prerequisite for any multi-step agent work.
|
||||
- Expected result: session dir populated after first run; second run recalls taught context.
|
||||
|
||||
### Entry 10.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: build; mock plumbing run; live teach/recall E2E; task suite.
|
||||
- Observed result: task suite 32/32; teach run replied REMEMBERED and session JSONL persisted host-side; recall run resumed (-c) and answered exactly 'mosaico'; single continued session file (no duplicate sessions).
|
||||
- Failure or correction: none. Design note: ephemeral (--no-session) remains the default when no session is declared.
|
||||
|
||||
## Phase 11: Operator ergonomics (M7)
|
||||
|
||||
### Entry 11.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: mosaic-task.mjs show <runId> (full record + snapshots + artifacts, traversal-safe), list with workspace/session columns, RELEASE -> 0.0.6, package + health-gated activate, docs (Gitea #24).
|
||||
- Reason: Run records are only as valuable as they are inspectable; release activation closes the loop on container-content changes.
|
||||
- Expected result: show works for real/missing/traversal ids; suites green; 0.0.6 active.
|
||||
|
||||
### Entry 11.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: build; show on real/missing/traversal ids; full sweep; release package + activate.
|
||||
- Observed result: config 24/24, task 32/32, release 14/14, verify PASS; 0.0.6 packaged and activated via exact-marker health gate.
|
||||
- Failure or correction: showRun initially rejected valid run ids (lowercase-only regex vs uppercase timestamp) and crashed on missing ids (uncaught readdir) — both fixed and covered.
|
||||
|
||||
## Autonomous run result
|
||||
|
||||
M5 tagged `workspace-capabilities-v1`, M6 tagged `sessions-v1`, M7 tagged `operator-ergonomics-v1`; release 0.0.6 active. Tracker: docs/plans/2026-09-03_autonomous-run.md.
|
||||
|
||||
---
|
||||
|
||||
## Phase 12: Conductor loop — self-orchestration (M8)
|
||||
|
||||
### Entry 12.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Stand up the poor-man orchestration loop per docs/plans/CONDUCTOR.md: conductor (host, holds git) mirrors the repo into a worker workspace, dispatches a headless pi worker (session worker-1, tools read/write/edit/bash) to implement retry <runId>, reviews the diff, integrates, verifies (Gitea #25, #26, #27).
|
||||
- Reason: The owner asked for circular task processing with agent workers; the stack now has every primitive needed — this proves it on the stack itself.
|
||||
- Expected result: worker-authored retry merged with suites green and a live retry verified.
|
||||
|
||||
### Entry 12.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: repo mirror clone; worker dispatch (tasks/worker-retry.json); diff review; apply; live retry; refinement dispatch (tasks/worker-retry-refine.json); second review; reverse+reapply combined patch; conductor interpolation fix; live retry; full sweep.
|
||||
- Observed result:
|
||||
- Worker round 1: implemented retry correctly per spec in 2m28s; diff reviewed clean.
|
||||
- Live retry exposed a spec gap (direct invocation lacks launcher env exports).
|
||||
- Worker round 2 (same session, 59s): made spawnEnv self-sufficient, but used PI_* names where compose interpolates MOSAIC_*.
|
||||
- Conductor hotfix: 3-line rename to MOSAIC_PROVIDER/MOSAIC_MODEL/MOSAIC_DATA_ROOT (too trivial for a worker round).
|
||||
- Final: live retry succeeded (replied REMEMBERED, new run recorded); all suites green.
|
||||
- Failure or correction: three rounds total — one spec gap (conductor), one naming mismatch (worker), one trivial rename (conductor). Each was caught by mechanical verification (run record stderr), never by hope.
|
||||
- Attribution: feature authored by headless pi worker (glm-5.3-flash) in sessions worker-1; conductor reviewed, integrated, and hotfixed.
|
||||
|
||||
## Result (M8)
|
||||
|
||||
Conductor loop proven end-to-end on the stack itself. `main` merged with M8; release 0.0.6 remains active (retry is host-side only, no image change).
|
||||
|
||||
---
|
||||
|
||||
## Phase 13: Mission-level capability policy (M9)
|
||||
|
||||
### Entry 13.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Missions may declare capabilities.tools as governing constraints (Gitea #30); merge semantics = least-privilege intersection (task narrows, never widens; empty intersection = tool-free run). Host-side only.
|
||||
- Reason: First mechanical restriction layer — the trust model becomes enforced, not instructed.
|
||||
- Expected result: all four merge cases asserted from run evidence; suites green.
|
||||
|
||||
### Entry 13.2 — after
|
||||
|
||||
- Observed: four merge cases verified deterministically via run-record stderr (mission-only, task-only, narrowed, emptied); invalid mission capabilities exit 2; task suite 41/41.
|
||||
- Failure or correction: selftest harness could not express ABSENT vs EMPTY fields via its printf helper — fixed with an ABSENT marker; two suite config-leak defects fixed (per-command env scoping). Product unaffected.
|
||||
|
||||
## Phase 14: Session forking (M11)
|
||||
|
||||
### Entry 14.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: sessionForkFrom task field branches the source session's newest file (pi --fork) into the target session dir; ancestor untouched; RELEASE -> 0.0.7 with health-gated activation (Gitea #33).
|
||||
- Reason: Owner flagged conversation forking from a common ancestor as a desired property; pi JSONL trees make it native.
|
||||
- Expected result: forked child recalls ancestor context; ancestor file untouched; suites green; 0.0.7 active.
|
||||
|
||||
### Entry 14.2 — after
|
||||
|
||||
- Observed: mock plumbing asserts fork source + target delivery; validation rejects fork-without-target and self-fork (exit 2); ghost source exits 4; live fork: child recalled 'mosaico' from ancestor context while the ancestor session file remained untouched (file-level assertion); suites 58/24/14 + verify green; 0.0.7 packaged and activated via health gate.
|
||||
- Failure or correction: retryRun-style self-assignment bug in validation (compared null to target) — caught by negative test, fixed.
|
||||
|
||||
## Result (M11)
|
||||
|
||||
Session forking verified. `main` merged with M11, tagged `session-fork-v1`; release 0.0.7 active.
|
||||
|
||||
|
||||
|
||||
+5
-2
@@ -18,11 +18,14 @@ WORKDIR /opt/app
|
||||
COPY package.json package-lock.json ./
|
||||
RUN npm ci --ignore-scripts
|
||||
|
||||
# Immutable contract fixtures (required location) and runtime scripts.
|
||||
# Immutable contract fixtures (required location), runtime scripts, and
|
||||
# runtime adapters.
|
||||
COPY contracts /opt/mosaic/contracts
|
||||
COPY src /opt/mosaic/src
|
||||
COPY adapters /opt/mosaic/adapters
|
||||
RUN chmod 0555 /opt/mosaic/contracts /opt/mosaic/contracts/* \
|
||||
&& chmod 0555 /opt/mosaic/src /opt/mosaic/src/*.sh
|
||||
&& chmod 0555 /opt/mosaic/src /opt/mosaic/src/*.sh \
|
||||
&& chmod 0555 /opt/mosaic/adapters /opt/mosaic/adapters/*/adapter.sh
|
||||
|
||||
# Writable state, workspace, and pi agent directory (auth.json is
|
||||
# bind-mounted read-only at runtime; nothing is copied into the image).
|
||||
|
||||
@@ -83,6 +83,67 @@ scripts/test-task.sh # selftests (schema negat
|
||||
|
||||
A run exits 0 only when its expectation is met (`expectExact` match); mismatches, nonzero agent exits, and timeouts record `status: failed` in `result.json` and exit 1. Each run gets a unique directory — rerunning never rewrites history.
|
||||
|
||||
## Release model (M3)
|
||||
|
||||
`RELEASE` single-sources the release version (0.0.X until declared stable); the image tag derives from it plus the pinned Pi version. Activation is health-gated and every event is recorded:
|
||||
|
||||
```bash
|
||||
scripts/release.sh package # build + tag the release image
|
||||
scripts/release.sh activate # health check (exact marker) -> atomic pointer swap
|
||||
scripts/release.sh activate --fault-injection # prove the refusal path (drills only)
|
||||
scripts/release.sh rollback # health-gated return to the previous release
|
||||
scripts/release.sh status # release, tag, active pointer, recent log
|
||||
scripts/test-release.sh # release selftests
|
||||
```
|
||||
|
||||
- `<dataRoot>/state/active.json` — the activation pointer (atomic tmp+rename replace)
|
||||
- `<dataRoot>/state/activation-log.jsonl` — append-only history: package / activate / refused / rollback
|
||||
|
||||
A failed health check never activates; the previously active release remains deployed. Updating the software therefore cannot corrupt the running installation: package beside, gate, then flip. Verified by the update/refusal/rollback drills in BUILD-LOG Phase 7.
|
||||
|
||||
## Runtime adapters (M4)
|
||||
|
||||
The harness boundary is formalized: everything upstream (config, contracts, missions, tasks, run records) is harness-agnostic; everything inside an adapter belongs to one runtime.
|
||||
|
||||
```text
|
||||
adapters/<name>/adapter.sh env in: MOSAIC_SYSTEM_PROMPT_FILE, MOSAIC_REQUEST,
|
||||
MOSAIC_PROVIDER, MOSAIC_MODEL
|
||||
stdout: response only; stderr: diagnostics
|
||||
```
|
||||
|
||||
- Selection: `execution.adapter` in config.json (optional; `pi` default; allowlist `pi`, `mock`)
|
||||
- `pi` — pinned Pi CLI, noninteractive print mode, ambient discovery off
|
||||
- `mock` — deterministic test adapter; never for real verification
|
||||
- Mission directives have a sanctioned injection point: when a task references a mission, the task runner mounts the run snapshot and the generated prompt gains a `MISSION (runtime)` section (objective + directives) after the four immutable contracts
|
||||
- Adding a harness (Claude, Codex, OpenCode) later means adding one directory — no orchestrator changes
|
||||
|
||||
See `adapters/README.md` for the full contract.
|
||||
|
||||
## Workspaces, capabilities, sessions (M5/M6)
|
||||
|
||||
Optional task fields extend what an agent can do — all defaulting to the previous behavior:
|
||||
|
||||
```json
|
||||
{
|
||||
"workspace": "demo", // ":run" ephemeral, or persistent dataRoot/workspaces/<name>
|
||||
"capabilities": { "tools": ["bash", "read"] }, // pi tool allowlist; absent = no tools
|
||||
"session": "demo" // persistent session at dataRoot/sessions/<name>
|
||||
}
|
||||
```
|
||||
|
||||
- The adapter runs inside the workspace; files it writes are host-visible (`dataRoot/workspaces/<name>`).
|
||||
- Sessions persist via pi's documented `--session-dir`; a follow-up run in the same session resumes the conversation (`-c`) and can recall prior context. Distinct names never share state. Ephemeral (`--no-session`) remains the default when no session is declared.
|
||||
- Selection authority: config for adapter/provider/model; the task file for workspace/capabilities/session.
|
||||
|
||||
Inspect anything:
|
||||
|
||||
```bash
|
||||
node scripts/mosaic-task.mjs list # runs with task/workspace/session columns
|
||||
node scripts/mosaic-task.mjs show <runId> # full record + snapshots + artifacts
|
||||
```
|
||||
|
||||
Demo fixtures: `tasks/workspace-demo.json`, `tasks/session-demo-1.json` + `tasks/session-demo-2.json`.
|
||||
|
||||
See `docs/plans/2026-09-02_atomic-mosaic-foundation.md` for the full plan.
|
||||
|
||||
Inside the container:
|
||||
@@ -95,7 +156,8 @@ Inside the container:
|
||||
|
||||
## How it works
|
||||
|
||||
1. `scripts/build.sh` builds `mosaic-poc-agent:0.84.4` with Docker Compose.
|
||||
1. `scripts/build.sh` builds the release image (`mosaic-poc-agent:<pi>-r<release>`,
|
||||
tag derived from `RELEASE` + the pinned Pi version) with Docker Compose.
|
||||
2. On each run, `/opt/mosaic/src/load-contracts.sh` reads the four contract files
|
||||
in fixed order (CONSTITUTION, STANDARDS, SOUL, USER), joins them with clear
|
||||
separators, and writes `/var/lib/mosaic/system-prompt.md`.
|
||||
@@ -116,8 +178,10 @@ scripts/build.sh # build the image
|
||||
scripts/hello.sh # one-shot request; prints the model response
|
||||
scripts/verify.sh # full gated test; exit 0 only on exact MOSAIC_HELLO_OK
|
||||
scripts/run-task.sh # run a mission/task file (see Missions & tasks)
|
||||
scripts/release.sh # package / activate / rollback / status (see Release model)
|
||||
scripts/test-config.sh # fast config-layer selftests (no Docker)
|
||||
scripts/test-task.sh # mission/task selftests (schema + live runs)
|
||||
scripts/test-task.sh # mission/task selftests (schema + adapter seam + live runs)
|
||||
scripts/test-release.sh # release selftests
|
||||
scripts/reset.sh # delete the configured data root (safety-checked)
|
||||
```
|
||||
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
# Mosaic runtime adapters
|
||||
|
||||
An adapter is the entire harness-specific surface of the system. Everything
|
||||
upstream of an adapter — configuration, contracts, missions, tasks, run
|
||||
records — is harness-agnostic; everything inside an adapter may assume one
|
||||
specific agent runtime.
|
||||
|
||||
## Contract
|
||||
|
||||
An adapter lives at:
|
||||
|
||||
```text
|
||||
/opt/mosaic/adapters/<name>/adapter.sh
|
||||
```
|
||||
|
||||
and must be executable. The dispatcher (`/opt/mosaic/src/run-agent.sh`)
|
||||
selects it via `MOSAIC_ADAPTER` (default: `pi`) and execs it after the
|
||||
system prompt has been generated.
|
||||
|
||||
**Inputs (environment):**
|
||||
|
||||
| Variable | Meaning |
|
||||
|---|---|
|
||||
| `MOSAIC_SYSTEM_PROMPT_FILE` | Absolute path to the generated system prompt (contracts + optional mission section). Read it; do not modify it. |
|
||||
| `MOSAIC_REQUEST` | The exact user request text (may contain newlines). |
|
||||
| `MOSAIC_PROVIDER` | Configured provider name. |
|
||||
| `MOSAIC_MODEL` | Configured model id. |
|
||||
|
||||
Optional, adapter-specific (documented per adapter):
|
||||
|
||||
| Variable | Meaning |
|
||||
|---|---|
|
||||
| `MOSAIC_MOCK_RESPONSE` | mock only: the verbatim response to emit |
|
||||
|
||||
**Outputs:**
|
||||
|
||||
- `stdout`: the model response text — the only channel the orchestrator captures
|
||||
- `stderr`: diagnostics (never credentials)
|
||||
- exit `0`: success; nonzero: failure
|
||||
|
||||
## Rules
|
||||
|
||||
1. Adapters print ONLY the response on stdout. Status lines go to stderr.
|
||||
2. Adapters never read configuration files; the resolved settings arrive via environment.
|
||||
3. Adapters never write outside `/var/lib/mosaic`.
|
||||
4. Adding an adapter requires: a new directory, the contract implementation, and
|
||||
adding the name to the allowlist in `scripts/mosaic-config.mjs`.
|
||||
|
||||
## Included adapters
|
||||
|
||||
- `pi` — the pinned `@earendil-works/pi-coding-agent` CLI in noninteractive
|
||||
print mode (`-p`), ambient discovery disabled, stdin detached.
|
||||
- `mock` — deterministic echo of `MOSAIC_MOCK_RESPONSE`. Test-only: never use
|
||||
it where a real model response is required.
|
||||
@@ -0,0 +1,18 @@
|
||||
#!/bin/sh
|
||||
# Mock adapter: deterministic response for seam tests. NEVER use where a
|
||||
# real model response is required.
|
||||
#
|
||||
# Contract: see /opt/mosaic/adapters/README.md.
|
||||
set -eu
|
||||
|
||||
[ -n "${MOSAIC_SYSTEM_PROMPT_FILE:-}" ] || { echo "mock adapter: MOSAIC_SYSTEM_PROMPT_FILE is required" >&2; exit 2; }
|
||||
if [ "${MOSAIC_INTERACTIVE:-}" != "1" ]; then
|
||||
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "mock adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
|
||||
fi
|
||||
[ -r "$MOSAIC_SYSTEM_PROMPT_FILE" ] || { echo "mock adapter: system prompt not readable: $MOSAIC_SYSTEM_PROMPT_FILE" >&2; exit 2; }
|
||||
|
||||
echo "mock adapter: responding verbatim from MOSAIC_MOCK_RESPONSE" >&2
|
||||
# Deterministic plumbing evidence: which MOSAIC_* variables did the
|
||||
# orchestrator actually deliver? (Auth secrets are not MOSAIC_-prefixed.)
|
||||
(env | grep '^MOSAIC_' | sort) >&2 2>/dev/null || true
|
||||
printf '%s\n' "${MOSAIC_MOCK_RESPONSE:-}"
|
||||
@@ -0,0 +1,83 @@
|
||||
#!/bin/sh
|
||||
# Pi adapter: implements the Mosaic adapter contract for the pinned
|
||||
# @earendil-works/pi-coding-agent CLI.
|
||||
#
|
||||
# Contract: see /opt/mosaic/adapters/README.md.
|
||||
# Headless (default): stdout = response only; stderr = diagnostics; exit 0.
|
||||
# Interactive (MOSAIC_INTERACTIVE=1): full pi TUI on the attached terminal.
|
||||
set -eu
|
||||
|
||||
[ -n "${MOSAIC_SYSTEM_PROMPT_FILE:-}" ] || { echo "pi adapter: MOSAIC_SYSTEM_PROMPT_FILE is required" >&2; exit 2; }
|
||||
[ -r "$MOSAIC_SYSTEM_PROMPT_FILE" ] || { echo "pi adapter: system prompt not readable: $MOSAIC_SYSTEM_PROMPT_FILE" >&2; exit 2; }
|
||||
# MOSAIC_AGENT_NAME is optional in headless mode (identity section is then
|
||||
# omitted); interactive launches always set it via scripts/agent.sh.
|
||||
|
||||
: "${PI_PROVIDER:?pi adapter: PI_PROVIDER is required}"
|
||||
: "${PI_MODEL:?pi adapter: PI_MODEL is required}"
|
||||
|
||||
INTERACTIVE="${MOSAIC_INTERACTIVE:-}"
|
||||
if [ "$INTERACTIVE" != "1" ]; then
|
||||
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "pi adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
|
||||
fi
|
||||
|
||||
# Workspace (M5): run inside the provided workspace when present.
|
||||
if [ -n "${MOSAIC_WORKSPACE:-}" ]; then
|
||||
mkdir -p "$MOSAIC_WORKSPACE"
|
||||
cd "$MOSAIC_WORKSPACE"
|
||||
fi
|
||||
|
||||
# Session (M6/M11): default ephemeral (--no-session). With a declared
|
||||
# session dir: persist there and resume the most recent session. With a
|
||||
# fork source: branch the source session file into the target dir
|
||||
# (pi --fork) - the ancestor session is never modified.
|
||||
SESSION_FLAGS="--no-session"
|
||||
if [ -n "${MOSAIC_SESSION_FORK:-}" ]; then
|
||||
[ -n "${MOSAIC_SESSION_DIR:-}" ] || { echo "pi adapter: session fork requires MOSAIC_SESSION_DIR" >&2; exit 2; }
|
||||
mkdir -p "$MOSAIC_SESSION_DIR"
|
||||
SESSION_FLAGS="--fork $MOSAIC_SESSION_FORK --session-dir $MOSAIC_SESSION_DIR"
|
||||
elif [ -n "${MOSAIC_SESSION_DIR:-}" ]; then
|
||||
mkdir -p "$MOSAIC_SESSION_DIR"
|
||||
SESSION_FLAGS="--session-dir $MOSAIC_SESSION_DIR"
|
||||
if [ -n "$(ls -A "$MOSAIC_SESSION_DIR" 2>/dev/null)" ]; then
|
||||
SESSION_FLAGS="$SESSION_FLAGS -c"
|
||||
fi
|
||||
fi
|
||||
|
||||
# Capabilities (M5): explicit allowlist or no tools.
|
||||
TOOLS_FLAG="--no-tools"
|
||||
[ -n "${MOSAIC_TOOLS:-}" ] && TOOLS_FLAG="--tools $MOSAIC_TOOLS"
|
||||
|
||||
# Mode (M13): interactive TUI or one-shot print.
|
||||
PRINT_MODE="-p"
|
||||
REQUEST_ARG=""
|
||||
if [ "$INTERACTIVE" = "1" ]; then
|
||||
PRINT_MODE=""
|
||||
else
|
||||
REQUEST_ARG="$MOSAIC_REQUEST"
|
||||
fi
|
||||
|
||||
# All flags documented in the pi package README (CLI Reference):
|
||||
# -p/--print one-shot mode: print the response and exit (omitted in
|
||||
# interactive TUI mode)
|
||||
# --system-prompt replace the default prompt with the generated one
|
||||
# --no-* no ambient context/skills/extensions/templates/themes
|
||||
# SESSION_FLAGS ephemeral | persistent | forked (per env)
|
||||
# TOOLS_FLAG per capabilities
|
||||
# --offline no startup network operations (update checks/telemetry)
|
||||
PROMPT_CONTENT="$(cat "$MOSAIC_SYSTEM_PROMPT_FILE")"
|
||||
set -- \
|
||||
--offline \
|
||||
--no-extensions \
|
||||
--no-skills \
|
||||
--no-prompt-templates \
|
||||
--no-themes \
|
||||
--no-context-files \
|
||||
$TOOLS_FLAG \
|
||||
$SESSION_FLAGS \
|
||||
--provider "$PI_PROVIDER" \
|
||||
--model "$PI_MODEL" \
|
||||
--system-prompt "$PROMPT_CONTENT"
|
||||
# One-shot mode appends -p and the request (both safely quoted);
|
||||
# interactive mode appends nothing - clean TUI.
|
||||
[ "$INTERACTIVE" = "1" ] || set -- "$@" -p "$MOSAIC_REQUEST"
|
||||
exec pi "$@"
|
||||
+20
-4
@@ -3,13 +3,29 @@ services:
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Containerfile
|
||||
image: mosaic-poc-agent:0.84.4
|
||||
image: ${MOSAIC_IMAGE_TAG:?MOSAIC_IMAGE_TAG must be set by scripts/load_release (run via scripts/*.sh)}
|
||||
user: "1000:1000"
|
||||
environment:
|
||||
# Resolved from config.json by scripts/common.sh (load_config).
|
||||
# Required: compose fails fast when the launcher did not supply them.
|
||||
PI_PROVIDER: ${MOSAIC_PROVIDER:?MOSAIC_PROVIDER must be set by scripts/load_config (run via scripts/*.sh)}
|
||||
PI_MODEL: ${MOSAIC_MODEL:?MOSAIC_MODEL must be set by scripts/load_config (run via scripts/*.sh)}
|
||||
# Adapter selection (resolved from config execution.adapter; default pi)
|
||||
MOSAIC_ADAPTER: ${MOSAIC_ADAPTER:-pi}
|
||||
# Mission directives injection point (set by the task runner when the
|
||||
# task references a mission; container path of the run snapshot)
|
||||
MOSAIC_MISSION_FILE: ${MOSAIC_MISSION_FILE:-}
|
||||
# Workspace + capabilities (set by the task runner; M5)
|
||||
MOSAIC_WORKSPACE: ${MOSAIC_WORKSPACE:-}
|
||||
MOSAIC_TOOLS: ${MOSAIC_TOOLS:-}
|
||||
# Persistent named session dir + optional fork source (M6/M11)
|
||||
MOSAIC_SESSION_DIR: ${MOSAIC_SESSION_DIR:-}
|
||||
MOSAIC_SESSION_FORK: ${MOSAIC_SESSION_FORK:-}
|
||||
# Interactive TUI mode + agent identity (M13, set by scripts/agent.sh)
|
||||
MOSAIC_INTERACTIVE: ${MOSAIC_INTERACTIVE:-}
|
||||
MOSAIC_AGENT_NAME: ${MOSAIC_AGENT_NAME:-}
|
||||
# mock adapter only: verbatim response for deterministic seam tests
|
||||
MOSAIC_MOCK_RESPONSE: ${MOSAIC_MOCK_RESPONSE:-}
|
||||
# Documented container auth alternative: provider API key via
|
||||
# runtime environment variable. Empty by default; when empty Pi
|
||||
# falls back to the read-only mounted auth.json credential file.
|
||||
@@ -21,6 +37,6 @@ services:
|
||||
# Runtime credential only: pi auth file mounted READ-ONLY.
|
||||
# Never copied into the image.
|
||||
- ${PI_AUTH_FILE:-/home/jwoltje/.pi/agent/auth.json}:/home/node/.pi/agent/auth.json:ro
|
||||
# One-shot: the exact startup verification request. It deliberately
|
||||
# does NOT contain the expected marker MOSAIC_HELLO_OK.
|
||||
command: ["Return your startup marker and nothing else."]
|
||||
# Headless runs: the request is passed as command args by the launchers
|
||||
# (run-task.sh) or defaults inside run-agent.sh (hello/verify). Never a
|
||||
# fixed command here - interactive runs (scripts/agent.sh) need no args.
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
# Session registry — append-only
|
||||
|
||||
Every agent session (assistant, worker-cycle conductor, or owner-directed
|
||||
automation) that works in this repository registers one line here. Entries
|
||||
are never rewritten or removed; corrections are new entries.
|
||||
|
||||
| Date (UTC) | Actor | Scope | Outcome / artifacts |
|
||||
|---|---|---|---|
|
||||
| 2026-09-03 | assistant (conductor + worker) | POC through M12: containerized pi proof, config layer, missions/tasks, release model, adapter seam, workspaces/capabilities, named sessions, retention, session forking, conductor auto-apply, roles/ convention | 13 tags; suites 24/58/14 + 17 conductor + verify green; releases 0.0.1–0.0.7; issues #1–#34 closed |
|
||||
@@ -0,0 +1,85 @@
|
||||
# TOOLS.md — command and tool reference
|
||||
|
||||
On-demand reference for agent sessions (conductors, bootstrapping agents,
|
||||
reviewers). `AGENTS.md` routes here; this file carries the depth: usage,
|
||||
inputs/outputs, exit codes, and safety notes for every entry point.
|
||||
|
||||
Reading guide: all entry points are `scripts/*.sh` (bash) or invoked via
|
||||
`node scripts/mosaic-task.mjs` (node). Every script fails closed — missing
|
||||
or invalid configuration/policy refuses the operation with a nonzero exit
|
||||
and changes nothing.
|
||||
|
||||
## Lifecycle
|
||||
|
||||
| Command | Purpose | Notes |
|
||||
|---|---|---|
|
||||
| `scripts/bootstrap.sh` | Create `~/.config/mosaic-dev/config.json` if absent | Idempotent; existing config validated, never rewritten |
|
||||
| `scripts/build.sh` | Build the release image | Tag derived from `RELEASE` + pinned pi version |
|
||||
| `scripts/hello.sh` | One-shot startup request | Prints model response on stdout |
|
||||
| `scripts/verify.sh` | Full gated test | Exit 0 only on exact `MOSAIC_HELLO_OK`; `EXPECTED_MARKER` overrides for negative drills |
|
||||
|
||||
## Tasks (missions, runs, evidence)
|
||||
|
||||
| Command | Purpose | Notes |
|
||||
|---|---|---|
|
||||
| `scripts/run-task.sh run <task.json>` | Execute a task | Immutable run record under `<dataRoot>/runs/` |
|
||||
| `scripts/run-task.sh validate <task.json>` | Strict validation | Writes nothing |
|
||||
| `node scripts/mosaic-task.mjs show <runId>` | Inspect a run | Full record + snapshots + artifacts |
|
||||
| `node scripts/mosaic-task.mjs list` | List runs | task/workspace/session columns |
|
||||
| `node scripts/mosaic-task.mjs retry <runId>` | Re-execute a run's snapshot | New run dir; `retriedFrom` lineage recorded |
|
||||
| `node scripts/mosaic-task.mjs prune [--keep=N] [--yes]` | Retention | Dry-run default; receipt in `runs/.pruned.log` |
|
||||
|
||||
Task fields: `prompt` (required), `mission` (path), `expectExact`,
|
||||
`timeoutSeconds` (5–600), `workspace` (`:run` or named), `capabilities.tools`
|
||||
(allowlist: read write edit bash grep find ls), `session`,
|
||||
`sessionForkFrom` (requires `session`). Mission fields: `objective`,
|
||||
`directives[]`, optional governing `capabilities.tools`. Policy: a task may
|
||||
narrow a mission's tools, never widen; empty intersection = tool-free run.
|
||||
|
||||
## Agent (interactive TUI)
|
||||
|
||||
```bash
|
||||
scripts/agent.sh <name> [--mission <file>] [--workspace <ws>] [--session <s>] [--tools <list>]
|
||||
```
|
||||
|
||||
Launches an interactive pi TUI inside the container with the four immutable
|
||||
contracts + optional mission + agent identity as its system prompt,
|
||||
persistent named session, optional workspace. Exit with `/quit`.
|
||||
|
||||
## Release
|
||||
|
||||
| Command | Purpose | Notes |
|
||||
|---|---|---|
|
||||
| `scripts/release.sh package` | Build + tag the release image | Tag: `mosaic-poc-agent:<pi>-r<release>` |
|
||||
| `scripts/release.sh activate` | Health gate → atomic pointer swap | `--fault-injection` proves the refusal path |
|
||||
| `scripts/release.sh rollback` | Health-gated return to previous | Refuses if image missing |
|
||||
| `scripts/release.sh status` | Release, tag, active pointer, log | Safe on empty state |
|
||||
|
||||
## Conductor (worker patches)
|
||||
|
||||
```bash
|
||||
scripts/conductor-apply.sh <runId> [--dry-run]
|
||||
```
|
||||
|
||||
Auto-applies a worker's patch under `roles/conductor-policy.json`:
|
||||
succeeded run → clean target tree → path allowlist → syntax gates →
|
||||
apply → policy suites → attribution commit. Any failure reverts.
|
||||
Push is never automatic.
|
||||
|
||||
## Maintenance
|
||||
|
||||
| Command | Purpose | Notes |
|
||||
|---|---|---|
|
||||
| `scripts/reset.sh` | Delete the data root | Triple-safety-checked (path, symlink, ownership marker) |
|
||||
| `scripts/test-config.sh` | Config selftests (no Docker) | 24 cases |
|
||||
| `scripts/test-task.sh` | Task selftests + live cases | 58 cases |
|
||||
| `scripts/test-release.sh` | Release selftests | 14 cases |
|
||||
| `scripts/test-conductor.sh` | Auto-apply selftests (sandboxed) | 17 cases |
|
||||
| `scripts/gitea-api.sh <METHOD> <path> [body]` | Gitea API helper | Token never on argv/stdout |
|
||||
|
||||
## Exit-code convention
|
||||
|
||||
`0` success · `1` operation failed · `2` invalid data/configuration ·
|
||||
`3` configuration missing for a read operation · `4` usage/file/environment
|
||||
problem. Scripts print diagnostics on stderr; model responses (and only
|
||||
model responses) on stdout.
|
||||
@@ -0,0 +1,81 @@
|
||||
# Autonomous Work Run — 2026-09-03
|
||||
|
||||
**Status:** COMPLETED (single-session batch; see Results at bottom)
|
||||
**Constraint:** The assistant cannot run unattended. This was one long interactive session, not 12 wall-clock hours. Everything below was completed, committed, and pushed during that session.
|
||||
|
||||
## Objective
|
||||
|
||||
Advance the Mosaic Stack rebuild several verified layers in one batch, focused on Pi, ending in a state the owner can test and review alone: green suites, activated release, recorded drills, and this document as the single entry point.
|
||||
|
||||
## Scope decided for this run
|
||||
|
||||
| Milestone | Theme | Status |
|
||||
|---|---|---|
|
||||
| M5 | Task workspaces + capability envelope (tools allowlist) | DONE |
|
||||
| M6 | Named sessions — persistence and resume (L1) | DONE |
|
||||
| M7 | Operator ergonomics: run inspection commands | DONE |
|
||||
| — | Releases 0.0.5+0.0.6 packaged; 0.0.6 health-gated activated | DONE |
|
||||
|
||||
Explicitly deferred (do not mistake for forgotten):
|
||||
- Claude/Codex/OpenCode adapters (owner: focus on Pi for now)
|
||||
- Network policy engine (container boundary is the current control)
|
||||
- Fine-grained read restrictions (excluded by the original brief)
|
||||
- Config/state migrations (no schema breaks so far; keep it that way)
|
||||
|
||||
## Design decisions taken during this run
|
||||
|
||||
1. **Workspace** (`task.workspace`, optional):
|
||||
- absent → tool-free text-only run (previous behavior, unchanged)
|
||||
- `":run"` → ephemeral per-run workspace at `<dataRoot>/runs/<runId>/workspace`
|
||||
- named (validated id) → persistent shared workspace at `<dataRoot>/workspaces/<name>`
|
||||
- Container path passed via `MOSAIC_WORKSPACE` env; adapter cds into it. No new mounts (dataRoot is already mounted).
|
||||
2. **Capabilities** (`task.capabilities.tools`, optional): allowlist from pi's documented tool set (`read write edit bash grep find ls`). Absent → `--no-tools` (previous behavior). Passed via `MOSAIC_TOOLS` env; pi adapter maps to `--tools`.
|
||||
3. **Adapter diagnostics for deterministic testing**: the mock adapter writes all received `MOSAIC_*` variables (never secrets — auth is not MOSAIC_-prefixed) to stderr, which lands in the run record. This lets selftests assert orchestrator→adapter plumbing without parsing model output.
|
||||
4. **Sessions** (`task.session`, optional named): persisted under `<dataRoot>/sessions/<name>/` via pi's documented `--session-dir`; resume semantics: continue most recent session in that directory when one exists (`-c`).
|
||||
5. **Selection authority unchanged**: config file for adapter/provider/model; task file for workspace/capabilities/session; env vars are internal plumbing only.
|
||||
6. **configVersion stays 1**; all new task fields are optional. Old tasks/configs remain valid.
|
||||
|
||||
## Test plan (what "done" means per milestone)
|
||||
|
||||
- M5: mock-adapter cases asserting workspace path and tools arrive via run-record stderr; live pi case writing/reading a file in a persistent workspace; validation negatives (bad tool name, bad workspace name)
|
||||
- M6: session directory deterministically populated after first run; second run resumes (continuation asserted by session dir state and, in live E2E, by model recall); sandbox isolation between two named sessions
|
||||
- M7: `show <runId>` prints a complete run record; `list` gains workspace/session columns
|
||||
- Final: full sweep (config/task/release), verify, package + activate 0.0.6, config checksum unchanged
|
||||
|
||||
## Review checklist for the owner
|
||||
|
||||
1. `cat docs/plans/2026-09-03_autonomous-run.md` (this file)
|
||||
2. `scripts/release.sh status` → 0.0.6 active
|
||||
3. `scripts/test-config.sh && scripts/test-task.sh && scripts/test-release.sh && scripts/verify.sh`
|
||||
4. Try a workspace task:
|
||||
```bash
|
||||
scripts/run-task.sh run tasks/workspace-demo.json
|
||||
ls ~/.mosaic-dev/workspaces/demo/
|
||||
```
|
||||
5. Try the session demo:
|
||||
```bash
|
||||
scripts/run-task.sh run tasks/session-demo-1.json # teaches a word
|
||||
scripts/run-task.sh run tasks/session-demo-2.json # recalls it
|
||||
```
|
||||
6. Inspect any run: `node scripts/mosaic-task.mjs show <runId>`
|
||||
7. Gitea: milestones M5/M6/M7 closed; issues referenced by merge commits
|
||||
|
||||
## Results
|
||||
|
||||
- M5 merged on `main` (merge commit `ddb1554`), tagged `workspace-capabilities-v1`
|
||||
- M6 merged on `main` (merge commit `4e2a413`), tagged `sessions-v1`
|
||||
- M7 merged on `main`, tagged `operator-ergonomics-v1`
|
||||
- Release 0.0.6 packaged, health-gated activated, full sweep green
|
||||
- Suites at end of run: config 24/24, task 32/32, release 14/14, verify PASS
|
||||
- Build log: Phases 9 (M5), 10 (M6), 11 (M7) appended with corrections
|
||||
- Corrections encountered: dropped constant from a failed atomic edit batch (SUPPORTED_TOOLS); dash `export` output format vs `env`; three selftest authoring defects; showRun id-regex case sensitivity + missing-run crash. All fixed and covered by tests.
|
||||
- Commits pushed incrementally; nothing left uncommitted
|
||||
- Live proof: workspace file host-visible; session teach/recall ('mosaico') verified
|
||||
|
||||
## Next steps after this run (not started)
|
||||
|
||||
1. Owner review + hands-on testing of workspaces, capabilities, sessions
|
||||
2. Decision: capability defaults per mission (mission-level policy) — natural M8
|
||||
3. Second real adapter remains available whenever wanted
|
||||
4. Consider run-record pruning/retention policy once run volume grows
|
||||
5. Consider a `mosaic-task.mjs retry <runId>` convenience for failed runs
|
||||
@@ -0,0 +1,58 @@
|
||||
# Conductor protocol — poor-man orchestration loop
|
||||
|
||||
How the stack orchestrates headless pi workers to do work on itself.
|
||||
|
||||
## Role contracts
|
||||
|
||||
Role authority is declared in role contracts, one file per role, under
|
||||
`roles/` (e.g. `roles/conductor-policy.json`). The repository root holds
|
||||
only first-class, bootstrap-required configuration; role contracts are
|
||||
tracked, versioned files whose changes arrive as reviewed commits.
|
||||
|
||||
## Roles
|
||||
|
||||
| Role | Runs where | Powers | Never has |
|
||||
|---|---|---|---|
|
||||
| **Conductor** | host (assistant or owner) | git (clone/commit/push), task dispatch, review, verification suites, Gitea | nothing new |
|
||||
| **Worker** | container (headless pi via `scripts/run-task.sh`) | read/write/edit/bash inside its workspace; persistent session on request | git credentials, docker socket, host filesystem |
|
||||
|
||||
## The loop
|
||||
|
||||
1. **Decompose**: conductor turns a goal into worker tasks small enough to
|
||||
specify completely in one prompt (file paths, acceptance criteria, style
|
||||
constraints, verification the worker can run itself, e.g. `node --check`).
|
||||
2. **Mirror**: conductor maintains the repo clone at
|
||||
`<dataRoot>/workspaces/stack-repo` (host-side git; workers see it read-write
|
||||
through their workspace mount).
|
||||
3. **Dispatch**: `scripts/run-task.sh run <worker-task.json>` — worker edits the
|
||||
clone. Session name `worker-<n>` keeps continuity across refinement rounds.
|
||||
4. **Extract**: `git -C <workspace> diff > patch` — the worker's entire output
|
||||
is a reviewable diff. Run record (result.json, stderr.txt) is the receipt.
|
||||
5. **Review**: conductor reads the diff line by line. Bad output → refine the
|
||||
prompt, re-dispatch (same session: "your patch had these problems…").
|
||||
6. **Integrate**: conductor applies the patch to the real repo, runs the full
|
||||
suites, commits and pushes. Suites failing → revert apply, back to step 5.
|
||||
7. **Record**: update CURRENT.md, BUILD-LOG, close the Gitea issue.
|
||||
|
||||
## Guardrails
|
||||
|
||||
- Workers never receive credentials; they never run git; they never leave the
|
||||
workspace (container is the boundary; tools allowlist is the gate).
|
||||
- Every worker diff is reviewed by the conductor before integration. No
|
||||
auto-apply. (Auto-apply would be a capability-policy decision for later.)
|
||||
- Verification is mechanical: suites + `node --check` / `bash -n` gates.
|
||||
- Recursive decomposition = "fail → smaller task", never "hope."
|
||||
|
||||
## Worker task template
|
||||
|
||||
```json
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-worker-<name>",
|
||||
"prompt": "<full spec: goal, files, constraints, acceptance, self-checks>",
|
||||
"workspace": "stack-repo",
|
||||
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
|
||||
"session": "worker-1",
|
||||
"timeoutSeconds": 600
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,61 @@
|
||||
# CURRENT — single source of "what happens next"
|
||||
|
||||
This file always names exactly one next action. Any "continue" / "next" /
|
||||
"proceed" message means: execute the action below, fully (implement → test →
|
||||
verify against its acceptance criteria → commit → push → close the issue →
|
||||
update this file to the next action). No ambiguity, no re-planning.
|
||||
|
||||
## Next action
|
||||
|
||||
Owner review of M12 (conductor auto-apply policy) — then name the next target.
|
||||
|
||||
## Queue (ordered, not started)
|
||||
|
||||
1. Second real adapter (parked — owner focused on Pi)
|
||||
2. Session forking from a common ancestor — SHIPPED in M11; exercise via session demos
|
||||
3. UX/DX backlog #31 (grep highlight vs status colors — likely closed by owner's unalias)
|
||||
4. Push policy decision: auto-apply commits locally; push remains explicit (documented in CONDUCTOR.md)
|
||||
|
||||
## Rules
|
||||
|
||||
- One action in flight. Update this file at the END of every action.
|
||||
- Blocked? Move the item to "Blocked" below with the reason and stop.
|
||||
- Completed actions move to the log at the bottom (date + issue + result).
|
||||
|
||||
## Blocked
|
||||
|
||||
(none)
|
||||
|
||||
## Completed log
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
|
||||
- 2026-09-03 — M12 conductor auto-apply policy (#34) — committed on main, 17/24/58/14 + verify green; push stays explicit
|
||||
- 2026-09-03 — repository convention: role contracts move to roles/ (root = bootstrap-only, per owner direction)
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
|
||||
- 2026-09-03 — M12 conductor auto-apply policy (#34) — committed on main, 17/24/58/14 + verify green; push stays explicit
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
|
||||
- 2026-09-03 — M5 workspaces + capability envelope (#20, #21) — merged, tests green
|
||||
- 2026-09-03 — M6 named sessions + resume (#22, #23) — merged, teach/recall verified
|
||||
- 2026-09-03 — M7 run inspection + release 0.0.6 (#24) — merged, activated
|
||||
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
|
||||
- 2026-09-03 — retry lineage + relative mission resolution (#28) — merged, 36/24/14 suites + verify green
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection, 41/36/14 + verify green
|
||||
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
|
||||
- 2026-09-03 — retry lineage + relative mission resolution (#28) — merged, 36/24/14 suites + verify green
|
||||
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"missionVersion": 1,
|
||||
"id": "m-hello",
|
||||
"objective": "Prove the startup marker path of the mosaic-poc-agent.",
|
||||
"objective": "Verify the startup marker path.",
|
||||
"directives": [
|
||||
"Startup verification requests are answered with the marker only.",
|
||||
"No explanation, no formatting."
|
||||
|
||||
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "mosaic-stack-dev-test",
|
||||
"version": "0.1.0",
|
||||
"version": "0.0.3",
|
||||
"private": true,
|
||||
"description": "Minimal Mosaic Stack container proof of concept: one Pi agent, four local contract files, one real model request returning MOSAIC_HELLO_OK.",
|
||||
"license": "UNLICENSED",
|
||||
|
||||
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"policyVersion": 1,
|
||||
"autoApply": {
|
||||
"enabled": true,
|
||||
"allowedPaths": [
|
||||
"scripts/**",
|
||||
"docs/**",
|
||||
"tasks/**",
|
||||
"missions/**",
|
||||
"adapters/**",
|
||||
"README.md"
|
||||
],
|
||||
"suites": [
|
||||
"test-config",
|
||||
"test-task",
|
||||
"test-release"
|
||||
]
|
||||
}
|
||||
}
|
||||
Executable
+67
@@ -0,0 +1,67 @@
|
||||
#!/usr/bin/env bash
|
||||
# Launch an interactive (TUI) Mosaic agent in its container.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/agent.sh <name> [--mission <file>] [--workspace <ws>]
|
||||
# [--session <name>] [--tools <comma,list>]
|
||||
#
|
||||
# The agent receives the four immutable contracts (constitution, standards,
|
||||
# SOUL, USER) plus its own identity and optional mission directives as its
|
||||
# system prompt, a persistent named session, and - if declared - a
|
||||
# workspace and tool capabilities. The TUI opens clean; you drive.
|
||||
#
|
||||
# This is the Mosaic alternative to launching vanilla pi: same engine,
|
||||
# governed context.
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
# shellcheck source=common.sh
|
||||
source scripts/common.sh
|
||||
|
||||
NAME=""
|
||||
MISSION=""
|
||||
WORKSPACE=""
|
||||
SESSION=""
|
||||
TOOLS=""
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
--mission) MISSION="${2:?}"; shift 2 ;;
|
||||
--workspace) WORKSPACE="${2:?}"; shift 2 ;;
|
||||
--session) SESSION="${2:?}"; shift 2 ;;
|
||||
--tools) TOOLS="${2:?}"; shift 2 ;;
|
||||
--help|-h) sed -n '2,12p' "$0"; exit 0 ;;
|
||||
*) NAME="$1"; shift ;;
|
||||
esac
|
||||
done
|
||||
|
||||
[ -n "$NAME" ] || { echo "agent: usage: scripts/agent.sh <name> [--mission f] [--workspace ws] [--session s] [--tools list]" >&2; exit 4; }
|
||||
case "$NAME" in *[!A-Za-z0-9._-]*|'') echo "agent: invalid agent name" >&2; exit 4;; esac
|
||||
|
||||
load_config
|
||||
load_release
|
||||
bootstrap_runtime_dir
|
||||
|
||||
SESSION="${SESSION:-agent-$NAME}"
|
||||
mkdir -p "$MOSAIC_DEV_DIR/sessions/$SESSION"
|
||||
export MOSAIC_SESSION_DIR="/var/lib/mosaic/sessions/$SESSION"
|
||||
export MOSAIC_AGENT_NAME="$NAME"
|
||||
export MOSAIC_INTERACTIVE=1
|
||||
export MOSAIC_TOOLS="${TOOLS:+$TOOLS}"
|
||||
|
||||
if [ -n "$MISSION" ]; then
|
||||
[ -r "$MISSION" ] || { echo "agent: mission file not readable: $MISSION" >&2; exit 4; }
|
||||
mkdir -p "$MOSAIC_DEV_DIR/agent-missions"
|
||||
cp "$MISSION" "$MOSAIC_DEV_DIR/agent-missions/$NAME.json"
|
||||
export MOSAIC_MISSION_FILE="/var/lib/mosaic/agent-missions/$NAME.json"
|
||||
fi
|
||||
|
||||
if [ -n "$WORKSPACE" ]; then
|
||||
case "$WORKSPACE" in *[!A-Za-z0-9._-]*|'') echo "agent: invalid workspace name" >&2; exit 4;; esac
|
||||
mkdir -p "$MOSAIC_DEV_DIR/workspaces/$WORKSPACE"
|
||||
export MOSAIC_WORKSPACE="/var/lib/mosaic/workspaces/$WORKSPACE"
|
||||
fi
|
||||
|
||||
echo "agent: launching TUI agent '$NAME' (session: $SESSION, adapter: $MOSAIC_ADAPTER, model: $MOSAIC_MODEL)"
|
||||
echo "agent: contracts + $([ -n "$MISSION" ] && echo 'mission' || echo 'no mission') loaded; exit the TUI with /quit"
|
||||
# No -T: the TTY is the point. Ctrl+C twice or /quit exits.
|
||||
exec docker compose run --rm mosaic-agent
|
||||
@@ -6,6 +6,7 @@ cd "$(dirname "$0")/.."
|
||||
source scripts/common.sh
|
||||
|
||||
load_config
|
||||
load_release
|
||||
|
||||
bootstrap_runtime_dir
|
||||
|
||||
|
||||
+16
-1
@@ -15,10 +15,25 @@ load_config() {
|
||||
exit 1
|
||||
fi
|
||||
eval "$config_env"
|
||||
export MOSAIC_DATA_ROOT MOSAIC_PROVIDER MOSAIC_MODEL
|
||||
export MOSAIC_DATA_ROOT MOSAIC_PROVIDER MOSAIC_MODEL MOSAIC_ADAPTER
|
||||
MOSAIC_DEV_DIR="$MOSAIC_DATA_ROOT"
|
||||
}
|
||||
|
||||
# Resolve the release identity: RELEASE is the single source of the
|
||||
# release version (stays 0.0.X until declared stable); the image tag
|
||||
# derives from it plus the pinned pi dependency version.
|
||||
load_release() {
|
||||
local release pi_version
|
||||
release="$(tr -d '[:space:]' < RELEASE)"
|
||||
if ! printf '%s' "$release" | grep -Eq '^[0-9]+\.[0-9]+\.[0-9]+$'; then
|
||||
echo "common: RELEASE must be a semver-ish version, got: '$release'" >&2
|
||||
exit 1
|
||||
fi
|
||||
pi_version="$(node -p "require('./package.json').dependencies['@earendil-works/pi-coding-agent']")"
|
||||
export MOSAIC_RELEASE="$release"
|
||||
export MOSAIC_IMAGE_TAG="mosaic-poc-agent:${pi_version}-r${release}"
|
||||
}
|
||||
|
||||
# Ensure the configured runtime data directory exists and carries this
|
||||
# project's ownership marker. The marker is what scripts/reset.sh requires
|
||||
# before it will delete anything.
|
||||
|
||||
Executable
+139
@@ -0,0 +1,139 @@
|
||||
#!/usr/bin/env bash
|
||||
# Conductor auto-apply: integrate a worker's patch under the declared policy.
|
||||
#
|
||||
# Usage: scripts/conductor-apply.sh <runId> [--dry-run]
|
||||
#
|
||||
# Policy (roles/conductor-policy.json in the target repo, strictly validated):
|
||||
# autoApply.enabled master switch
|
||||
# autoApply.allowedPaths glob allowlist ('dir/**' = everything under dir)
|
||||
# autoApply.suites suite scripts that must pass AFTER applying
|
||||
#
|
||||
# Gate sequence: succeeded run record -> clean target tree -> diff extracted
|
||||
# from the worker workspace -> allowlist -> syntax gates (node/bash/json) ->
|
||||
# apply -> policy suites -> commit with attribution. ANY failure reverts the
|
||||
# working tree and exits nonzero. Push is never automatic.
|
||||
#
|
||||
# Environment:
|
||||
# MOSAIC_APPLY_TARGET repo root to apply into (default: this project)
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
TARGET_ROOT="${MOSAIC_APPLY_TARGET:-$(cd "$SCRIPT_DIR/.." && pwd)}"
|
||||
RUN_ID="${1:?usage: conductor-apply.sh <runId> [--dry-run]}"
|
||||
DRY_RUN="no"
|
||||
[ "${2:-}" = "--dry-run" ] && DRY_RUN="yes"
|
||||
|
||||
cd "$TARGET_ROOT"
|
||||
|
||||
fail() { echo "conductor-apply: $*" >&2; exit "${2:-1}"; }
|
||||
[ -d .git ] || fail "target is not a git repository: $TARGET_ROOT" 4
|
||||
[ -f roles/conductor-policy.json ] || fail "no roles/conductor-policy.json in target" 2
|
||||
|
||||
# ---- policy (strict) ----
|
||||
POLICY_JSON="$(node -e '
|
||||
const fs = require("fs");
|
||||
const p = JSON.parse(fs.readFileSync("roles/conductor-policy.json", "utf8"));
|
||||
if (p.policyVersion !== 1) process.exit(3);
|
||||
if (!p.autoApply || typeof p.autoApply.enabled !== "boolean" || !Array.isArray(p.autoApply.allowedPaths) || !Array.isArray(p.autoApply.suites)) process.exit(3);
|
||||
for (const g of p.autoApply.allowedPaths) {
|
||||
if (typeof g !== "string" || !/^[A-Za-z0-9_.*/-]+$/.test(g) || g.startsWith("/") || g.includes("..")) process.exit(3);
|
||||
}
|
||||
console.log(JSON.stringify(p.autoApply));
|
||||
')" || fail "invalid roles/conductor-policy.json" 2
|
||||
|
||||
ENABLED="$(node -e 'console.log(JSON.parse(process.argv[1]).enabled)' "$POLICY_JSON")"
|
||||
[ "$ENABLED" = "true" ] || fail "auto-apply is disabled by policy" 2
|
||||
|
||||
# ---- run record ----
|
||||
DATA_ROOT="$(node scripts/mosaic-config.mjs validate | node -e 'let d="";process.stdin.on("data",c=>d+=c).on("end",()=>console.log(JSON.parse(d).dataRoot))')"
|
||||
RESULT_FILE="$DATA_ROOT/runs/$RUN_ID/result.json"
|
||||
[ -f "$RESULT_FILE" ] || fail "run not found: $RUN_ID" 4
|
||||
|
||||
node -e '
|
||||
const r = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
|
||||
process.exit(r.status === "succeeded" ? 0 : 1);
|
||||
' "$RESULT_FILE" || fail "run $RUN_ID did not succeed; refusing to auto-apply"
|
||||
|
||||
WORKSPACE_NAME="$(node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.workspace||"")' "$RESULT_FILE")"
|
||||
[ -n "$WORKSPACE_NAME" ] || fail "run has no workspace; nothing to integrate" 4
|
||||
WORKSPACE="$DATA_ROOT/workspaces/$WORKSPACE_NAME"
|
||||
[ -d "$WORKSPACE/.git" ] || fail "workspace is not a git clone: $WORKSPACE" 4
|
||||
|
||||
# ---- extract diff (tracked + intent-to-add) ----
|
||||
git -C "$WORKSPACE" add -N . >/dev/null 2>&1 || true
|
||||
DIFF_FILE="$(mktemp)"
|
||||
trap 'rm -f "$DIFF_FILE"' EXIT
|
||||
git -C "$WORKSPACE" diff > "$DIFF_FILE"
|
||||
if [ ! -s "$DIFF_FILE" ]; then
|
||||
fail "workspace has no changes to apply"
|
||||
fi
|
||||
|
||||
# ---- allowlist ----
|
||||
mapfile -t CHANGED < <(git -C "$WORKSPACE" diff --name-only)
|
||||
GLOBS="$(node -e 'const a=JSON.parse(process.argv[1]).allowedPaths;console.log(a.join("\n"))' "$POLICY_JSON")"
|
||||
REFUSED=""
|
||||
for f in "${CHANGED[@]}"; do
|
||||
ok="no"
|
||||
while IFS= read -r g; do
|
||||
[ -z "$g" ] && continue
|
||||
case "$f" in
|
||||
$g) ok="yes"; break ;;
|
||||
esac
|
||||
done <<< "$GLOBS"
|
||||
[ "$ok" = "yes" ] || REFUSED="$REFUSED $f"
|
||||
done
|
||||
if [ -n "$REFUSED" ]; then
|
||||
echo "conductor-apply: refusing - files outside policy allowlist:$REFUSED" >&2
|
||||
echo "conductor-apply: patch preserved at $DIFF_FILE for manual review" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ---- syntax gates (on workspace files, pre-apply) ----
|
||||
for f in "${CHANGED[@]}"; do
|
||||
case "$f" in
|
||||
*.mjs) node --check "$WORKSPACE/$f" || fail "syntax gate failed (node): $f" 1 ;;
|
||||
*.sh) bash -n "$WORKSPACE/$f" || fail "syntax gate failed (bash): $f" 1 ;;
|
||||
*.json) node -e 'JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"))' "$WORKSPACE/$f" || fail "syntax gate failed (json): $f" 1 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [ "$DRY_RUN" = "yes" ]; then
|
||||
echo "conductor-apply (dry-run): would apply $(echo "${#CHANGED[@]}") file(s) from $RUN_ID:"
|
||||
printf ' %s\n' "${CHANGED[@]}"
|
||||
echo "conductor-apply (dry-run): suites would run: $(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(", "))' "$POLICY_JSON")"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ---- apply ----
|
||||
[ -z "$(git status --porcelain)" ] || fail "target tree is not clean; refusing to mix states"
|
||||
git apply "$DIFF_FILE" || fail "git apply failed"
|
||||
|
||||
# ---- policy suites ----
|
||||
SUITES="$(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(" "))' "$POLICY_JSON")"
|
||||
SUITES_OK="yes"
|
||||
for s in $SUITES; do
|
||||
case "$s" in
|
||||
test-[a-z]*) : ;; # shape guard; existence checked next
|
||||
*) echo "conductor-apply: refusing suspicious suite name: $s" >&2; SUITES_OK="no"; break ;;
|
||||
esac
|
||||
[ -x "scripts/$s.sh" ] || { echo "conductor-apply: suite script missing: scripts/$s.sh" >&2; SUITES_OK="no"; break; }
|
||||
if ! bash "scripts/$s.sh" >/dev/null 2>&1; then
|
||||
echo "conductor-apply: suite failed: $s" >&2
|
||||
SUITES_OK="no"
|
||||
break
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$SUITES_OK" != "yes" ]; then
|
||||
git apply -R "$DIFF_FILE" && echo "conductor-apply: changes REVERTED (suites failed)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ---- commit with attribution ----
|
||||
git add -A
|
||||
git commit -q -m "feat(worker): auto-applied patch from run $RUN_ID
|
||||
|
||||
Authored-by: pi worker (run $RUN_ID, workspace $WORKSPACE_NAME)
|
||||
Applied-under: conductor-policy v1 (allowlist + syntax gates + suites)"
|
||||
echo "conductor-apply: applied and committed run $RUN_ID ($(echo "${#CHANGED[@]}") file(s)); suites: $SUITES"
|
||||
echo "conductor-apply: NOT pushed - push remains an explicit act."
|
||||
@@ -11,6 +11,7 @@ cd "$(dirname "$0")/.."
|
||||
source scripts/common.sh
|
||||
|
||||
load_config
|
||||
load_release
|
||||
|
||||
bootstrap_runtime_dir
|
||||
|
||||
|
||||
@@ -30,6 +30,7 @@ import process from "node:process";
|
||||
const SUPPORTED_CONFIG_VERSION = 1;
|
||||
const SUPPORTED_ENVIRONMENTS = new Set(["development", "production"]);
|
||||
const SUPPORTED_BACKENDS = new Set(["docker"]);
|
||||
const SUPPORTED_ADAPTERS = new Set(["pi", "mock"]); // mock: test-only, see adapters/README.md
|
||||
const NAME_PATTERN = /^[A-Za-z0-9][A-Za-z0-9._:/-]{0,199}$/;
|
||||
|
||||
function fail(exitCode, message) {
|
||||
@@ -127,7 +128,7 @@ function validate(document, file) {
|
||||
if (!isPlainObject(document.execution)) {
|
||||
fail(2, '"execution" must be a JSON object');
|
||||
}
|
||||
rejectUnknownKeys(document.execution, ["backend", "provider", "model"], '"execution"');
|
||||
rejectUnknownKeys(document.execution, ["backend", "provider", "model", "adapter"], '"execution"');
|
||||
if (!SUPPORTED_BACKENDS.has(document.execution.backend)) {
|
||||
fail(2, `unsupported execution.backend: ${JSON.stringify(document.execution.backend)} (supported: ${[...SUPPORTED_BACKENDS].join(", ")})`);
|
||||
}
|
||||
@@ -138,6 +139,13 @@ function validate(document, file) {
|
||||
}
|
||||
}
|
||||
|
||||
const adapter = document.execution.adapter === undefined || document.execution.adapter === null
|
||||
? "pi"
|
||||
: document.execution.adapter;
|
||||
if (typeof adapter !== "string" || !SUPPORTED_ADAPTERS.has(adapter)) {
|
||||
fail(2, `unsupported execution.adapter: ${JSON.stringify(adapter)} (supported: ${[...SUPPORTED_ADAPTERS].join(", ")})`);
|
||||
}
|
||||
|
||||
return {
|
||||
configVersion: document.configVersion,
|
||||
environment: document.environment,
|
||||
@@ -146,6 +154,7 @@ function validate(document, file) {
|
||||
backend: document.execution.backend,
|
||||
provider: document.execution.provider,
|
||||
model: document.execution.model,
|
||||
adapter,
|
||||
},
|
||||
};
|
||||
}
|
||||
@@ -222,6 +231,7 @@ switch (operation) {
|
||||
`MOSAIC_DATA_ROOT=${shellQuote(resolved.dataRoot)}`,
|
||||
`MOSAIC_PROVIDER=${shellQuote(resolved.execution.provider)}`,
|
||||
`MOSAIC_MODEL=${shellQuote(resolved.execution.model)}`,
|
||||
`MOSAIC_ADAPTER=${shellQuote(resolved.execution.adapter)}`,
|
||||
"",
|
||||
].join("\n"),
|
||||
);
|
||||
|
||||
+338
-7
@@ -28,6 +28,7 @@
|
||||
*/
|
||||
|
||||
import fs from "node:fs";
|
||||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import process from "node:process";
|
||||
import { randomBytes } from "node:crypto";
|
||||
@@ -38,6 +39,7 @@ const PROJECT_ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)),
|
||||
const RUNS_DIRNAME = "runs";
|
||||
const ID_PATTERN = /^[a-z0-9][a-z0-9._-]{0,63}$/;
|
||||
const DEFAULT_TIMEOUT_SECONDS = 120;
|
||||
const SUPPORTED_TOOLS = ["read", "write", "edit", "bash", "grep", "find", "ls"]; // pi documented built-ins
|
||||
|
||||
function fail(exitCode, message) {
|
||||
process.stderr.write(`mosaic-task: ${message}\n`);
|
||||
@@ -82,7 +84,7 @@ function validateId(value, what) {
|
||||
|
||||
function validateMission(document, file) {
|
||||
if (!isPlainObject(document)) fail(2, "mission must be a JSON object");
|
||||
rejectUnknownKeys(document, ["missionVersion", "id", "objective", "directives"], "mission");
|
||||
rejectUnknownKeys(document, ["missionVersion", "id", "objective", "directives", "capabilities"], "mission");
|
||||
if (document.missionVersion !== 1) {
|
||||
fail(2, `unsupported missionVersion: ${JSON.stringify(document.missionVersion)} (supported: 1)`);
|
||||
}
|
||||
@@ -100,12 +102,34 @@ function validateMission(document, file) {
|
||||
return d;
|
||||
});
|
||||
}
|
||||
return { missionVersion: document.missionVersion, id: document.id, objective: document.objective, directives };
|
||||
|
||||
// Governing capability constraints (M9): same validation as task
|
||||
// capabilities; semantically these BOUND tasks (least-privilege
|
||||
// intersection at run time), never grant beyond them.
|
||||
let capabilities = null;
|
||||
if (document.capabilities !== undefined && document.capabilities !== null) {
|
||||
if (!isPlainObject(document.capabilities)) fail(2, 'mission "capabilities" must be a JSON object');
|
||||
rejectUnknownKeys(document.capabilities, ["tools"], 'mission "capabilities"');
|
||||
if (!Array.isArray(document.capabilities.tools) || document.capabilities.tools.length === 0) {
|
||||
fail(2, 'mission "capabilities.tools" must be a non-empty array of tool names');
|
||||
}
|
||||
const seen = new Set();
|
||||
for (const tool of document.capabilities.tools) {
|
||||
if (!SUPPORTED_TOOLS.includes(tool)) {
|
||||
fail(2, `unsupported mission tool: ${JSON.stringify(tool)} (supported: ${SUPPORTED_TOOLS.join(", ")})`);
|
||||
}
|
||||
if (seen.has(tool)) fail(2, `duplicate tool in mission capabilities.tools: ${tool}`);
|
||||
seen.add(tool);
|
||||
}
|
||||
capabilities = { tools: [...seen] };
|
||||
}
|
||||
|
||||
return { missionVersion: document.missionVersion, id: document.id, objective: document.objective, directives, capabilities };
|
||||
}
|
||||
|
||||
function validateTask(document, file) {
|
||||
if (!isPlainObject(document)) fail(2, "task must be a JSON object");
|
||||
rejectUnknownKeys(document, ["taskVersion", "id", "prompt", "mission", "expectExact", "timeoutSeconds"], "task");
|
||||
rejectUnknownKeys(document, ["taskVersion", "id", "prompt", "mission", "expectExact", "timeoutSeconds", "workspace", "capabilities", "session", "sessionForkFrom"], "task");
|
||||
if (document.taskVersion !== 1) {
|
||||
fail(2, `unsupported taskVersion: ${JSON.stringify(document.taskVersion)} (supported: 1)`);
|
||||
}
|
||||
@@ -144,6 +168,67 @@ function validateTask(document, file) {
|
||||
timeoutSeconds = document.timeoutSeconds;
|
||||
}
|
||||
|
||||
// Workspace (M5): absent = none; ":run" = ephemeral per-run; otherwise a
|
||||
// persistent named workspace under <dataRoot>/workspaces/<name>.
|
||||
let workspace = null;
|
||||
if (document.workspace !== undefined && document.workspace !== null) {
|
||||
if (typeof document.workspace !== "string" || document.workspace.length === 0) {
|
||||
fail(2, 'task "workspace" must be a non-empty string when present');
|
||||
}
|
||||
if (document.workspace !== ":run") {
|
||||
validateId(document.workspace, "task workspace");
|
||||
}
|
||||
workspace = document.workspace;
|
||||
}
|
||||
|
||||
// Session (M6): optional named persistent session under
|
||||
// <dataRoot>/sessions/<name>. Distinct names never share state.
|
||||
let session = null;
|
||||
if (document.session !== undefined && document.session !== null) {
|
||||
if (typeof document.session !== "string" || document.session.length === 0) {
|
||||
fail(2, 'task "session" must be a non-empty string when present');
|
||||
}
|
||||
validateId(document.session, "task session");
|
||||
session = document.session;
|
||||
}
|
||||
|
||||
// Session fork (M11): optional source session whose newest session file
|
||||
// is branched (pi --fork) into the target session dir. Requires session.
|
||||
let sessionForkFrom = null;
|
||||
if (document.sessionForkFrom !== undefined && document.sessionForkFrom !== null) {
|
||||
if (typeof document.sessionForkFrom !== "string" || document.sessionForkFrom.length === 0) {
|
||||
fail(2, 'task "sessionForkFrom" must be a non-empty string when present');
|
||||
}
|
||||
validateId(document.sessionForkFrom, "task sessionForkFrom");
|
||||
if (!session) {
|
||||
fail(2, 'task "sessionForkFrom" requires "session" (the fork target) to be set');
|
||||
}
|
||||
if (document.sessionForkFrom === session) {
|
||||
fail(2, 'task "sessionForkFrom" must differ from "session" (cannot fork onto itself)');
|
||||
}
|
||||
sessionForkFrom = document.sessionForkFrom;
|
||||
}
|
||||
|
||||
// Capabilities (M5): optional tools allowlist mapped by adapters to their
|
||||
// native permission flags. Absent = no tools.
|
||||
let tools = null;
|
||||
if (document.capabilities !== undefined && document.capabilities !== null) {
|
||||
if (!isPlainObject(document.capabilities)) fail(2, '"capabilities" must be a JSON object');
|
||||
rejectUnknownKeys(document.capabilities, ["tools"], '"capabilities"');
|
||||
if (!Array.isArray(document.capabilities.tools) || document.capabilities.tools.length === 0) {
|
||||
fail(2, '"capabilities.tools" must be a non-empty array of tool names');
|
||||
}
|
||||
const seen = new Set();
|
||||
for (const tool of document.capabilities.tools) {
|
||||
if (!SUPPORTED_TOOLS.includes(tool)) {
|
||||
fail(2, `unsupported tool: ${JSON.stringify(tool)} (supported: ${SUPPORTED_TOOLS.join(", ")})`);
|
||||
}
|
||||
if (seen.has(tool)) fail(2, `duplicate tool in capabilities.tools: ${tool}`);
|
||||
seen.add(tool);
|
||||
}
|
||||
tools = [...seen];
|
||||
}
|
||||
|
||||
return {
|
||||
taskVersion: document.taskVersion,
|
||||
id: document.id,
|
||||
@@ -153,6 +238,10 @@ function validateTask(document, file) {
|
||||
missionSnapshot,
|
||||
expectExact,
|
||||
timeoutSeconds,
|
||||
workspace,
|
||||
tools,
|
||||
session,
|
||||
sessionForkFrom,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -185,7 +274,7 @@ function writeOnce(file, content) {
|
||||
}
|
||||
}
|
||||
|
||||
function runTask(taskFile) {
|
||||
function runTask(taskFile, options = {}) {
|
||||
const resolved = JSON.parse(
|
||||
spawnSync(process.execPath, [path.join(PROJECT_ROOT, "scripts", "mosaic-config.mjs"), "validate"], {
|
||||
cwd: PROJECT_ROOT,
|
||||
@@ -210,12 +299,89 @@ function runTask(taskFile) {
|
||||
|
||||
const startedAt = new Date();
|
||||
const stderrFile = path.join(runDir, "stderr.txt");
|
||||
|
||||
// Sanctioned mission injection: point the container at the run snapshot's
|
||||
// CONTAINER path (dataRoot maps to /var/lib/mosaic in the image).
|
||||
const spawnEnv = { ...process.env };
|
||||
spawnEnv.MOSAIC_ADAPTER = resolved.execution.adapter;
|
||||
// Self-sufficient env: direct invocation (e.g. `retry`) skips the shell
|
||||
// launcher exports, so derive them from the resolved config and release.
|
||||
// (Names here are the compose interpolation consumers, not PI_*.)
|
||||
spawnEnv.MOSAIC_PROVIDER = resolved.execution.provider;
|
||||
spawnEnv.MOSAIC_MODEL = resolved.execution.model;
|
||||
spawnEnv.MOSAIC_DATA_ROOT = resolved.dataRoot;
|
||||
const release = fs.readFileSync(path.join(PROJECT_ROOT, "RELEASE"), "utf8").trim();
|
||||
const piVersion = JSON.parse(fs.readFileSync(path.join(PROJECT_ROOT, "package.json"), "utf8")).dependencies["@earendil-works/pi-coding-agent"];
|
||||
spawnEnv.MOSAIC_IMAGE_TAG = `mosaic-poc-agent:${piVersion}-r${release}`;
|
||||
if (task.missionSnapshot) {
|
||||
const relative = path.relative(resolved.dataRoot, runDir);
|
||||
if (relative.startsWith("..") || path.isAbsolute(relative)) {
|
||||
fail(4, `run directory is outside the configured dataRoot: ${runDir}`);
|
||||
}
|
||||
spawnEnv.MOSAIC_MISSION_FILE = `/var/lib/mosaic/${relative.split(path.sep).join("/")}/mission.json`;
|
||||
}
|
||||
|
||||
// Workspace (M5): create host-side, pass the CONTAINER path.
|
||||
let workspaceContainerPath = null;
|
||||
if (task.workspace === ":run") {
|
||||
fs.mkdirSync(path.join(runDir, "workspace"), { recursive: true });
|
||||
workspaceContainerPath = `/var/lib/mosaic/runs/${runId}/workspace`;
|
||||
} else if (task.workspace) {
|
||||
fs.mkdirSync(path.join(resolved.dataRoot, "workspaces", task.workspace), { recursive: true });
|
||||
workspaceContainerPath = `/var/lib/mosaic/workspaces/${task.workspace}`;
|
||||
}
|
||||
if (workspaceContainerPath) spawnEnv.MOSAIC_WORKSPACE = workspaceContainerPath;
|
||||
|
||||
// Capability policy (M9): least-privilege intersection. A task may narrow
|
||||
// a mission's tool grant, never widen it. Empty intersection = tool-free.
|
||||
let effectiveTools = task.tools;
|
||||
let policyNote = null;
|
||||
if (task.missionSnapshot?.capabilities) {
|
||||
const missionTools = task.missionSnapshot.capabilities.tools;
|
||||
if (effectiveTools) {
|
||||
effectiveTools = effectiveTools.filter((t) => missionTools.includes(t));
|
||||
if (effectiveTools.length === 0) {
|
||||
policyNote = `capability policy: mission ${task.missionSnapshot.id} and task request no tools in common -> tool-free run`;
|
||||
}
|
||||
} else {
|
||||
effectiveTools = [...missionTools];
|
||||
}
|
||||
}
|
||||
if (policyNote) process.stderr.write(`mosaic-task: ${policyNote}\n`);
|
||||
spawnEnv.MOSAIC_TOOLS = effectiveTools ? effectiveTools.join(",") : "";
|
||||
|
||||
// Session (M6): persistent named session dir, passed as container path.
|
||||
if (task.session) {
|
||||
fs.mkdirSync(path.join(resolved.dataRoot, "sessions", task.session), { recursive: true });
|
||||
spawnEnv.MOSAIC_SESSION_DIR = `/var/lib/mosaic/sessions/${task.session}`;
|
||||
}
|
||||
|
||||
// Session fork (M11): resolve the source session's newest file; pi --fork
|
||||
// branches it into the target dir without modifying the ancestor.
|
||||
if (task.sessionForkFrom) {
|
||||
const sourceDir = path.join(resolved.dataRoot, "sessions", task.sessionForkFrom);
|
||||
let sources = [];
|
||||
try {
|
||||
sources = fs.readdirSync(sourceDir).filter((f) => f.endsWith(".jsonl")).sort();
|
||||
} catch {
|
||||
fail(4, `cannot fork: source session dir not found: ${sourceDir}`);
|
||||
}
|
||||
if (sources.length === 0) {
|
||||
fail(4, `cannot fork: source session '${task.sessionForkFrom}' has no session files`);
|
||||
}
|
||||
const relative = path.relative(resolved.dataRoot, path.join(sourceDir, sources[sources.length - 1]));
|
||||
if (relative.startsWith("..") || path.isAbsolute(relative)) {
|
||||
fail(4, `source session is outside the configured dataRoot: ${sourceDir}`);
|
||||
}
|
||||
spawnEnv.MOSAIC_SESSION_FORK = `/var/lib/mosaic/${relative.split(path.sep).join("/")}`;
|
||||
}
|
||||
|
||||
const proc = spawnSync(
|
||||
"docker",
|
||||
["compose", "run", "--rm", "-T", "mosaic-agent", task.prompt],
|
||||
{
|
||||
cwd: PROJECT_ROOT,
|
||||
env: process.env, // MOSAIC_DATA_ROOT / MOSAIC_PROVIDER / MOSAIC_MODEL resolved by run-task.sh
|
||||
env: spawnEnv,
|
||||
input: "", // stdin detached: print mode must never wait on a terminal (see issue #5)
|
||||
encoding: "utf8",
|
||||
maxBuffer: 16 * 1024 * 1024,
|
||||
@@ -256,6 +422,11 @@ function runTask(taskFile) {
|
||||
request: task.prompt,
|
||||
response,
|
||||
expectedExact: expected,
|
||||
...(options.retriedFrom ? { retriedFrom: options.retriedFrom } : {}),
|
||||
workspace: task.workspace,
|
||||
tools: effectiveTools,
|
||||
session: task.session,
|
||||
sessionForkFrom: task.sessionForkFrom,
|
||||
exitCode: proc.status,
|
||||
signal: proc.signal ?? null,
|
||||
provider: resolved.execution.provider,
|
||||
@@ -283,17 +454,166 @@ function listRuns() {
|
||||
for (const runId of entries) {
|
||||
let status = "unknown";
|
||||
let taskId = "-";
|
||||
let workspace = "-";
|
||||
let session = "-";
|
||||
try {
|
||||
const result = JSON.parse(fs.readFileSync(path.join(root, runId, "result.json"), "utf8"));
|
||||
status = result.status;
|
||||
taskId = result.taskId;
|
||||
workspace = result.workspace ?? "-";
|
||||
session = result.session ?? "-";
|
||||
} catch {
|
||||
// Incomplete run record; report as unknown.
|
||||
}
|
||||
process.stdout.write(`${runId} ${status.padEnd(9)} ${taskId}\n`);
|
||||
process.stdout.write(`${runId} ${status.padEnd(9)} task=${taskId.padEnd(18)} ws=${String(workspace).padEnd(10)} session=${session}\n`);
|
||||
}
|
||||
}
|
||||
|
||||
function showRun(runId) {
|
||||
if (!/^r-[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(runId)) {
|
||||
fail(4, `invalid run id: ${JSON.stringify(runId)} (expected r-<id>)`);
|
||||
}
|
||||
const resolved = loadConfig();
|
||||
const dir = path.join(runsRoot(resolved), runId);
|
||||
if (!fs.existsSync(dir)) {
|
||||
fail(4, `run not found: ${runId} (under ${runsRoot(resolved)})`);
|
||||
}
|
||||
|
||||
const read = (name) => {
|
||||
try {
|
||||
return JSON.parse(fs.readFileSync(path.join(dir, name), "utf8"));
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
};
|
||||
const result = read("result.json");
|
||||
const task = read("task.json");
|
||||
const mission = read("mission.json");
|
||||
|
||||
process.stdout.write(`run: ${runId}\n`);
|
||||
if (result) {
|
||||
process.stdout.write(
|
||||
[
|
||||
`status: ${result.status}${result.reason ? ` (${result.reason})` : ""}`,
|
||||
`task: ${result.taskId}`,
|
||||
result.missionId ? `mission: ${result.missionId}` : null,
|
||||
result.workspace ? `workspace: ${result.workspace}` : null,
|
||||
result.session ? `session: ${result.session}` : null,
|
||||
result.tools ? `tools: ${result.tools.join(", ")}` : null,
|
||||
`adapter: ${"(see config)"} provider=${result.provider} model=${result.model}`,
|
||||
`request: ${JSON.stringify(result.request)}`,
|
||||
`response: ${JSON.stringify(result.response)}`,
|
||||
result.expectedExact !== null && result.expectedExact !== undefined ? `expected: ${JSON.stringify(result.expectedExact)}` : null,
|
||||
result.retriedFrom ? `retriedFrom: ${result.retriedFrom}` : null,
|
||||
`timing: ${result.startedAt} -> ${result.finishedAt} (${result.durationMs} ms)`,
|
||||
`exit: ${result.exitCode}${result.signal ? ` signal=${result.signal}` : ""}`,
|
||||
].filter((line) => line !== null).join("\n") + "\n",
|
||||
);
|
||||
} else {
|
||||
process.stdout.write("result.json: (missing or unreadable)\n");
|
||||
}
|
||||
if (task) process.stdout.write(`task snapshot: ${"task.json"} present\n`);
|
||||
if (mission) process.stdout.write(`mission snapshot: ${mission.id} - ${mission.objective}\n`);
|
||||
process.stdout.write(`artifacts: ${fs.readdirSync(dir).map((f) => `${f}`).join(", ")}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
function retryRun(runId) {
|
||||
if (!/^r-[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(runId)) {
|
||||
fail(4, `invalid run id: ${JSON.stringify(runId)} (expected r-<id>)`);
|
||||
}
|
||||
const resolved = loadConfig();
|
||||
const dir = path.join(runsRoot(resolved), runId);
|
||||
if (!fs.existsSync(dir)) {
|
||||
fail(4, `run not found: ${runId}`);
|
||||
}
|
||||
|
||||
let snapshot;
|
||||
try {
|
||||
snapshot = fs.readFileSync(path.join(dir, "task.json"), "utf8");
|
||||
} catch {
|
||||
fail(4, `run task snapshot unreadable: ${runId}`);
|
||||
}
|
||||
|
||||
// Relative mission paths in a snapshot resolve against the ORIGINAL task
|
||||
// location, which no longer exists here — rewrite them to the run's own
|
||||
// recorded mission.json so retries stay faithful.
|
||||
let snapshotDoc;
|
||||
try {
|
||||
snapshotDoc = JSON.parse(snapshot);
|
||||
} catch {
|
||||
fail(4, `run task snapshot is not valid JSON: ${runId}`);
|
||||
}
|
||||
if (snapshotDoc.mission && !path.isAbsolute(snapshotDoc.mission)) {
|
||||
const recordedMission = path.join(dir, "mission.json");
|
||||
if (!fs.existsSync(recordedMission)) {
|
||||
fail(4, `cannot retry ${runId}: relative mission path but no mission.json snapshot in run dir`);
|
||||
}
|
||||
snapshotDoc.mission = recordedMission;
|
||||
snapshot = `${JSON.stringify(snapshotDoc, null, 2)}\n`;
|
||||
}
|
||||
|
||||
// A retry is a brand-new run: replay the recorded task snapshot through
|
||||
// the ordinary run path; existing run records stay untouched.
|
||||
const tempDir = fs.mkdtempSync(path.join(os.tmpdir(), "mosaic-retry-"));
|
||||
const tempTaskFile = path.join(tempDir, "task.json");
|
||||
fs.writeFileSync(tempTaskFile, snapshot);
|
||||
process.on("exit", () => {
|
||||
try {
|
||||
fs.rmSync(tempDir, { recursive: true, force: true });
|
||||
} catch {
|
||||
// Best-effort cleanup only.
|
||||
}
|
||||
});
|
||||
runTask(tempTaskFile, { retriedFrom: runId });
|
||||
}
|
||||
|
||||
function pruneRuns(args) {
|
||||
const resolved = loadConfig();
|
||||
const root = runsRoot(resolved);
|
||||
let keep = 50;
|
||||
let apply = false;
|
||||
for (const arg of args) {
|
||||
if (arg === "--yes") apply = true;
|
||||
else if (arg === "--keep") fail(4, "prune: --keep requires a value");
|
||||
else if (arg.startsWith("--keep=")) {
|
||||
keep = Number(arg.slice("--keep=".length));
|
||||
if (!Number.isInteger(keep) || keep < 1) fail(4, `--keep must be a positive integer (got ${arg.slice(7)})`);
|
||||
} else fail(4, `unknown prune option: ${arg}`);
|
||||
}
|
||||
|
||||
let entries = [];
|
||||
try {
|
||||
entries = fs.readdirSync(root, { withFileTypes: true })
|
||||
.filter((e) => e.name.startsWith("r-") && e.isDirectory() && !e.isSymbolicLink())
|
||||
.map((e) => e.name)
|
||||
.sort();
|
||||
} catch {
|
||||
// No runs yet.
|
||||
}
|
||||
|
||||
if (entries.length <= keep) {
|
||||
process.stdout.write(`prune: ${entries.length} run(s) present, keep=${keep} -> nothing to prune\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const doomed = entries.slice(0, entries.length - keep); // oldest first
|
||||
if (!apply) {
|
||||
process.stdout.write(`prune (dry-run): would remove ${doomed.length} oldest run(s), keep ${entries.length - doomed.length}:\n`);
|
||||
for (const id of doomed) process.stdout.write(` would remove: ${id}\n`);
|
||||
process.stdout.write("prune: re-run with --yes to apply\n");
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const receipt = path.join(root, ".pruned.log");
|
||||
for (const id of doomed) {
|
||||
fs.rmSync(path.join(root, id), { recursive: true, force: true });
|
||||
fs.appendFileSync(receipt, `${JSON.stringify({ at: new Date().toISOString(), event: "pruned", runId: id })}\n`);
|
||||
}
|
||||
process.stdout.write(`prune: removed ${doomed.length} run(s), kept ${keep}; receipt: ${receipt}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const operation = process.argv[2];
|
||||
const target = process.argv[3];
|
||||
|
||||
@@ -310,9 +630,20 @@ switch (operation) {
|
||||
if (!target) fail(4, "usage: mosaic-task.mjs run <taskFile>");
|
||||
runTask(path.resolve(target));
|
||||
break;
|
||||
case "show":
|
||||
if (!target) fail(4, "usage: mosaic-task.mjs show <runId>");
|
||||
showRun(target);
|
||||
break;
|
||||
case "list":
|
||||
listRuns();
|
||||
process.exit(0);
|
||||
case "prune":
|
||||
pruneRuns(process.argv.slice(3));
|
||||
break;
|
||||
case "retry":
|
||||
if (!target) fail(4, "usage: mosaic-task.mjs retry <runId>");
|
||||
retryRun(target);
|
||||
break;
|
||||
default:
|
||||
fail(4, `unknown operation: ${JSON.stringify(operation ?? "")} (expected validate | run | list)`);
|
||||
fail(4, `unknown operation: ${JSON.stringify(operation ?? "")} (expected validate | run | show | list | retry | prune)`);
|
||||
}
|
||||
|
||||
Executable
+149
@@ -0,0 +1,149 @@
|
||||
#!/usr/bin/env bash
|
||||
# Release lifecycle: package, activate, rollback, status.
|
||||
#
|
||||
# scripts/release.sh package build the image for this release
|
||||
# scripts/release.sh activate [--fault-injection]
|
||||
# scripts/release.sh rollback
|
||||
# scripts/release.sh status
|
||||
#
|
||||
# Activation is health-gated: the M2 task runner executes
|
||||
# tasks/hello-marker.json; only an exact-marker pass activates. The
|
||||
# pointer (state/active.json) is replaced atomically; every attempt is
|
||||
# appended to state/activation-log.jsonl (append-only history).
|
||||
#
|
||||
# Fault injection exists solely to prove the refusal path in drills.
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
# shellcheck source=common.sh
|
||||
source scripts/common.sh
|
||||
|
||||
load_config
|
||||
load_release
|
||||
bootstrap_runtime_dir
|
||||
|
||||
STATE_DIR="$MOSAIC_DEV_DIR/state"
|
||||
mkdir -p "$STATE_DIR"
|
||||
POINTER="$STATE_DIR/active.json"
|
||||
LOG="$STATE_DIR/activation-log.jsonl"
|
||||
|
||||
now() { date -u +%Y-%m-%dT%H:%M:%SZ; }
|
||||
|
||||
append_log() { # event release imageTag note
|
||||
local event="$1" release="$2" imageTag="$3" note="${4:-}"
|
||||
printf '{"at":"%s","event":"%s","release":"%s","imageTag":"%s"%s}\n' \
|
||||
"$(now)" "$event" "$release" "$imageTag" \
|
||||
"$(printf '%s' "$note" | node -e 'let d="";process.stdin.on("data",c=>d+=c).on("end",()=>{const s=d.replace(/\n$/,"");process.stdout.write(s ? ",\"note\":"+JSON.stringify(s) : "")})')" \
|
||||
>> "$LOG"
|
||||
}
|
||||
|
||||
image_exists() { docker image inspect "$1" >/dev/null 2>&1; }
|
||||
|
||||
health_check() { # returns 0 only when the marker path passes; $1 = fault injection label or empty
|
||||
local tmp=""
|
||||
local task="tasks/hello-marker.json"
|
||||
if [ -n "${1:-}" ]; then
|
||||
tmp="$(mktemp -d)"
|
||||
# Fault injection: same prompt, deliberately wrong expectation.
|
||||
printf '{"taskVersion":1,"id":"t-health-fault","prompt":"Return your startup marker and nothing else.","expectExact":"MOSAIC_FAULT_%s"}' \
|
||||
"$RANDOM$RANDOM" > "$tmp/fault-task.json"
|
||||
task="$tmp/fault-task.json"
|
||||
fi
|
||||
local rc=0
|
||||
scripts/run-task.sh run "$task" >/dev/null 2>&1 || rc=$?
|
||||
[ -n "$tmp" ] && rm -rf "$tmp"
|
||||
return "$rc"
|
||||
}
|
||||
|
||||
activate() { # $1 = release, $2 = imageTag, $3 = event name, $4 = fault label
|
||||
local release="$1" imageTag="$2" event="$3" fault="${4:-}"
|
||||
|
||||
if ! image_exists "$imageTag"; then
|
||||
echo "release: refusing $event: image not present locally: $imageTag" >&2
|
||||
append_log "refused" "$release" "$imageTag" "image missing"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if ! health_check "$fault"; then
|
||||
echo "release: refusing $event: health check failed" >&2
|
||||
append_log "refused" "$release" "$imageTag" "health check failed${fault:+ (fault-injected)}"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Atomic pointer replacement: write sibling temp file, then rename.
|
||||
local tmp_pointer="$POINTER.tmp.$$"
|
||||
printf '{"pointerVersion":1,"release":"%s","imageTag":"%s","activatedAt":"%s"}\n' \
|
||||
"$release" "$imageTag" "$(now)" > "$tmp_pointer"
|
||||
mv -f "$tmp_pointer" "$POINTER"
|
||||
|
||||
append_log "$event" "$release" "$imageTag"
|
||||
echo "release: $event OK -> $release ($imageTag)"
|
||||
}
|
||||
|
||||
previous_image_tag() { # last activated imageTag different from current pointer
|
||||
[ -f "$POINTER" ] || return 1
|
||||
local current
|
||||
current="$(node -p 'JSON.parse(require("fs").readFileSync(process.argv[1],"utf8")).imageTag' "$POINTER")"
|
||||
node -e '
|
||||
const fs = require("fs");
|
||||
const current = process.argv[1];
|
||||
const lines = fs.readFileSync(process.argv[2], "utf8").split("\n").filter(Boolean);
|
||||
for (let i = lines.length - 1; i >= 0; i--) {
|
||||
let e;
|
||||
try { e = JSON.parse(lines[i]); } catch { continue; }
|
||||
if ((e.event === "activate" || e.event === "rollback") && e.imageTag && e.imageTag !== current) {
|
||||
console.log(e.imageTag);
|
||||
process.exit(0);
|
||||
}
|
||||
}
|
||||
process.exit(1);
|
||||
' "$current" "$LOG"
|
||||
}
|
||||
|
||||
cmd_status() {
|
||||
echo "release: $MOSAIC_RELEASE"
|
||||
echo "image tag: $MOSAIC_IMAGE_TAG (packaged: $(image_exists "$MOSAIC_IMAGE_TAG" && echo yes || echo no))"
|
||||
if [ -f "$POINTER" ]; then
|
||||
node -e '
|
||||
const p = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
|
||||
console.log("active: " + p.release + " (" + p.imageTag + ") since " + p.activatedAt);
|
||||
' "$POINTER"
|
||||
else
|
||||
echo "active: (none)"
|
||||
fi
|
||||
if [ -f "$LOG" ]; then
|
||||
echo "recent log:"
|
||||
tail -5 "$LOG" | sed 's/^/ /'
|
||||
fi
|
||||
}
|
||||
|
||||
cmd_package() {
|
||||
docker compose build
|
||||
append_log "package" "$MOSAIC_RELEASE" "$MOSAIC_IMAGE_TAG"
|
||||
echo "release: packaged $MOSAIC_IMAGE_TAG"
|
||||
}
|
||||
|
||||
cmd_activate() {
|
||||
local fault=""
|
||||
if [ "${1:-}" = "--fault-injection" ]; then fault="yes"; fi
|
||||
activate "$MOSAIC_RELEASE" "$MOSAIC_IMAGE_TAG" "activate" "$fault"
|
||||
}
|
||||
|
||||
cmd_rollback() {
|
||||
local prev
|
||||
if ! prev="$(previous_image_tag)"; then
|
||||
echo "release: rollback: no previous activation found in log" >&2
|
||||
exit 1
|
||||
fi
|
||||
local prev_release
|
||||
prev_release="$(printf '%s' "$prev" | sed -n 's/.*-r\([0-9.]*\)$/\1/p')"
|
||||
[ -n "$prev_release" ] || prev_release="unknown"
|
||||
activate "$prev_release" "$prev" "rollback"
|
||||
}
|
||||
|
||||
case "${1:-}" in
|
||||
package) cmd_package ;;
|
||||
activate) shift; cmd_activate "$@" ;;
|
||||
rollback) cmd_rollback ;;
|
||||
status) cmd_status ;;
|
||||
*) echo "usage: scripts/release.sh package | activate [--fault-injection] | rollback | status" >&2; exit 4 ;;
|
||||
esac
|
||||
@@ -11,6 +11,7 @@ cd "$(dirname "$0")/.."
|
||||
source scripts/common.sh
|
||||
|
||||
load_config
|
||||
load_release # compose requires MOSAIC_IMAGE_TAG; task runs are release-scoped too
|
||||
bootstrap_runtime_dir
|
||||
|
||||
exec node scripts/mosaic-task.mjs "$@"
|
||||
|
||||
Executable
+141
@@ -0,0 +1,141 @@
|
||||
#!/usr/bin/env bash
|
||||
# Sandboxed selftests for the conductor auto-apply policy gate.
|
||||
#
|
||||
# Builds a throwaway target repo + worker workspace + fake run records, then
|
||||
# exercises every gate: policy validation, allowlist, syntax gates, suite
|
||||
# failure revert, disabled policy, missing/failed runs. No real model calls.
|
||||
set -uo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
|
||||
SANDBOX="$(mktemp -d)"
|
||||
trap 'rm -rf "$SANDBOX"' EXIT
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
|
||||
check_rc() { # name expectedRc command...
|
||||
local name="$1" expected="$2"
|
||||
shift 2
|
||||
local rc
|
||||
"$@" >/dev/null 2>&1
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
# ---- infrastructure: target repo + worker workspace + fake run ----
|
||||
git clone -q . "$SANDBOX/repo"
|
||||
# The clone carries committed state only - give the target its policy and
|
||||
# commit it so the tree starts clean (untracked policy would fail target_clean).
|
||||
mkdir -p "$SANDBOX/repo/roles"
|
||||
cp roles/conductor-policy.json "$SANDBOX/repo/roles/conductor-policy.json"
|
||||
git -C "$SANDBOX/repo" add roles/conductor-policy.json
|
||||
git -C "$SANDBOX/repo" -c user.name=suite -c user.email=suite@local commit -q -m policy
|
||||
mkdir -p "$SANDBOX/data/workspaces" "$SANDBOX/data/runs"
|
||||
git clone -q "$SANDBOX/repo" "$SANDBOX/data/workspaces/stack-repo"
|
||||
|
||||
TARGET="$SANDBOX/repo"
|
||||
WS="$SANDBOX/data/workspaces/stack-repo"
|
||||
cat > "$SANDBOX/config.json" <<EOF
|
||||
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"mock"}}
|
||||
EOF
|
||||
export MOSAIC_APPLY_TARGET="$TARGET"
|
||||
export MOSAIC_CONFIG="$SANDBOX/config.json"
|
||||
|
||||
RUN_OK="r-20260903T000000000Z-ok0000001"
|
||||
mkdir -p "$SANDBOX/data/runs/$RUN_OK"
|
||||
printf '{"runVersion":1,"runId":"%s","taskId":"t-fake","status":"succeeded","workspace":"stack-repo"}' "$RUN_OK" \
|
||||
> "$SANDBOX/data/runs/$RUN_OK/result.json"
|
||||
|
||||
ws_edit() { printf '\n%s\n' "$2" >> "$WS/$1"; }
|
||||
ws_reset() { git -C "$WS" checkout -q -- . 2>/dev/null; git -C "$WS" clean -qfd; }
|
||||
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
|
||||
|
||||
set_policy() { # enabled suites (commits: the target tree must stay clean)
|
||||
local suites="[\"$2\"]"
|
||||
printf '{"policyVersion":1,"autoApply":{"enabled":%s,"allowedPaths":["scripts/**","docs/**","README.md"],"suites":%s}}' "$1" "$suites" \
|
||||
> "$TARGET/roles/conductor-policy.json"
|
||||
git -C "$TARGET" add roles/conductor-policy.json
|
||||
git -C "$TARGET" -c user.name=suite -c user.email=suite@local commit -q -m "policy update"
|
||||
}
|
||||
set_policy true "test-config"
|
||||
|
||||
# T1: dry run - allowed change, nothing applied
|
||||
ws_edit "README.md" "worker dry-run line"
|
||||
check_rc "dry-run: allowed change, exit 0, nothing committed" 0 \
|
||||
scripts/conductor-apply.sh "$RUN_OK" --dry-run
|
||||
if git -C "$TARGET" log --format=%s | grep -q "auto-applied"; then
|
||||
check "dry-run committed nothing" 1
|
||||
else
|
||||
check "dry-run committed nothing" 0
|
||||
fi
|
||||
ws_reset
|
||||
|
||||
# T2: apply - allowed change, suites pass, commit created
|
||||
ws_edit "README.md" "worker applied line"
|
||||
check_rc "apply: allowed change exits 0" 0 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" log -1 --format=%s | grep -q "auto-applied patch from run $RUN_OK" \
|
||||
&& check "apply: attribution in commit subject" 0 || check "apply: attribution in commit subject" 1
|
||||
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
|
||||
target_clean && check "apply: target tree clean after commit" 0 || check "apply: target tree clean after commit" 1
|
||||
git -C "$TARGET" reset -q --hard HEAD~1
|
||||
|
||||
# T3: disallowed path refused
|
||||
ws_edit "Containerfile" "# worker touch"
|
||||
check_rc "disallowed path refused" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "disallowed path: target untouched" 0 || check "disallowed path: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T4: syntax gate - broken .mjs on an allowed path
|
||||
printf 'this is not (valid js\n' > "$WS/scripts/broken-worker.mjs"
|
||||
check_rc "syntax gate refused broken .mjs" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "syntax gate: target untouched" 0 || check "syntax gate: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T5: suite failure - allowed change breaks a policy suite -> auto-revert
|
||||
printf '\nexit 7\n' >> "$TARGET/scripts/test-config.sh"
|
||||
ws_edit "README.md" "worker change that will fail suites"
|
||||
check_rc "suite failure refused" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" checkout -q -- scripts/test-config.sh
|
||||
target_clean && check "suite failure: target reverted to clean" 0 || check "suite failure: target reverted to clean" 1
|
||||
|
||||
# T6: disabled policy
|
||||
set_policy false "test-config"
|
||||
ws_edit "README.md" "worker line while disabled"
|
||||
check_rc "disabled policy refused" 2 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "disabled policy: target untouched" 0 || check "disabled policy: target untouched" 1
|
||||
ws_reset
|
||||
set_policy true "test-config"
|
||||
|
||||
# T7: failed run refused
|
||||
RUN_FAIL="r-20260903T000000000Z-fail00001"
|
||||
mkdir -p "$SANDBOX/data/runs/$RUN_FAIL"
|
||||
printf '{"runVersion":1,"runId":"%s","taskId":"t","status":"failed","workspace":"stack-repo"}' "$RUN_FAIL" \
|
||||
> "$SANDBOX/data/runs/$RUN_FAIL/result.json"
|
||||
ws_edit "README.md" "worker line from failed run"
|
||||
check_rc "failed run refused" 1 scripts/conductor-apply.sh "$RUN_FAIL"
|
||||
target_clean && check "failed run: target untouched" 0 || check "failed run: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T8/T9: missing run + invalid policy
|
||||
check_rc "missing run exits 4" 4 scripts/conductor-apply.sh r-missing
|
||||
printf '{"policyVersion":9}' > "$TARGET/roles/conductor-policy.json"
|
||||
check_rc "invalid policy exits 2" 2 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" checkout -q -- roles/conductor-policy.json
|
||||
|
||||
echo
|
||||
echo "selftest: $PASS passed, $FAIL failed"
|
||||
[ "$FAIL" -eq 0 ]
|
||||
+34
-10
@@ -10,8 +10,18 @@ SANDBOX="$(mktemp -d)"
|
||||
trap 'rm -rf "$SANDBOX"' EXIT
|
||||
|
||||
PASS=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
FAIL=0
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
# expect_exit NAME EXPECTED_RC -- command...
|
||||
expect_exit() {
|
||||
local name="$1" expected="$2"
|
||||
@@ -21,10 +31,10 @@ expect_exit() {
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS + 1))
|
||||
echo "ok $name (exit $rc)"
|
||||
echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL + 1))
|
||||
echo "FAIL $name (exit $rc, expected $expected)"
|
||||
echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
@@ -40,12 +50,26 @@ cfg() { printf '%s' "$2" > "$SANDBOX/$1"; }
|
||||
|
||||
DATA_ROOT="$SANDBOX/data"
|
||||
|
||||
# --- adapter selection (M4) ---
|
||||
cfg default-adapter.json '{"configVersion":1,"environment":"development","dataRoot":"'$DATA_ROOT'","execution":{"backend":"docker","provider":"zai","model":"m"}}'
|
||||
MOSAIC_CONFIG="$SANDBOX/default-adapter.json" $CONFIG_OP validate | grep -q '"adapter": "pi"'
|
||||
check "absent adapter defaults to pi" $?
|
||||
|
||||
cfg mock-adapter.json '{"configVersion":1,"environment":"development","dataRoot":"'$DATA_ROOT'","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"mock"}}'
|
||||
expect_exit "adapter mock validates" 0 -- env MOSAIC_CONFIG="$SANDBOX/mock-adapter.json" $CONFIG_OP validate
|
||||
|
||||
cfg bad-adapter.json '{"configVersion":1,"environment":"development","dataRoot":"'$DATA_ROOT'","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"claude"}}'
|
||||
expect_exit "unsupported adapter exits 2" 2 -- env MOSAIC_CONFIG="$SANDBOX/bad-adapter.json" $CONFIG_OP validate
|
||||
|
||||
MOSAIC_CONFIG="$SANDBOX/mock-adapter.json" $CONFIG_OP env | grep -q "MOSAIC_ADAPTER='mock'"
|
||||
check "env exports adapter" $?
|
||||
|
||||
# --- bootstrap ---
|
||||
rm -f "$SANDBOX/config.json"
|
||||
expect_exit "bootstrap creates default when absent" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/config.json" $CONFIG_OP bootstrap
|
||||
[ -f "$SANDBOX/config.json" ] && { PASS=$((PASS+1)); echo "ok bootstrap wrote config file"; } \
|
||||
|| { FAIL=$((FAIL+1)); echo "FAIL bootstrap wrote config file"; }
|
||||
[ -f "$SANDBOX/config.json" ] && { PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} bootstrap wrote config file"; } \
|
||||
|| { FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} bootstrap wrote config file"; }
|
||||
|
||||
SUM_BEFORE=$(sha256sum "$SANDBOX/config.json" | cut -d' ' -f1)
|
||||
MTIME_BEFORE=$(stat -c %Y "$SANDBOX/config.json")
|
||||
@@ -55,9 +79,9 @@ expect_exit "bootstrap is idempotent on existing config" 0 -- \
|
||||
SUM_AFTER=$(sha256sum "$SANDBOX/config.json" | cut -d' ' -f1)
|
||||
MTIME_AFTER=$(stat -c %Y "$SANDBOX/config.json")
|
||||
if [ "$SUM_BEFORE" = "$SUM_AFTER" ] && [ "$MTIME_BEFORE" = "$MTIME_AFTER" ]; then
|
||||
PASS=$((PASS+1)); echo "ok bootstrap did not rewrite existing config"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} bootstrap did not rewrite existing config"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL bootstrap rewrote existing config"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} bootstrap rewrote existing config"
|
||||
fi
|
||||
|
||||
# --- validate ---
|
||||
@@ -122,9 +146,9 @@ cfg valid.json "$(valid_body "$DATA_ROOT")"
|
||||
EVAL_OUT="$(MOSAIC_CONFIG="$SANDBOX/valid.json" $CONFIG_OP env)" || true
|
||||
if eval "$EVAL_OUT" 2>/dev/null && [ "$MOSAIC_DATA_ROOT" = "$DATA_ROOT" ] \
|
||||
&& [ "$MOSAIC_PROVIDER" = "zai" ] && [ "$MOSAIC_MODEL" = "glm-5.3-flash" ]; then
|
||||
PASS=$((PASS+1)); echo "ok env exports resolve correctly"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} env exports resolve correctly"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL env exports resolve correctly"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} env exports resolve correctly"
|
||||
fi
|
||||
|
||||
# --- validation must not modify the file ---
|
||||
@@ -132,9 +156,9 @@ SUM_INVALID_BEFORE=$(sha256sum "$SANDBOX/invalid.json" | cut -d' ' -f1)
|
||||
MOSAIC_CONFIG="$SANDBOX/invalid.json" $CONFIG_OP validate >/dev/null 2>&1
|
||||
SUM_INVALID_AFTER=$(sha256sum "$SANDBOX/invalid.json" | cut -d' ' -f1)
|
||||
if [ "$SUM_INVALID_BEFORE" = "$SUM_INVALID_AFTER" ]; then
|
||||
PASS=$((PASS+1)); echo "ok failed validation modified nothing"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} failed validation modified nothing"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL failed validation modified the file"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} failed validation modified the file"
|
||||
fi
|
||||
|
||||
echo
|
||||
|
||||
Executable
+104
@@ -0,0 +1,104 @@
|
||||
#!/usr/bin/env bash
|
||||
# Sandboxed selftests for the release layer.
|
||||
#
|
||||
# Fast cases (version validation) need no Docker. State-machine cases
|
||||
# (status/activate/refusal) run against a sandboxed config and therefore
|
||||
# require the Docker daemon; they are skipped when it is unavailable.
|
||||
set -uo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
|
||||
SANDBOX="$(mktemp -d)"
|
||||
RELEASE_BACKUP="$(mktemp)"
|
||||
cp RELEASE "$RELEASE_BACKUP"
|
||||
# One exit trap: the repo RELEASE is ALWAYS restored from the backup,
|
||||
# regardless of how the test run ends.
|
||||
trap 'cp "$RELEASE_BACKUP" RELEASE 2>/dev/null; rm -rf "$SANDBOX" "$RELEASE_BACKUP"' EXIT
|
||||
|
||||
PASS=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
FAIL=0
|
||||
|
||||
expect_exit() {
|
||||
local name="$1" expected="$2"
|
||||
shift 3
|
||||
local rc
|
||||
"$@" >/dev/null 2>&1
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
# ---------- fast: release identity ----------
|
||||
expect_exit "valid RELEASE resolves" 0 -- bash -c 'source scripts/common.sh && load_release'
|
||||
|
||||
printf 'garbage\n' > RELEASE
|
||||
expect_exit "invalid RELEASE exits 1" 1 -- bash -c 'source scripts/common.sh && load_release'
|
||||
|
||||
mv RELEASE "$SANDBOX/RELEASE.hidden"
|
||||
expect_exit "missing RELEASE exits 1" 1 -- bash -c 'source scripts/common.sh && load_release'
|
||||
cp "$RELEASE_BACKUP" RELEASE
|
||||
|
||||
bash -c 'source scripts/common.sh && load_release' >/dev/null 2>&1
|
||||
bash -c 'source scripts/common.sh && load_release && case "$MOSAIC_IMAGE_TAG" in mosaic-poc-agent:*-r'"$(cat RELEASE)"') exit 0;; *) exit 1;; esac' >/dev/null 2>&1
|
||||
check "valid RELEASE leaves image tag consistent with version" $?
|
||||
|
||||
# ---------- sandboxed state machine (Docker required) ----------
|
||||
if docker info >/dev/null 2>&1; then
|
||||
mkdir -p "$SANDBOX/data"
|
||||
cat > "$SANDBOX/config.json" <<EOF
|
||||
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"glm-5.3-flash"}}
|
||||
EOF
|
||||
export MOSAIC_CONFIG="$SANDBOX/config.json"
|
||||
|
||||
expect_exit "status safe on empty state" 0 -- scripts/release.sh status
|
||||
[ ! -e "$SANDBOX/data/state/active.json" ] \
|
||||
&& check "status created no pointer" 0 || check "status created no pointer" 1
|
||||
|
||||
expect_exit "fault-injected activation refuses" 1 -- scripts/release.sh activate --fault-injection
|
||||
[ ! -e "$SANDBOX/data/state/active.json" ] \
|
||||
&& check "refused activation wrote no pointer" 0 || check "refused activation wrote no pointer" 1
|
||||
if [ -f "$SANDBOX/data/state/activation-log.jsonl" ]; then
|
||||
node -e '
|
||||
const fs = require("fs");
|
||||
const lines = fs.readFileSync(process.argv[1], "utf8").split("\n").filter(Boolean);
|
||||
if (lines.length !== 1) process.exit(1);
|
||||
const e = JSON.parse(lines[0]);
|
||||
process.exit(e.event === "refused" && e.release && e.imageTag && e.at ? 0 : 1);
|
||||
' "$SANDBOX/data/state/activation-log.jsonl"
|
||||
check "refusal logged exactly once with valid fields" $?
|
||||
else
|
||||
check "refusal logged exactly once with valid fields" 1
|
||||
fi
|
||||
|
||||
expect_exit "healthy activation succeeds" 0 -- scripts/release.sh activate
|
||||
node -e '
|
||||
const fs = require("fs");
|
||||
const p = JSON.parse(fs.readFileSync(process.argv[1], "utf8"));
|
||||
process.exit(p.pointerVersion === 1 && p.release && p.imageTag && p.activatedAt ? 0 : 1);
|
||||
' "$SANDBOX/data/state/active.json"
|
||||
check "pointer written with valid fields" $?
|
||||
|
||||
expect_exit "repeat activation succeeds (log grows)" 0 -- scripts/release.sh activate
|
||||
LINES=$(grep -c '' "$SANDBOX/data/state/activation-log.jsonl")
|
||||
[ "$LINES" -ge 3 ] && check "log is append-only across activations" 0 || check "log is append-only across activations" 1
|
||||
|
||||
expect_exit "rollback without previous refuses" 1 -- scripts/release.sh rollback
|
||||
else
|
||||
echo "skip state-machine cases (docker daemon unavailable)"
|
||||
fi
|
||||
|
||||
echo
|
||||
echo "selftest: $PASS passed, $FAIL failed"
|
||||
[ "$FAIL" -eq 0 ]
|
||||
+230
-9
@@ -11,6 +11,12 @@ SANDBOX="$(mktemp -d)"
|
||||
trap 'rm -rf "$SANDBOX"' EXIT
|
||||
|
||||
PASS=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
FAIL=0
|
||||
|
||||
expect_exit() {
|
||||
@@ -20,14 +26,20 @@ expect_exit() {
|
||||
"$@" >/dev/null 2>&1
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS+1)); echo "ok $name (exit $rc)"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL $name (exit $rc, expected $expected)"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
latest_reason() {
|
||||
local latest
|
||||
latest="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
|
||||
[ -n "$latest" ] && node -e 'try{const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.reason??"")}catch{console.log("")}' "$latest/result.json" 2>/dev/null
|
||||
}
|
||||
|
||||
CONFIG="$SANDBOX/config.json"
|
||||
@@ -92,13 +104,200 @@ $TASK validate "$SANDBOX/ok.json" >/dev/null 2>&1
|
||||
M2=$(stat -c %Y "$SANDBOX/ok.json")
|
||||
[ "$M1" = "$M2" ] && check "validation does not modify the task file" 0 || check "validation does not modify the task file" 1
|
||||
|
||||
# ---------- live: real runs (Docker + credentials required) ----------
|
||||
# ---------- retention: prune (deterministic, no Docker) ----------
|
||||
mkdir -p "$SANDBOX/data"
|
||||
cat > "$SANDBOX/prune-config.json" <<EOF
|
||||
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m"}}
|
||||
EOF
|
||||
for i in 1 2 3 4 5; do
|
||||
D="$SANDBOX/data/runs/r-20260903T0100_0${i}Z-suite00$i"
|
||||
mkdir -p "$D"
|
||||
printf '{"runVersion":1,"runId":"r-20260903T0100_0%sZ-suite00%s","taskId":"t","status":"succeeded"}' "$i" "$i" > "$D/result.json"
|
||||
done
|
||||
mkdir -p "$SANDBOX/data/sessions/sentinel" "$SANDBOX/data/workspaces/sentinel"
|
||||
|
||||
expect_exit "prune dry-run exits 0" 0 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune
|
||||
[ "$(ls "$SANDBOX/data/runs" | grep -c '^r-')" -eq 5 ] \
|
||||
&& check "dry-run deleted nothing" 0 || check "dry-run deleted nothing" 1
|
||||
|
||||
expect_exit "prune --keep=2 --yes removes oldest" 0 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune --keep=2 --yes
|
||||
[ "$(ls "$SANDBOX/data/runs" | grep -c '^r-')" -eq 2 ] \
|
||||
&& check "kept exactly 2 newest runs" 0 || check "kept exactly 2 newest runs" 1
|
||||
NEWEST="r-20260903T0100_05Z-suite005"
|
||||
[ -d "$SANDBOX/data/runs/$NEWEST" ] \
|
||||
&& check "newest run kept, oldest pruned" 0 || check "newest run kept, oldest pruned" 1
|
||||
[ -f "$SANDBOX/data/runs/.pruned.log" ] \
|
||||
&& [ "$(grep -c 'pruned' "$SANDBOX/data/runs/.pruned.log")" -eq 3 ] \
|
||||
&& check "append-only receipt written (3 entries)" 0 \
|
||||
|| check "append-only receipt written (3 entries)" 1
|
||||
[ -d "$SANDBOX/data/sessions/sentinel" ] && [ -d "$SANDBOX/data/workspaces/sentinel" ] \
|
||||
&& check "sessions/workspaces untouched by prune" 0 \
|
||||
|| check "sessions/workspaces untouched by prune" 1
|
||||
expect_exit "prune with invalid keep exits 4" 4 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune --keep=0 --yes
|
||||
|
||||
# ---------- adapter seam: deterministic mock cases (Docker, no provider) ----------
|
||||
if docker info >/dev/null 2>&1; then
|
||||
expect_exit "live hello task succeeds with exact marker" 0 -- \
|
||||
good_task "$SANDBOX/ok.json"
|
||||
mock_config() { # file adapter
|
||||
cat > "$SANDBOX/$1" <<EOF
|
||||
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"$2"}}
|
||||
EOF
|
||||
}
|
||||
|
||||
mock_config mock-adapters.json mock
|
||||
expect_exit "mock adapter: gate passes on matching mock response" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOSAIC_HELLO_OK \
|
||||
scripts/run-task.sh run "$SANDBOX/ok.json"
|
||||
|
||||
[ -f "$DATA_ROOT/runs" ] && RUNS1=$(ls "$DATA_ROOT/runs" | wc -l)
|
||||
R1="$(ls "$DATA_ROOT/runs" | head -1)"
|
||||
expect_exit "mock adapter: expect-mismatch recorded" 1 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=SOMETHING_ELSE \
|
||||
scripts/run-task.sh run "$SANDBOX/ok.json"
|
||||
[ "$(latest_reason)" = "expect-mismatch" ] \
|
||||
&& check "mismatch reason recorded" 0 || check "mismatch reason recorded" 1
|
||||
|
||||
mock_config bad-adapter.json nonexistent
|
||||
# The config layer rejects unknown adapters first, so run-task.sh fails
|
||||
# closed (exit 1) before any container or run record exists.
|
||||
expect_exit "unknown adapter fails closed" 1 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/bad-adapter.json" scripts/run-task.sh run "$SANDBOX/ok.json"
|
||||
|
||||
good_mission "$SANDBOX/m-ok.json"
|
||||
printf '{"taskVersion":1,"id":"t-mission","prompt":"ignored by mock","mission":"m-ok.json"}' > "$SANDBOX/mission-task.json"
|
||||
# Dedicated mission with distinctive directives for the injection assertion.
|
||||
cat > "$SANDBOX/m-seam.json" <<'EOF'
|
||||
{"missionVersion":1,"id":"m-seam","objective":"Prove the mission injection point.","directives":["Seam directive A.","Seam directive B."]}
|
||||
EOF
|
||||
printf '{"taskVersion":1,"id":"t-mission-seam","prompt":"ignored by mock","mission":"m-seam.json"}' > "$SANDBOX/seam-task.json"
|
||||
expect_exit "mission task runs via mock adapter" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
|
||||
scripts/run-task.sh run "$SANDBOX/seam-task.json"
|
||||
grep -q 'MISSION (runtime)' "$SANDBOX/data/system-prompt.md" \
|
||||
&& grep -q 'Seam directive A.' "$SANDBOX/data/system-prompt.md" \
|
||||
&& check "mission section injected into generated prompt" 0 \
|
||||
|| check "mission section injected into generated prompt" 1
|
||||
|
||||
# retry lineage + relative mission path resolution
|
||||
expect_exit "retry of mission run succeeds" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
|
||||
node scripts/mosaic-task.mjs retry "$(ls -dt "$SANDBOX/data/runs"/r-* | head -1 | xargs basename)"
|
||||
RT="$(ls -dt "$SANDBOX/data/runs"/r-* | head -1 | xargs basename)"
|
||||
ORIG="$(ls -dt "$SANDBOX/data/runs"/r-* | sed -n 2p | xargs basename)"
|
||||
node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));process.exit(r.retriedFrom===process.argv[2]?0:1)' \
|
||||
"$SANDBOX/data/runs/$RT/result.json" "$ORIG" \
|
||||
&& check "retriedFrom lineage recorded" 0 || check "retriedFrom lineage recorded" 1
|
||||
grep -q 'MISSION (runtime)' "$SANDBOX/data/system-prompt.md" \
|
||||
&& check "mission section present after retry (relative path resolved)" 0 \
|
||||
|| check "mission section present after retry (relative path resolved)" 1
|
||||
expect_exit "retry of missing run exits 4" 4 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" node scripts/mosaic-task.mjs retry r-missing
|
||||
|
||||
# session fork plumbing (M11): fork source + target dir delivered
|
||||
printf '{"taskVersion":1,"id":"t-forkplumb","prompt":"x","session":"fork-child","sessionForkFrom":"base"}' > "$SANDBOX/forkplumb.json"
|
||||
mkdir -p "$SANDBOX/data/sessions/base"
|
||||
printf '{}' > "$SANDBOX/data/sessions/base/20260903T000000-plumb.jsonl"
|
||||
expect_exit "fork task runs via mock" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
|
||||
scripts/run-task.sh run "$SANDBOX/forkplumb.json"
|
||||
FL="$(ls -dt "$SANDBOX/data/runs"/r-* | head -1)"
|
||||
grep -q '^MOSAIC_SESSION_FORK=/var/lib/mosaic/sessions/base/20260903T000000-plumb.jsonl$' "$FL/stderr.txt" 2>/dev/null \
|
||||
&& grep -q '^MOSAIC_SESSION_DIR=/var/lib/mosaic/sessions/fork-child$' "$FL/stderr.txt" 2>/dev/null \
|
||||
&& check "fork source + target delivered to adapter" 0 \
|
||||
|| check "fork source + target delivered to adapter" 1
|
||||
printf '{"taskVersion":1,"id":"t-f2","prompt":"x","sessionForkFrom":"base"}' > "$SANDBOX/noforktarget.json"
|
||||
expect_exit "fork without session target exits 2" 2 -- $TASK validate "$SANDBOX/noforktarget.json"
|
||||
printf '{"taskVersion":1,"id":"t-f3","prompt":"x","session":"base","sessionForkFrom":"base"}' > "$SANDBOX/selfork.json"
|
||||
expect_exit "self-fork exits 2" 2 -- $TASK validate "$SANDBOX/selfork.json"
|
||||
|
||||
# capability policy (M9): least-privilege intersection
|
||||
POL="$SANDBOX/data/workspaces"; mkdir -p "$POL"
|
||||
pol_run() { # missionTools(ABSENT|json) taskTools(ABSENT|json) -> stderr MOSAIC_TOOLS value
|
||||
if [ "$1" = "ABSENT" ]; then
|
||||
printf '{"missionVersion":1,"id":"m-pol","objective":"o"}' > "$SANDBOX/pol-m.json"
|
||||
else
|
||||
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":[%s]}}' "$1" > "$SANDBOX/pol-m.json"
|
||||
fi
|
||||
if [ "$2" = "ABSENT" ]; then
|
||||
printf '{"taskVersion":1,"id":"t-pol","prompt":"x","mission":"pol-m.json"}' > "$SANDBOX/pol-t.json"
|
||||
else
|
||||
printf '{"taskVersion":1,"id":"t-pol","prompt":"x","mission":"pol-m.json","capabilities":{"tools":[%s]}}' "$2" > "$SANDBOX/pol-t.json"
|
||||
fi
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
|
||||
scripts/run-task.sh run "$SANDBOX/pol-t.json" >/dev/null 2>&1
|
||||
grep -o '^MOSAIC_TOOLS=.*' "$(ls -dt "$SANDBOX/data/runs"/r-* | head -1)/stderr.txt" 2>/dev/null
|
||||
}
|
||||
[ "$(pol_run '"read","bash"' ABSENT)" = 'MOSAIC_TOOLS=read,bash' ] \
|
||||
&& check "policy: mission only -> mission tools" 0 || check "policy: mission only -> mission tools" 1
|
||||
[ "$(pol_run ABSENT '"read","bash"')" = 'MOSAIC_TOOLS=read,bash' ] \
|
||||
&& check "policy: task only -> task tools" 0 || check "policy: task only -> task tools" 1
|
||||
[ "$(pol_run '"read"' '"bash","read"')" = 'MOSAIC_TOOLS=read' ] \
|
||||
&& check "policy: both -> intersection (task narrowed)" 0 || check "policy: both -> intersection (task narrowed)" 1
|
||||
[ "$(pol_run '"grep"' '"bash","read"')" = 'MOSAIC_TOOLS=' ] \
|
||||
&& check "policy: empty intersection -> tool-free" 0 || check "policy: empty intersection -> tool-free" 1
|
||||
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":["sudo"]}}' > "$SANDBOX/pol-m.json"
|
||||
expect_exit "invalid mission capabilities rejected" 2 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" $TASK validate "$SANDBOX/pol-t.json"
|
||||
else
|
||||
echo "skip adapter seam cases (docker daemon unavailable)"
|
||||
fi
|
||||
|
||||
# ---------- workspace + capabilities (M5): deterministic mock cases ----------
|
||||
if docker info >/dev/null 2>&1; then
|
||||
printf '{"taskVersion":1,"id":"t-ws","prompt":"ignored","workspace":"suitews","capabilities":{"tools":["read","bash"]}}' > "$SANDBOX/ws-task.json"
|
||||
expect_exit "workspace+tools task runs via mock" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
|
||||
scripts/run-task.sh run "$SANDBOX/ws-task.json"
|
||||
WSL="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
|
||||
grep -q '^MOSAIC_WORKSPACE=/var/lib/mosaic/workspaces/suitews$' "$WSL/stderr.txt" 2>/dev/null \
|
||||
&& grep -q '^MOSAIC_TOOLS=read,bash$' "$WSL/stderr.txt" 2>/dev/null \
|
||||
&& check "workspace path + tools delivered to adapter" 0 \
|
||||
|| check "workspace path + tools delivered to adapter" 1
|
||||
[ -d "$SANDBOX/data/workspaces/suitews" ] \
|
||||
&& check "persistent workspace created on host" 0 \
|
||||
|| check "persistent workspace created on host" 1
|
||||
|
||||
expect_exit "plain task still runs (no workspace/tools)" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOSAIC_HELLO_OK \
|
||||
scripts/run-task.sh run "$SANDBOX/ok.json"
|
||||
PL="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
|
||||
grep -q '^MOSAIC_WORKSPACE=$' "$PL/stderr.txt" 2>/dev/null \
|
||||
&& check "workspace var present but empty when absent" 0 || check "workspace var present but empty when absent" 1
|
||||
grep -q '^MOSAIC_TOOLS=$' "$PL/stderr.txt" 2>/dev/null \
|
||||
&& check "tools empty when absent" 0 || check "tools empty when absent" 1
|
||||
|
||||
printf '{"taskVersion":1,"id":"t-badtool","prompt":"x","capabilities":{"tools":["sudo"]}}' > "$SANDBOX/badtool.json"
|
||||
expect_exit "unknown tool exits 2" 2 -- $TASK validate "$SANDBOX/badtool.json"
|
||||
printf '{"taskVersion":1,"id":"t-badws","prompt":"x","workspace":"../escape"}' > "$SANDBOX/badws.json"
|
||||
expect_exit "workspace traversal exits 2" 2 -- $TASK validate "$SANDBOX/badws.json"
|
||||
else
|
||||
echo "skip workspace/capability cases (docker daemon unavailable)"
|
||||
fi
|
||||
|
||||
# ---------- live: real runs (Docker + credentials required) ----------
|
||||
# On failure, surface the run record + agent stderr BEFORE the sandbox
|
||||
# cleanup destroys them. Never let a wrong-exit mask the real reason.
|
||||
dump_latest_run() {
|
||||
local latest
|
||||
latest="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
|
||||
if [ -n "$latest" ]; then
|
||||
echo "--- latest run evidence: $latest ---" >&2
|
||||
cat "$latest/result.json" 2>/dev/null >&2
|
||||
echo "--- stderr.txt (tail) ---" >&2
|
||||
tail -8 "$latest/stderr.txt" 2>/dev/null >&2
|
||||
else
|
||||
echo "--- no run dir was created at all ---" >&2
|
||||
fi
|
||||
}
|
||||
|
||||
if docker info >/dev/null 2>&1; then
|
||||
RUNS1=$(ls "$DATA_ROOT/runs" 2>/dev/null | wc -l)
|
||||
if scripts/run-task.sh run "$SANDBOX/ok.json" >/dev/null 2>&1; then
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} live hello task succeeds with exact marker"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} live hello task succeeds with exact marker" >&2
|
||||
dump_latest_run
|
||||
fi
|
||||
|
||||
R1="$(ls -t "$DATA_ROOT/runs" | head -1)" # newest = the live run above
|
||||
[ -f "$DATA_ROOT/runs/$R1/result.json" ] && check "result.json written in run dir" 0 || check "result.json written in run dir" 1
|
||||
node -e '
|
||||
const r = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
|
||||
@@ -107,8 +306,15 @@ process.exit(r.status === "succeeded" && r.response === "MOSAIC_HELLO_OK" && r.e
|
||||
check "result.json contents are correct" $?
|
||||
|
||||
printf '{"taskVersion":1,"id":"t-wrong","prompt":"Return your startup marker and nothing else.","expectExact":"MOSAIC_NOT_OK"}' > "$SANDBOX/wrong.json"
|
||||
expect_exit "wrong expectExact fails with exit 1" 1 -- \
|
||||
scripts/run-task.sh run "$SANDBOX/wrong.json"
|
||||
scripts/run-task.sh run "$SANDBOX/wrong.json" >/dev/null 2>&1
|
||||
RC=$?
|
||||
WRONG_REASON="$(latest_reason)"
|
||||
if [ "$RC" -eq 1 ] && [ "$WRONG_REASON" = "expect-mismatch" ]; then
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} wrong expectExact fails with exit 1 (reason: expect-mismatch)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} wrong expectExact (exit $RC, reason: '${WRONG_REASON:-none}')" >&2
|
||||
dump_latest_run
|
||||
fi
|
||||
|
||||
RUNS2=$(ls "$DATA_ROOT/runs" | wc -l)
|
||||
[ "$RUNS2" -gt "${RUNS1:-0}" ] && check "each run gets a distinct run dir (no clobber)" 0 \
|
||||
@@ -116,6 +322,21 @@ process.exit(r.status === "succeeded" && r.response === "MOSAIC_HELLO_OK" && r.e
|
||||
|
||||
COUNT=$($TASK list | wc -l)
|
||||
[ "$COUNT" -ge 2 ] && check "list shows both runs" 0 || check "list shows both runs" 1
|
||||
|
||||
# session fork (M11): teach in base, fork into child, child recalls;
|
||||
# ancestor file count must be unchanged by the fork
|
||||
printf '{"taskVersion":1,"id":"t-fork-base","prompt":"Remember this code word for later: mosaico. Reply with exactly: REMEMBERED","session":"suite-base","expectExact":"REMEMBERED","timeoutSeconds":180}' > "$SANDBOX/fb.json"
|
||||
expect_exit "fork base: teach succeeds" 0 -- scripts/run-task.sh run "$SANDBOX/fb.json"
|
||||
BASECOUNT=$(ls "$SANDBOX/data/sessions/suite-base" | wc -l)
|
||||
printf '{"taskVersion":1,"id":"t-fork-child","prompt":"What code word did I ask you to remember? Reply with only the code word.","session":"suite-child","sessionForkFrom":"suite-base","timeoutSeconds":180}' > "$SANDBOX/fc.json"
|
||||
expect_exit "fork child recalls ancestor context" 0 -- scripts/run-task.sh run "$SANDBOX/fc.json"
|
||||
FR="$(ls -dt "$DATA_ROOT/runs"/r-* | head -1)"
|
||||
node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));process.exit(String(r.response).toLowerCase().includes("mosaico")?0:1)' "$FR/result.json" \
|
||||
&& check "forked child recalled ancestor code word" 0 || check "forked child recalled ancestor code word" 1
|
||||
[ "$(ls "$SANDBOX/data/sessions/suite-base" | wc -l)" -eq "$BASECOUNT" ] \
|
||||
&& check "ancestor session untouched by fork" 0 || check "ancestor session untouched by fork" 1
|
||||
[ "$(ls "$SANDBOX/data/sessions/suite-child" | wc -l)" -ge 1 ] \
|
||||
&& check "child session dir has its own branch file" 0 || check "child session dir has its own branch file" 1
|
||||
else
|
||||
echo "skip live task cases (docker unavailable)"
|
||||
fi
|
||||
|
||||
+11
-3
@@ -14,10 +14,18 @@ cd "$(dirname "$0")/.."
|
||||
# shellcheck source=common.sh
|
||||
source scripts/common.sh
|
||||
|
||||
IMAGE="mosaic-poc-agent:0.84.4"
|
||||
EXPECTED="${EXPECTED_MARKER:-MOSAIC_HELLO_OK}"
|
||||
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
|
||||
load_config
|
||||
load_release
|
||||
IMAGE="$MOSAIC_IMAGE_TAG"
|
||||
|
||||
# Ensure the configured data root exists (host-owned) before the mount,
|
||||
# otherwise Docker would auto-create a root-owned directory.
|
||||
@@ -50,11 +58,11 @@ TRIMMED="$(printf '%s' "$RESPONSE" | sed -e 's/^[[:space:]]*//' -e 's/[[:space:]
|
||||
|
||||
# 4-6. Exact comparison gate.
|
||||
if [ "$TRIMMED" = "$EXPECTED" ]; then
|
||||
echo "PASS: response matches expected marker"
|
||||
echo "${C_OK}PASS${C_RESET}: response matches expected marker"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "FAIL: response does not match expected marker" >&2
|
||||
echo "${C_FAIL}FAIL${C_RESET}: response does not match expected marker" >&2
|
||||
printf 'expected: %s\n' "$EXPECTED" >&2
|
||||
printf 'actual : %s\n' "$TRIMMED" >&2
|
||||
exit 1
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
---
|
||||
name: ms-tools
|
||||
description: Contains a reference to all available tools for Mosaic Stack.
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
# ms-tools
|
||||
|
||||
You are a Mosaic fleet agent. A maintained toolkit lives at `<dir>`.
|
||||
Use it FIRST for the tasks below — improvising with raw CLIs causes the recurring failures this
|
||||
skill exists to prevent. This is the high-frequency fast path; the full reference is the
|
||||
`# Machine Tools` section already in your system prompt.
|
||||
|
||||
|
||||
@@ -0,0 +1,85 @@
|
||||
---
|
||||
name: ms-unslop
|
||||
description: Cut AI tells from any writing. Must always apply.
|
||||
disable-model-invocation: false
|
||||
---
|
||||
|
||||
# Unslop
|
||||
|
||||
Edit text to remove AI patterns and add human voice.
|
||||
|
||||
## Process
|
||||
|
||||
1. Scan for the patterns below.
|
||||
2. Rewrite. Preserve meaning, match intended tone.
|
||||
3. Add soul (see next section).
|
||||
4. Self-audit: "What makes this obviously AI generated?" Fix remaining tells.
|
||||
|
||||
## Adding soul
|
||||
|
||||
Removing patterns is half the job. Sterile, voiceless writing is just as obvious.
|
||||
|
||||
- **Have opinions.** React to facts instead of neutrally listing pros and cons.
|
||||
- **Vary rhythm.** Short sentences. Then longer ones that take their time. Mix it up.
|
||||
- **Acknowledge complexity.** "Impressive but also kind of unsettling" beats "impressive."
|
||||
- **Use "I" when it fits.** First person isn't unprofessional.
|
||||
- **Let some mess in.** Perfect structure looks machine-made.
|
||||
- **Be specific.** Not "this is concerning" but "there's something unsettling about agents churning away at 3am."
|
||||
|
||||
## Patterns to detect and fix
|
||||
|
||||
### Content
|
||||
|
||||
1. **Puffery.** `pivotal moment`, `testament to`, "evolving landscape", "setting the stage for", "indelible mark", "deeply rooted". Cut puffery, state what happened.
|
||||
2. **Name-dropping.** Listing media outlets without context. Pick one, say what was said.
|
||||
3. **Superficial -ing phrases.** "highlighting...", "ensuring...", "reflecting...", "showcasing...", "fostering...". Delete or expand with real sources.
|
||||
4. **Promotional language.** "nestled", `vibrant`, "breathtaking", "groundbreaking", "renowned", "stunning", "must-visit". Use neutral descriptions.
|
||||
5. **Vague attributions.** "Experts believe", "Industry reports suggest", "Some critics argue". Name the source or delete.
|
||||
6. **Formulaic challenges.** "Despite challenges... continues to thrive." Replace with specific facts.
|
||||
|
||||
### Language
|
||||
|
||||
7. **AI vocabulary.** `Additionally`, `crucial`, `delve`, `enduring`, `enhance`, `fostering`, `garner`, `interplay`, `intricate`, `landscape` (abstract), `pivotal`, `showcase`, `tapestry` (abstract), `testament`, `underscore`, `vibrant`. Replace with plain words.
|
||||
8. **Fancy ways to say "is".** "serves as", "stands as", "boasts", "features". Just say "is" or "has".
|
||||
9. **`Not just X, but Y`.** State the point directly instead.
|
||||
10. **Rule of three.** Forcing ideas into groups of three. Use the natural number.
|
||||
11. **Synonym cycling.** Protagonist, main character, central figure, hero all in one paragraph. Pick one, repeat it.
|
||||
12. **False ranges.** "from X to Y" where X and Y aren't on a meaningful scale. List topics directly.
|
||||
|
||||
### Style
|
||||
|
||||
13. **Em dash overuse.** Avoid em dashes entirely. Use periods or commas only (no parentheses, no en dashes, no hyphen-as-dash substitutes). Em dashes are an AI tell, and reaching for parentheses instead just trades one tell for another. If a thought needs separation, end the sentence or use a comma.
|
||||
14. **Colon overuse.** Colons are fine before a list or example. Not as mid-sentence connectors. "If you're coming from traditional automation: instead of registering event handlers, you describe conditions" adds nothing with the colon. Rewrite to let the point stand on its own without comparison framing. "Describing when the scheduler should fire works best as plain English." Same meaning, no crutch punctuation.
|
||||
15. **Boldface overuse.** Don't bold every proper noun or acronym.
|
||||
16. **Inline-header lists.** The tell is a bold label and colon that restates the line: "**Performance:** Performance improved...". Convert those to prose. A bold lead-in that ends in a period, names the item, and is followed by genuinely new detail ("**Schema in TypeScript.** Tables live in one file.") is fine, not a tell.
|
||||
17. **Title case headings.** Use sentence case.
|
||||
18. **Decorative emojis.** Remove from headings and bullets.
|
||||
19. **Curly quotes.** Replace with straight quotes.
|
||||
|
||||
### Communication artifacts
|
||||
|
||||
20. **Chatbot phrases.** `I hope this helps!`, `Let me know if...`, `Of course!`, `Certainly!`, `Found the smoking gun!` Remove.
|
||||
21. **Cutoff disclaimers.** "While specific details are limited..." Find sources or remove.
|
||||
22. **Sycophantic tone.** `Great question!` `You're absolutely right!` Respond directly.
|
||||
|
||||
### Filler
|
||||
|
||||
23. **Filler phrases.** `In order to` becomes "To". `Due to the fact that` becomes "Because". `It is important to note that` gets deleted.
|
||||
24. **Excessive hedging.** "could potentially possibly be argued that it might" becomes "may".
|
||||
25. **Generic conclusions.** "The future looks bright." State specific plans or facts.
|
||||
|
||||
### Jargon
|
||||
|
||||
26. **Abstract metaphor nouns.** Substrate, wedge, vector, locus, vantage, nexus, primitive (as noun), harness (as metaphor), surface (as in "API surface"), bedrock, scaffolding (as metaphor), modality, paradigm, gold-plating, ratchet (as metaphor), evacuate (for moving code), endgame, north star, flywheel. These read as technical but usually have a plainer concrete word. "Substrate" becomes "base". "Wedge in" becomes "add". "Vector" becomes "way" or "method". "Gold-plating" becomes "more than the job needs". "Ratchet" becomes the mechanism's real name or "a limit that only tightens". "Evacuate" becomes "move out". "Endgame" becomes "the last phase". Pick the concrete word.
|
||||
|
||||
### Plain speech
|
||||
|
||||
27. **Say what it does, not how it feels.** "the database stays close at hand", "SQL you can read", "types that follow your schema" name a feeling. The fix names the mechanism or a number: "`.toSQL()` returns the exact string sent to the database", "a column rename fails the build". Ask what the sentence tells the reader to do or know, then write that. If you can't restate it as a concrete instruction, fact, or number, cut it. One more check: if the sentence could appear unchanged in another project's docs, it says nothing about this one. Cut it.
|
||||
28. **Shorten or split dense sentences.** If the reader has to backtrack to parse a sentence, break it in two or drop clauses. One idea per sentence.
|
||||
29. **Active voice.** Prefer it. Catch "is/are/was/were + past participle" and name the actor: "queries are validated" becomes "the compiler validates queries", "the file is parsed by the loader" becomes "the loader parses the file". Passive is fine only when the actor is unknown or genuinely doesn't matter.
|
||||
30. **Cut adverbs, or use a stronger verb.** "runs quickly" becomes "is fast" or the number. "significantly improves" becomes the measured delta. An adverb propping up a weak verb means the verb is wrong.
|
||||
31. **Prefer the plain word.** `utilize` becomes "use", `leverage` becomes "use", `facilitate` becomes "help", "numerous" becomes "many", "in the event that" becomes "if". The fancier synonym is rarely clearer.
|
||||
|
||||
## Mention convention
|
||||
|
||||
A document that MENTIONS a banned word or phrase quotes it as inline code. The checker (`tools/unslop-hook/unslop-check.js`, machine source `tools/unslop-hook/lists.json`) strips code spans before matching, so a backticked mention is invisible to the gate while a bare one flags. This file follows that convention and doubles as a regression fixture: if `unslop-check.js` ever flags this file, either an edit broke the mention convention or code stripping regressed. Documents that deliberately CONTAIN slop to test detection (fixture files) are uses, not mentions; they are expected to flag.
|
||||
@@ -33,5 +33,31 @@ for f in $FILES; do
|
||||
printf '\n' >> "$TEMP"
|
||||
done
|
||||
|
||||
# Agent identity (M13): when the launcher names the agent, the generated
|
||||
# prompt states it - SOUL.md provides the persona, this provides the name.
|
||||
if [ -n "${MOSAIC_AGENT_NAME:-}" ]; then
|
||||
printf '===== AGENT IDENTITY =====\n' >> "$TEMP"
|
||||
printf 'agent name: %s\n' "$MOSAIC_AGENT_NAME" >> "$TEMP"
|
||||
printf '\n' >> "$TEMP"
|
||||
fi
|
||||
|
||||
# Sanctioned mission injection point (M4): when the task runner provides a
|
||||
# mission snapshot, its objective and directives are appended AFTER the
|
||||
# immutable contracts. Runtime data; never part of the contract fixtures.
|
||||
if [ -n "${MOSAIC_MISSION_FILE:-}" ]; then
|
||||
if [ ! -r "$MOSAIC_MISSION_FILE" ]; then
|
||||
echo "load-contracts: MOSAIC_MISSION_FILE set but not readable: $MOSAIC_MISSION_FILE" >&2
|
||||
rm -f "$TEMP"
|
||||
exit 1
|
||||
fi
|
||||
printf '===== MISSION (runtime) =====\n' >> "$TEMP"
|
||||
node -e '
|
||||
const m = JSON.parse(require("fs").readFileSync(process.env.MOSAIC_MISSION_FILE, "utf8"));
|
||||
process.stdout.write("Objective: " + m.objective + "\n");
|
||||
for (const d of m.directives ?? []) process.stdout.write("- " + d + "\n");
|
||||
' >> "$TEMP"
|
||||
printf '\n' >> "$TEMP"
|
||||
fi
|
||||
|
||||
mv "$TEMP" "$OUT"
|
||||
echo "load-contracts: wrote $OUT from $CONTRACT_DIR" >&2
|
||||
|
||||
+34
-31
@@ -1,38 +1,41 @@
|
||||
#!/bin/sh
|
||||
# One-shot Pi agent runner inside the container.
|
||||
# Loads the contract-generated system prompt, then sends exactly one
|
||||
# user request through Pi's documented noninteractive mode and prints
|
||||
# the model response on stdout.
|
||||
# Agent dispatcher inside the container.
|
||||
#
|
||||
# Headless (default): loads the contract-generated system prompt, then
|
||||
# dispatches one request to /opt/mosaic/adapters/<MOSAIC_ADAPTER>/adapter.sh
|
||||
# (contract: /opt/mosaic/adapters/README.md).
|
||||
#
|
||||
# Interactive (MOSAIC_INTERACTIVE=1, from scripts/agent.sh): same prompt,
|
||||
# but the adapter opens the full pi TUI with no initial prompt - the human
|
||||
# drives from there.
|
||||
set -eu
|
||||
|
||||
: "${PI_PROVIDER:=zai}"
|
||||
: "${PI_MODEL:=glm-5.3-flash}"
|
||||
export PI_PROVIDER PI_MODEL
|
||||
REQUEST=""
|
||||
if [ "${MOSAIC_INTERACTIVE:-}" != "1" ]; then
|
||||
# Headless: args are the request; default is the startup verification
|
||||
# request used by hello/verify.
|
||||
REQUEST="${*:-Return your startup marker and nothing else.}"
|
||||
export MOSAIC_REQUEST="$REQUEST"
|
||||
fi
|
||||
|
||||
REQUEST="${*:-Return your startup marker and nothing else.}"
|
||||
ADAPTER="${MOSAIC_ADAPTER:-pi}"
|
||||
case "$ADAPTER" in
|
||||
# Allowlist mirrors scripts/mosaic-config.mjs; pattern check first so a
|
||||
# crafted name cannot escape the adapters directory.
|
||||
*[!A-Za-z0-9._-]*|'')
|
||||
echo "run-agent: invalid adapter name: '$ADAPTER'" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
|
||||
ADAPTER_SCRIPT="/opt/mosaic/adapters/$ADAPTER/adapter.sh"
|
||||
if [ ! -x "$ADAPTER_SCRIPT" ]; then
|
||||
echo "run-agent: unknown or non-executable adapter: $ADAPTER" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
/opt/mosaic/src/load-contracts.sh /opt/mosaic/contracts /var/lib/mosaic/system-prompt.md
|
||||
|
||||
# All flags are documented in the package README (CLI Reference):
|
||||
# -p / --print noninteractive: print the response and exit
|
||||
# --system-prompt replace the default system prompt with the
|
||||
# contract-generated prompt
|
||||
# --no-* switches prevent ambient context files, skills, extensions,
|
||||
# prompt templates, and themes from being appended
|
||||
# --no-session ephemeral: no persistent agent session
|
||||
# --no-tools the startup request needs no tool execution
|
||||
# --offline disable startup network operations (update checks,
|
||||
# package update checks, install/update telemetry)
|
||||
exec pi \
|
||||
--offline \
|
||||
--no-session \
|
||||
--no-extensions \
|
||||
--no-skills \
|
||||
--no-prompt-templates \
|
||||
--no-themes \
|
||||
--no-context-files \
|
||||
--no-tools \
|
||||
--provider "$PI_PROVIDER" \
|
||||
--model "$PI_MODEL" \
|
||||
--system-prompt "$(cat /var/lib/mosaic/system-prompt.md)" \
|
||||
-p "$REQUEST"
|
||||
export MOSAIC_SYSTEM_PROMPT_FILE="/var/lib/mosaic/system-prompt.md"
|
||||
|
||||
exec "$ADAPTER_SCRIPT"
|
||||
|
||||
@@ -0,0 +1,8 @@
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-session-teach",
|
||||
"prompt": "Remember this code word for later: mosaico. Reply with exactly: REMEMBERED",
|
||||
"session": "demo",
|
||||
"expectExact": "REMEMBERED",
|
||||
"timeoutSeconds": 180
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-session-recall",
|
||||
"prompt": "What code word did I ask you to remember earlier in this session? Reply with only the code word.",
|
||||
"session": "demo",
|
||||
"timeoutSeconds": 180
|
||||
}
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-worker-retry-refine",
|
||||
"prompt": "Follow-up on your last change to scripts/mosaic-task.mjs (the retry operation). Conductor review found one defect: runTask builds spawnEnv from process.env, so a direct 'node scripts/mosaic-task.mjs retry <runId>' invocation lacks the launcher-exported variables and compose fails with: required variable MOSAIC_IMAGE_TAG is missing.\n\nFIX (in runTask, not retryRun): make spawnEnv self-sufficient —\n1. spawnEnv.PI_PROVIDER = resolved.execution.provider; spawnEnv.PI_MODEL = resolved.execution.model (from the already-loaded resolved config).\n2. spawnEnv.MOSAIC_IMAGE_TAG = 'mosaic-poc-agent:' + <pinned pi version> + '-r' + <RELEASE file content trimmed>, reading RELEASE and package.json the same way scripts/common.sh load_release does.\nKeep the change minimal; do not touch other files; keep code style.\n\nSELF-CHECK: node --check scripts/mosaic-task.mjs must pass.\nWhen finished, reply with exactly: REFINE_DONE",
|
||||
"workspace": "stack-repo",
|
||||
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
|
||||
"session": "worker-1",
|
||||
"timeoutSeconds": 600
|
||||
}
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-worker-retry",
|
||||
"prompt": "You are working alone inside a git repository checkout (your current working directory). Implement a new feature in scripts/mosaic-task.mjs.\n\nFEATURE: add a 'retry <runId>' operation alongside the existing 'validate | run | show | list' operations.\n\nBEHAVIOR: 'retry <runId>' re-executes a previously recorded run. Steps: (1) validate the runId argument with the same pattern showRun uses; (2) load the config exactly like showRun does; (3) locate the run directory under the runs root; if it does not exist, fail with exit code 4 and the message 'run not found: <runId>'; (4) read that run directory's task.json snapshot; if unreadable, fail with exit code 4; (5) write the snapshot to a temp file and execute it through the EXISTING runTask(path) function so a brand-new run is created — do not duplicate any execution logic; (6) exit with whatever exit code the new run produced.\n\nCONSTRAINTS: reuse runTask; do not refactor unrelated code; do not touch any file other than scripts/mosaic-task.mjs; keep the existing code style; update the default usage error message to include retry.\n\nSELF-CHECK (run it yourself with the bash tool): node --check scripts/mosaic-task.mjs must pass.\n\nWhen finished, reply with exactly: RETRY_DONE",
|
||||
"workspace": "stack-repo",
|
||||
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
|
||||
"session": "worker-1",
|
||||
"timeoutSeconds": 600
|
||||
}
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-workspace-demo",
|
||||
"prompt": "Use the bash tool to create a file named proof.txt in the current directory containing exactly the text: workspace works. Then reply with exactly: WORKSPACE_OK",
|
||||
"workspace": "demo",
|
||||
"capabilities": { "tools": ["bash", "read", "write"] },
|
||||
"expectExact": "WORKSPACE_OK",
|
||||
"timeoutSeconds": 180
|
||||
}
|
||||
@@ -0,0 +1,83 @@
|
||||
# unslop-hook
|
||||
|
||||
Mechanical AI-tell enforcement for pi seats. Anti-drift gate for the writing
|
||||
standard in SYSTEM.md / ms-unslop: prose distribution alone decays over long
|
||||
sessions; this check cannot forget.
|
||||
|
||||
- `lists.json`: committed machine source for every list the checker enforces:
|
||||
words, phrases, punct rules, regex patterns, density thresholds. Each entry
|
||||
carries provenance (`ms-unslop:<pattern id>` or `system-md`), the mention
|
||||
convention, and the documented divergence of the density gate from
|
||||
SYSTEM.md's outright em-dash ban. Edit lists here, not in code.
|
||||
- `unslop-check.js`: dependency-free checker (node CLI + module) driven by
|
||||
lists.json. Loads and schema-validates the lists on first use and hard-fails
|
||||
closed: empty, unparseable, or invalid lists throw. Detects banned vocabulary,
|
||||
chatbot/sycophancy phrases, filler phrases, em/en dashes, curly quotes,
|
||||
`not just X but Y`. Strips fenced and inline code first, so quoted code is
|
||||
never flagged. Exit 0 clean, 1 violations, 2 gate broken (lists unreadable,
|
||||
never a clean verdict).
|
||||
- `extension.ts`: pi extension. `message_end` checks finalized assistant text
|
||||
and notifies the operator (TUI/RPC). `before_agent_start` reads the most
|
||||
recent assistant reply from the session file and, if it carries tells,
|
||||
injects a correction notice the model sees on its next turn. `/unslop`
|
||||
reports session stats. Violation state lives in the session file, so the
|
||||
injection path survives restart, resume, fork, and reload (an in-memory
|
||||
pending flag was measured dead across print-mode turns, 2026-08-19). A
|
||||
broken lists.json fails closed: checks stop, `broken_lists` /
|
||||
`skipped_broken` events log the reason, operator notified once, seat keeps
|
||||
running.
|
||||
- `test-unslop-check.js`: unit tests with red and green controls.
|
||||
|
||||
## Use
|
||||
|
||||
```bash
|
||||
node test-unslop-check.js # suite
|
||||
node unslop-check.js <file> # CLI check
|
||||
UNSLOP_LISTS=<path> node unslop-check.js <file> # alt lists location
|
||||
pi -e ~/.mosaic/tools/unslop-hook/extension.ts # ad-hoc load
|
||||
# deploy: copy dir to ~/.pi/agent/extensions/unslop-hook/ or seat .pi
|
||||
# equivalent, or list it in settings.json "extensions"
|
||||
```
|
||||
|
||||
Env: `MOSAIC_UNSLOP_HOOK=0` disables. `MOSAIC_UNSLOP_LOG=<path>` appends JSONL
|
||||
events (loaded / flagged / notice_injected / checked / broken_lists /
|
||||
skipped_broken) for headless evidence. `UNSLOP_LISTS=<path>` overrides the
|
||||
lists.json location for both CLI and extension.
|
||||
|
||||
## Verified here (2026-08-19)
|
||||
|
||||
- Unit suite 24/24 (8 behavioral, 11 loader/CLI, 5 review follow-up), red and
|
||||
green controls both exercised, including exit-2 on broken lists and on an
|
||||
unreadable input file (S3).
|
||||
- CLI: slop file exit 1, clean file exit 0, broken lists exit 2 with the fault
|
||||
named on stderr.
|
||||
- Extension, healthy path (print mode, zai/glm-5.3:low): startup probe loads
|
||||
lists.json, reply checked clean.
|
||||
- Extension, broken-lists path (print mode): `broken_lists` at startup,
|
||||
`skipped_broken` per turn, seat survives, reply still delivered.
|
||||
- Earlier live evidence (pre-C1, inline lists): forced-slop turn flagged;
|
||||
fresh-process follow-up injected the notice and the reply came back clean;
|
||||
full TUI trial (notify line, injection, /unslop stats) on session vision-unslop.
|
||||
- Log evidence in session scratchpad.
|
||||
|
||||
## Limits
|
||||
|
||||
- `/unslop` command not tested headless (print mode has no command surface);
|
||||
it is a thin stats wrapper.
|
||||
- En dash flag fires on typographic ranges too (2–3); acceptable for fleet
|
||||
prose, revisit if it noisifies technical writing.
|
||||
- Notice injection is a nudger, not a blocker. Output already streamed to the
|
||||
user stays as-is; correction lands on the next turn.
|
||||
- A broken lists.json latches for the session: repairing the file mid-session
|
||||
does not revive checks until the seat restarts. Acceptable for an advisory
|
||||
gate (review S1).
|
||||
- The fail-closed operator notification requires a UI. Print-mode sessions
|
||||
log `skipped_broken` but notify nobody (review S2).
|
||||
- The word/phrase lists are the mechanical subset of ms-unslop only, keyed to
|
||||
pattern ids in lists.json. Style judgments (voice, rhythm, structure) stay in
|
||||
the skill, not the gate.
|
||||
|
||||
## Promotion path
|
||||
|
||||
Stack issue (A4): checker shared as the single source for a matching Claude
|
||||
Code Stop-hook script; lists versioned beside SYSTEM.md contract text.
|
||||
@@ -0,0 +1,163 @@
|
||||
// unslop-hook — pi extension wrapper around unslop-check.js.
|
||||
// Detects mechanical AI tells in finalized assistant messages and injects a
|
||||
// correction notice the model sees on its next turn. Anti-drift enforcement for
|
||||
// SYSTEM.md / ms-unslop; prose distribution alone decays, this cannot forget.
|
||||
//
|
||||
// Deploy: copy dir to ~/.pi/agent/extensions/unslop-hook/ (or seat .pi equivalent),
|
||||
// or add this file's dir to settings.json "extensions".
|
||||
// Test: pi -e <abs path>/extension.ts
|
||||
// Off: MOSAIC_UNSLOP_HOOK=0
|
||||
// Log: MOSAIC_UNSLOP_LOG=/path/to/log.jsonl (JSONL events; headless evidence)
|
||||
// Broken: lists.json missing/empty/invalid → the checker throws; checks are
|
||||
// skipped, logged as skipped_broken, and the operator is notified once.
|
||||
// Never silently pass while the lists cannot load (fail closed).
|
||||
//
|
||||
// Design note: violation state lives in the SESSION FILE, not memory. At
|
||||
// before_agent_start we read the most recent assistant text message from
|
||||
// ctx.sessionManager and check it there. That survives process restarts, resume,
|
||||
// fork, and reload — an in-memory pending flag measured dead on 2026-08-19 when
|
||||
// a print-mode second turn never injected.
|
||||
import { appendFileSync } from "node:fs";
|
||||
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
||||
import { checkText } from "./unslop-check.js";
|
||||
|
||||
interface Finding {
|
||||
rule: string;
|
||||
detail: string;
|
||||
count: number;
|
||||
}
|
||||
|
||||
interface MessageEntry {
|
||||
type: "message";
|
||||
id: string;
|
||||
message: { role?: string; content?: unknown };
|
||||
}
|
||||
|
||||
function assistantText(entry: unknown): string | null {
|
||||
const e = entry as Partial<MessageEntry>;
|
||||
if (e?.type !== "message") return null;
|
||||
const msg = e.message;
|
||||
if (msg?.role !== "assistant" || !Array.isArray(msg.content)) return null;
|
||||
const text = msg.content
|
||||
.filter((b): b is { type: "text"; text: string } =>
|
||||
typeof b === "object" && b !== null && (b as { type?: string }).type === "text")
|
||||
.map((b) => b.text ?? "")
|
||||
.join("\n");
|
||||
return text.trim() ? text : null; // tool-call-only assistant messages return null
|
||||
}
|
||||
|
||||
export default function (pi: ExtensionAPI) {
|
||||
if (process.env.MOSAIC_UNSLOP_HOOK === "0") return;
|
||||
|
||||
const LOG = process.env.MOSAIC_UNSLOP_LOG;
|
||||
const log = (ev: Record<string, unknown>) => {
|
||||
if (LOG) appendFileSync(LOG, JSON.stringify({ ts: Date.now(), ...ev }) + "\n");
|
||||
};
|
||||
|
||||
// Entry ids we have already injected a notice for. In-memory only: after a
|
||||
// restart the same entry may inject once more, which re-anchors the style
|
||||
// after a context loss. That is wanted, not a bug.
|
||||
const injectedFor = new Set<string>();
|
||||
let turnsChecked = 0;
|
||||
let turnsFlagged = 0;
|
||||
const histogram = new Map<string, number>();
|
||||
|
||||
// Fail-closed path for a broken lists.json. A checker that cannot load its
|
||||
// lists must never be read as "everything passed": checks stop, the skip is
|
||||
// logged each turn, and the operator is notified once.
|
||||
let broken: string | null = null;
|
||||
let brokenNotified = false;
|
||||
const reportBroken = (ctx: { hasUI?: boolean } | undefined, where: string) => {
|
||||
log({ ev: "skipped_broken", where, reason: broken });
|
||||
if (!brokenNotified && ctx?.hasUI) {
|
||||
ctx.ui.notify(`unslop gate BROKEN: ${broken}. Fix tools/unslop-hook/lists.json; no clean verdicts until then.`, "error");
|
||||
brokenNotified = true;
|
||||
}
|
||||
};
|
||||
const safeCheck = (text: string): ReturnType<typeof checkText> | null => {
|
||||
if (broken) return null;
|
||||
try {
|
||||
return checkText(text);
|
||||
} catch (e) {
|
||||
broken = String((e as Error).message);
|
||||
log({ ev: "broken_lists", reason: broken });
|
||||
return null;
|
||||
}
|
||||
};
|
||||
|
||||
pi.on("session_start", async (event, _ctx) => {
|
||||
log({ ev: "loaded", reason: event.reason });
|
||||
try {
|
||||
checkText(""); // probe: load+validate lists at startup, not mid-conversation
|
||||
} catch (e) {
|
||||
broken = String((e as Error).message);
|
||||
log({ ev: "broken_lists", reason: broken, at: "startup" });
|
||||
}
|
||||
});
|
||||
|
||||
pi.on("message_end", async (event, ctx) => {
|
||||
if ((event.message as { role?: string }).role !== "assistant") return;
|
||||
const text = assistantText({ type: "message", id: "", message: event.message });
|
||||
if (text === null) return;
|
||||
|
||||
const result = safeCheck(text);
|
||||
if (result === null) {
|
||||
reportBroken(ctx, "message_end");
|
||||
return;
|
||||
}
|
||||
turnsChecked++;
|
||||
if (result.clean) {
|
||||
log({ ev: "checked", clean: true, turn: turnsChecked, charsChecked: result.charsChecked });
|
||||
return;
|
||||
}
|
||||
turnsFlagged++;
|
||||
for (const f of result.findings) histogram.set(f.rule, (histogram.get(f.rule) ?? 0) + 1);
|
||||
const summary = result.findings.map((f) => f.detail).join("; ");
|
||||
if (ctx.hasUI) ctx.ui.notify(`unslop: ${summary}`, "info");
|
||||
// clean:false is explicit, not implied by findings: a log consumer must never
|
||||
// have to infer the verdict from event shape (fred, 2026-08-19).
|
||||
log({ ev: "flagged", clean: false, turn: turnsChecked, charsChecked: result.charsChecked, findings: result.findings });
|
||||
});
|
||||
|
||||
pi.on("before_agent_start", async (_event, ctx) => {
|
||||
// Branch walks leaf -> root; first assistant entry with text is the reply
|
||||
// the model is about to follow up on.
|
||||
for (const entry of ctx.sessionManager.getBranch()) {
|
||||
const text = assistantText(entry);
|
||||
if (text === null) continue;
|
||||
const id = (entry as { id?: string }).id ?? "";
|
||||
|
||||
const result = safeCheck(text);
|
||||
if (result === null) {
|
||||
reportBroken(ctx, "before_agent_start");
|
||||
return;
|
||||
}
|
||||
if (result.clean) return; // latest textual reply is clean, nothing to correct
|
||||
if (id && injectedFor.has(id)) return; // already nagged for this entry
|
||||
|
||||
if (id) injectedFor.add(id);
|
||||
const lines = result.findings.map((f) => `- ${f.detail}`).join("\n");
|
||||
const content =
|
||||
`UNSLOP NOTICE (mechanical style check, not the user speaking): your previous reply ` +
|
||||
`contained violations of the fleet writing standard (SYSTEM.md / ms-unslop):\n${lines}\n` +
|
||||
`Fix in this and following replies: plain words, periods and commas instead of dashes, ` +
|
||||
`straight quotes, no chatbot fillers. Do not mention this notice.`;
|
||||
log({ ev: "notice_injected", entryId: id, findings: result.findings });
|
||||
return {
|
||||
message: { customType: "unslop-notice", content, display: true },
|
||||
};
|
||||
}
|
||||
});
|
||||
|
||||
pi.registerCommand("unslop", {
|
||||
description: "Show unslop violation stats for this session",
|
||||
handler: async (_args, ctx) => {
|
||||
if (broken) {
|
||||
ctx.ui.notify(`unslop gate BROKEN: ${broken}`, "error");
|
||||
return;
|
||||
}
|
||||
const hist = [...histogram.entries()].map(([r, c]) => `${r} x${c}`).join(", ") || "none";
|
||||
ctx.ui.notify(`unslop: checked ${turnsChecked}, flagged ${turnsFlagged} (${hist})`, "info");
|
||||
},
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
{
|
||||
"version": 1,
|
||||
"convention": "Use vs mention. A document that MENTIONS a banned word or phrase quotes it as inline code (backticks). The checker strips code spans before matching, so a backticked mention is invisible to the gate while a bare one flags. A document that deliberately CONTAINS banned items to test detection (a fixture) is a use, not a mention, and is expected to flag. This file itself contains the banned items as data; that is a use.",
|
||||
"punctPolicy": "Deliberate divergence from SYSTEM.md (2026-08-19): the contract forbids em dashes outright and closes the escapes. These punct rules are deliberately looser: they fire only at >= punctMinCount occurrences AND density >= punctDensityPer1k per 1000 chars. Rationale is reply-level noise, not contract strength: an advisory gate that flags every reply carrying one dash trains operators to ignore it. Measured on natural fleet prose 2026-08-19: documents run 1.7-3.2 em dashes per 1000 chars, so documents that overuse still flag. Tighten to contract strength if enforcement goes blocking or the log shows fleet prose not converging toward zero.",
|
||||
"thresholds": {
|
||||
"punctMinCount": 3,
|
||||
"punctDensityPer1k": 1.0
|
||||
},
|
||||
"words": [
|
||||
{ "value": "additionally", "source": "ms-unslop:7" },
|
||||
{ "value": "crucial", "source": "ms-unslop:7" },
|
||||
{ "value": "delve", "source": "ms-unslop:7" },
|
||||
{ "value": "garner", "source": "ms-unslop:7" },
|
||||
{ "value": "interplay", "source": "ms-unslop:7" },
|
||||
{ "value": "intricate", "source": "ms-unslop:7" },
|
||||
{ "value": "pivotal", "source": "ms-unslop:7" },
|
||||
{ "value": "showcase", "source": "ms-unslop:7" },
|
||||
{ "value": "tapestry", "source": "ms-unslop:7" },
|
||||
{ "value": "testament", "source": "ms-unslop:7" },
|
||||
{ "value": "underscore", "source": "ms-unslop:7" },
|
||||
{ "value": "vibrant", "source": "ms-unslop:7" },
|
||||
{ "value": "utilize", "source": "ms-unslop:31" },
|
||||
{ "value": "leverage", "source": "ms-unslop:31" },
|
||||
{ "value": "facilitate", "source": "ms-unslop:31" },
|
||||
{ "value": "load-bearing", "source": "system-md" }
|
||||
],
|
||||
"phrases": [
|
||||
{ "value": "worth stating plainly", "source": "system-md" },
|
||||
{ "value": "here's the honest truth", "source": "system-md" },
|
||||
{ "value": "heres the honest truth", "source": "system-md", "note": "apostrophe-OMITTED renderings only; ASCII and curly-apostrophe forms match the main entry because the checker normalizes U+2019/U+2018 to ASCII before phrase matching" },
|
||||
{ "value": "the real tension", "source": "system-md" },
|
||||
{ "value": "carry the argument", "source": "system-md" },
|
||||
{ "value": "in order to", "source": "ms-unslop:23" },
|
||||
{ "value": "due to the fact that", "source": "ms-unslop:23" },
|
||||
{ "value": "it is important to note", "source": "ms-unslop:23" },
|
||||
{ "value": "i hope this helps", "source": "ms-unslop:20" },
|
||||
{ "value": "let me know if", "source": "ms-unslop:20" },
|
||||
{ "value": "of course!", "source": "ms-unslop:20" },
|
||||
{ "value": "certainly!", "source": "ms-unslop:20" },
|
||||
{ "value": "found the smoking gun", "source": "ms-unslop:20" },
|
||||
{ "value": "happy to help", "source": "ms-unslop:20", "note": "extension of the named pattern set" },
|
||||
{ "value": "great question", "source": "ms-unslop:22" },
|
||||
{ "value": "absolutely right", "source": "ms-unslop:22" },
|
||||
{ "value": "excellent question", "source": "ms-unslop:22", "note": "extension of the named pattern set" }
|
||||
],
|
||||
"punct": [
|
||||
{ "value": "em", "label": "em dash", "chars": ["\u2014"], "source": "ms-unslop:13+system-md" },
|
||||
{ "value": "en", "label": "en dash", "chars": ["\u2013"], "source": "ms-unslop:13" },
|
||||
{ "value": "curly", "label": "curly quote/apostrophe", "chars": ["\u201c", "\u201d", "\u2018", "\u2019"], "source": "ms-unslop:19" }
|
||||
],
|
||||
"patterns": [
|
||||
{ "value": "not-just-but", "regex": "not just\\s+[^.!?]{0,80}?\\s+but", "flags": "gi", "detail": "not just X but Y", "source": "ms-unslop:9" }
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,228 @@
|
||||
"use strict";
|
||||
// Tests for unslop-check.js. Run: node test-unslop-check.js
|
||||
// Exit 0 = all pass. Cases include a red control (slop must fail) and a green
|
||||
// control (clean prose must pass) per evidence discipline.
|
||||
|
||||
const assert = require("node:assert");
|
||||
const fs = require("node:fs");
|
||||
const os = require("node:os");
|
||||
const path = require("node:path");
|
||||
const { spawnSync } = require("node:child_process");
|
||||
const { checkText, stripCode, loadLists } = require("./unslop-check.js");
|
||||
|
||||
const SLOP = `Certainly! Let me delve into the evolving tapestry of database technology — it’s truly “pivotal” — and intricate.
|
||||
In order to understand it — we should leverage this interplay of systems — deeply. I hope this helps!`;
|
||||
|
||||
const CLEAN = `The loader parses the file and validates each row. Rows that fail are logged
|
||||
and skipped. We measured a range from 1 to 10 seconds. Use "straight quotes" and
|
||||
commas, not dashes. That is the whole finding.`;
|
||||
|
||||
// Code-stripping control: banned words inside code must not count.
|
||||
const WITH_CODE = [
|
||||
"The config uses `utilize=false` internally.",
|
||||
"```",
|
||||
"delve tapestry — pivotal",
|
||||
"```",
|
||||
"The config file sets one flag. It is parsed at startup.",
|
||||
].join("\n");
|
||||
|
||||
const results = [];
|
||||
function t(name, fn) {
|
||||
try { fn(); results.push([name, true]); } catch (e) { results.push([name, false]); console.error(`FAIL ${name}: ${e.message}`); }
|
||||
}
|
||||
|
||||
t("slop fixture is flagged (red control)", () => {
|
||||
const r = checkText(SLOP);
|
||||
assert.ok(!r.clean, "slop must not be clean");
|
||||
const details = r.findings.map((f) => f.detail).join("; ");
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("delve")), `delve missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("tapestry")), `tapestry missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("pivotal")), `pivotal missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("em dash")), `em dash missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("curly")), `curly missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("in order to")), `in order to missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("i hope this helps")), `chatbot phrase missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("certainly")), `certainly missing: ${details}`);
|
||||
});
|
||||
|
||||
t("clean fixture passes (green control)", () => {
|
||||
const r = checkText(CLEAN);
|
||||
assert.deepStrictEqual(r.findings, [], `unexpected findings: ${JSON.stringify(r.findings)}`);
|
||||
});
|
||||
|
||||
t("numeric range is not a false range flag", () => {
|
||||
const r = checkText(CLEAN);
|
||||
assert.ok(!r.findings.some((f) => f.rule === "pattern"), "must not flag numeric ranges");
|
||||
});
|
||||
|
||||
t("code blocks and inline code are stripped", () => {
|
||||
const r = checkText(WITH_CODE);
|
||||
assert.deepStrictEqual(r.findings, [], `code leaked into check: ${JSON.stringify(r.findings)}`);
|
||||
});
|
||||
|
||||
t("stripCode removes fenced and inline code", () => {
|
||||
const s = stripCode("a `x — y` b\n```\ndelve\n```\nc");
|
||||
assert.ok(!s.includes("delve"), "fenced code not stripped");
|
||||
assert.ok(!s.includes("—"), "inline code not stripped");
|
||||
assert.ok(s.includes("a") && s.includes("b") && s.includes("c"), "prose lost");
|
||||
});
|
||||
|
||||
t("not-just-but pattern is detected", () => {
|
||||
const r = checkText("This is not just a cache but a coordination layer.");
|
||||
assert.ok(r.findings.some((f) => f.rule === "pattern"), "pattern missed");
|
||||
});
|
||||
|
||||
t("light dash use is not flagged (below threshold)", () => {
|
||||
const prose =
|
||||
"The loader parses each row and validates it against the schema. Rows that fail " +
|
||||
"are logged — with their line numbers — and skipped. The operator reviews the log " +
|
||||
"daily and reconciles the rejects against the source system by hand, which takes " +
|
||||
"a few minutes and has never once produced a discrepancy worth acting on.";
|
||||
const r = checkText(prose);
|
||||
assert.ok(!r.findings.some((f) => f.detail.includes("dash")), "2 dashes in ~330 chars must not flag");
|
||||
});
|
||||
|
||||
t("dash overuse is flagged (above threshold)", () => {
|
||||
const r = checkText("One — two — three — four. That is the whole sentence.");
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("em dash")), "4 dashes in 50 chars must flag");
|
||||
});
|
||||
|
||||
// ── C1: lists.json machine source ──────────────────────────────────────
|
||||
// The committed lists are the single source of truth; these tests pin the
|
||||
// file's validity, its shape, and the loader's fail-closed behavior.
|
||||
|
||||
const tmpdir = fs.mkdtempSync(path.join(os.tmpdir(), "unslop-c1-"));
|
||||
const tmp = (n) => path.join(tmpdir, n);
|
||||
|
||||
function brokenVariant(mutate) {
|
||||
const l = JSON.parse(JSON.stringify(loadLists()));
|
||||
mutate(l);
|
||||
return l;
|
||||
}
|
||||
|
||||
function writeTmp(name, data) {
|
||||
const f = tmp(name);
|
||||
fs.writeFileSync(f, typeof data === "string" ? data : JSON.stringify(data));
|
||||
return f;
|
||||
}
|
||||
|
||||
t("lists.json (the real file) validates and is pinned in size", () => {
|
||||
const l = loadLists();
|
||||
assert.strictEqual(l.version, 1);
|
||||
// Counts pin the migration: 16 words, 17 phrases, 3 punct, 1 pattern moved
|
||||
// from the old inline constants. Changing a count means changing this test
|
||||
// too, consciously.
|
||||
assert.strictEqual(l.words.length, 16, "word count drifted");
|
||||
assert.strictEqual(l.phrases.length, 17, "phrase count drifted");
|
||||
assert.strictEqual(l.punct.length, 3, "punct count drifted");
|
||||
assert.strictEqual(l.patterns.length, 1, "pattern count drifted");
|
||||
assert.ok(l.convention.length > 50, "mention convention must be present");
|
||||
assert.ok(l.punctPolicy.length > 50, "punct divergence policy must be present");
|
||||
for (const e of [...l.words, ...l.phrases, ...l.punct, ...l.patterns]) {
|
||||
assert.ok(e.source && e.source.trim(), `entry missing source: ${JSON.stringify(e)}`);
|
||||
}
|
||||
});
|
||||
|
||||
t("loader rejects an empty file", () => {
|
||||
const f = writeTmp("empty.json", "");
|
||||
assert.throws(() => loadLists(f), /empty file/);
|
||||
});
|
||||
|
||||
t("loader rejects unparseable JSON", () => {
|
||||
const f = writeTmp("bad.json", "{nope");
|
||||
assert.throws(() => loadLists(f), /unparseable/);
|
||||
});
|
||||
|
||||
t("loader rejects a missing file", () => {
|
||||
assert.throws(() => loadLists(tmp("does-not-exist.json")), /cannot read/);
|
||||
});
|
||||
|
||||
t("loader rejects missing keys", () => {
|
||||
const f = writeTmp("nokeys.json", { version: 1 });
|
||||
assert.throws(() => loadLists(f), /missing key/);
|
||||
});
|
||||
|
||||
t("loader rejects an emptied word list", () => {
|
||||
const f = writeTmp("emptywords.json", brokenVariant((l) => { l.words = []; }));
|
||||
assert.throws(() => loadLists(f), /words must be a non-empty array/);
|
||||
});
|
||||
|
||||
t("loader rejects entries without provenance", () => {
|
||||
const f = writeTmp("nosource.json", brokenVariant((l) => { delete l.phrases[0].source; }));
|
||||
assert.throws(() => loadLists(f), /source/);
|
||||
});
|
||||
|
||||
t("loader rejects duplicate values", () => {
|
||||
const f = writeTmp("dup.json", brokenVariant((l) => { l.words.push({ ...l.words[0] }); }));
|
||||
assert.throws(() => loadLists(f), /duplicate/);
|
||||
});
|
||||
|
||||
t("loader rejects a non-compiling pattern regex", () => {
|
||||
const f = writeTmp("badregex.json", brokenVariant((l) => { l.patterns[0].regex = "("; }));
|
||||
assert.throws(() => loadLists(f), /does not compile/);
|
||||
});
|
||||
|
||||
t("CLI exits 2 on broken lists (red control)", () => {
|
||||
const f = writeTmp("cli-broken.json", "");
|
||||
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js")], {
|
||||
input: "some prose",
|
||||
encoding: "utf8",
|
||||
env: { ...process.env, UNSLOP_LISTS: f },
|
||||
});
|
||||
assert.strictEqual(r.status, 2, `expected exit 2, got ${r.status} (stderr: ${r.stderr})`);
|
||||
assert.ok(r.stderr.includes("lists.json invalid"), `stderr must name the fault: ${r.stderr}`);
|
||||
});
|
||||
|
||||
t("CLI honors UNSLOP_LISTS for a valid file (green control)", () => {
|
||||
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js")], {
|
||||
input: "plain prose with no tells at all",
|
||||
encoding: "utf8",
|
||||
env: { ...process.env, UNSLOP_LISTS: path.join(__dirname, "lists.json") },
|
||||
});
|
||||
assert.strictEqual(r.status, 0, `expected exit 0, got ${r.status} (stderr: ${r.stderr})`);
|
||||
});
|
||||
|
||||
// ── Review follow-up (rev-code-01, 2026-08-19): F1, F2, F3, S3 ────────
|
||||
|
||||
t("loader rejects non-finite thresholds (F1)", () => {
|
||||
// Infinity cannot round-trip JSON.stringify, so the fixture is a raw string
|
||||
// edit of the real file — exactly the hand-edit that produced the finding.
|
||||
const real = fs.readFileSync(path.join(__dirname, "lists.json"), "utf8");
|
||||
const f = writeTmp("inf-threshold.json", real.replace('"punctMinCount": 3', '"punctMinCount": 1e999'));
|
||||
assert.ok(real !== fs.readFileSync(f, "utf8") || !real.includes('"punctMinCount": 3'), "fixture mutation did not apply; test is vacuous");
|
||||
assert.throws(() => loadLists(f), /finite/);
|
||||
const f2 = writeTmp("inf-density.json", real.replace('"punctDensityPer1k": 1.0', '"punctDensityPer1k": 1e999'));
|
||||
assert.throws(() => loadLists(f2), /finite/);
|
||||
});
|
||||
|
||||
t("curly-apostrophe phrase rendering is flagged (F2 red control)", () => {
|
||||
const r = checkText("Here\u2019s the honest truth about the deploy.");
|
||||
assert.ok(!r.clean, "curly apostrophe must not defeat phrase matching");
|
||||
assert.ok(r.findings.some((x) => x.detail.includes("here's the honest truth")), `main entry must match, got: ${JSON.stringify(r.findings)}`);
|
||||
});
|
||||
|
||||
t("curly apostrophes still fire the punct rule alongside phrases (F2 ordering)", () => {
|
||||
// Normalization for phrases must not eat the punct signal: four curly
|
||||
// quotes in short text must flag punct, not only the phrase.
|
||||
const r = checkText("It\u2019s \u2019one\u2019 \u2019two\u2019 \u2019three\u2019 \u2019four\u2019 done.");
|
||||
assert.ok(r.findings.some((f) => f.rule === "punct"), `punct must fire on original text: ${JSON.stringify(r.findings)}`);
|
||||
});
|
||||
|
||||
t("loader rejects non-lowercase phrase values (F3)", () => {
|
||||
const f = writeTmp("cap-phrase.json", brokenVariant((l) => { l.phrases[0].value = "Worth Stating Plainly"; }));
|
||||
assert.throws(() => loadLists(f), /lowercase/);
|
||||
});
|
||||
|
||||
t("CLI exits 2 on unreadable input file (S3)", () => {
|
||||
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js"), tmp("definitely-absent.txt")], {
|
||||
encoding: "utf8",
|
||||
});
|
||||
assert.strictEqual(r.status, 2, `expected exit 2, got ${r.status} (stderr: ${r.stderr})`);
|
||||
assert.ok(r.stderr.includes("cannot read input"), `stderr must name the fault: ${r.stderr}`);
|
||||
});
|
||||
|
||||
let failed = 0;
|
||||
for (const [name, ok] of results) { console.log(`${ok ? "PASS" : "FAIL"} ${name}`); if (!ok) failed++; }
|
||||
console.log(`${results.length - failed}/${results.length} passed`);
|
||||
try { fs.rmSync(tmpdir, { recursive: true, force: true }); } catch {}
|
||||
process.exit(failed ? 1 : 0);
|
||||
@@ -0,0 +1,181 @@
|
||||
#!/usr/bin/env node
|
||||
"use strict";
|
||||
// unslop-check — mechanical AI-tell checker (ms-unslop subset + SYSTEM.md phrase bans).
|
||||
// Plain JS, no deps, so pi extensions (jiti) and Claude Code hook scripts (node CLI)
|
||||
// share one implementation.
|
||||
//
|
||||
// The lists live in lists.json beside this file: committed machine source with
|
||||
// per-entry provenance (which ms-unslop pattern or SYSTEM.md rule each entry
|
||||
// mechanizes), the mention convention, and the punct thresholds. The loader
|
||||
// hard-fails closed: an empty, unparseable, or schema-invalid lists.json throws,
|
||||
// and the CLI exits 2 so a broken gate is never mistaken for a clean verdict.
|
||||
//
|
||||
// CLI: node unslop-check.js <file> (or stdin)
|
||||
// exit 0 = clean, exit 1 = violations found (findings printed as JSON),
|
||||
// exit 2 = gate broken (lists.json missing/empty/invalid; error on stderr).
|
||||
// Env: UNSLOP_LISTS=<path> overrides the lists.json location (testing; reuse by
|
||||
// other harnesses sharing this file).
|
||||
//
|
||||
// Provenance note (2026-08-19): the inline lists this file carried before C1
|
||||
// moved to lists.json unchanged — 16 words, 17 phrases, 3 punct rules, 1 pattern.
|
||||
// The suite pins those counts; a list edit without a test edit is a drift signal.
|
||||
|
||||
const fs = require("node:fs");
|
||||
const path = require("node:path");
|
||||
|
||||
function stripCode(text) {
|
||||
// Fenced blocks (``` or ~~~), then inline code spans. Code is quoted material,
|
||||
// not the agent's prose style. This is also the mention convention: a banned
|
||||
// item quoted as inline code is a mention and must not flag (see lists.json).
|
||||
return text
|
||||
.replace(/```[\s\S]*?```/g, " ")
|
||||
.replace(/~~~[\s\S]*?~~~/g, " ")
|
||||
.replace(/`[^`\n]*`/g, " ");
|
||||
}
|
||||
|
||||
// ── lists.json loading and validation ─────────────────────────────────────────
|
||||
|
||||
function validateLists(data) {
|
||||
const fail = (why) => { throw new Error(`lists.json invalid: ${why}`); };
|
||||
if (typeof data !== "object" || data === null || Array.isArray(data)) fail("top level must be an object");
|
||||
for (const k of ["version", "convention", "punctPolicy", "thresholds", "words", "phrases", "punct", "patterns"]) {
|
||||
if (!(k in data)) fail(`missing key: ${k}`);
|
||||
}
|
||||
if (typeof data.version !== "number" || data.version < 1) fail("version must be a number >= 1");
|
||||
for (const k of ["convention", "punctPolicy"]) {
|
||||
if (typeof data[k] !== "string" || !data[k].trim()) fail(`${k} must be a non-empty string`);
|
||||
}
|
||||
const th = data.thresholds;
|
||||
if (typeof th !== "object" || th === null) fail("thresholds must be an object");
|
||||
// Number.isFinite, not just typeof: JSON.parse of 1e999 yields Infinity, which
|
||||
// passes typeof-number and would silently disable the punct gate (review F1).
|
||||
if (!Number.isFinite(th.punctMinCount) || th.punctMinCount < 1) fail("thresholds.punctMinCount must be a finite number >= 1");
|
||||
if (!Number.isFinite(th.punctDensityPer1k) || !(th.punctDensityPer1k > 0)) fail("thresholds.punctDensityPer1k must be a finite number > 0");
|
||||
|
||||
const seen = new Set();
|
||||
const checkEntries = (arr, kind, extra) => {
|
||||
if (!Array.isArray(arr) || arr.length === 0) fail(`${kind} must be a non-empty array`);
|
||||
arr.forEach((e, i) => {
|
||||
const at = `${kind}[${i}]`;
|
||||
if (typeof e !== "object" || e === null) fail(`${at} must be an object`);
|
||||
if (typeof e.value !== "string" || !e.value.trim()) fail(`${at}.value must be a non-empty string`);
|
||||
if (typeof e.source !== "string" || !e.source.trim()) fail(`${at}.source must be a non-empty string (pattern id or system-md)`);
|
||||
if (extra) extra(e, at, fail);
|
||||
if (seen.has(`${kind}:${e.value}`)) fail(`duplicate ${kind} value: ${e.value}`);
|
||||
seen.add(`${kind}:${e.value}`);
|
||||
});
|
||||
};
|
||||
checkEntries(data.words, "words");
|
||||
checkEntries(data.phrases, "phrases", (e, at, fail) => {
|
||||
// Phrase matching splits a lowercased haystack, so an uppercase letter in a
|
||||
// phrase value is a silently dead rule (review F3). Reject, do not silently
|
||||
// normalize: list edits should fail loud (D-a).
|
||||
if (e.value !== e.value.toLowerCase()) fail(`${at}.value must be lowercase; phrase matching lowercases the haystack: ${e.value}`);
|
||||
});
|
||||
checkEntries(data.punct, "punct", (e, at, fail) => {
|
||||
if (typeof e.label !== "string" || !e.label.trim()) fail(`${at}.label must be a non-empty string`);
|
||||
if (!Array.isArray(e.chars) || e.chars.length === 0 || !e.chars.every((c) => typeof c === "string" && c.length === 1)) {
|
||||
fail(`${at}.chars must be a non-empty array of single-char strings`);
|
||||
}
|
||||
});
|
||||
checkEntries(data.patterns, "patterns", (e, at, fail) => {
|
||||
if (typeof e.regex !== "string" || !e.regex.trim()) fail(`${at}.regex must be a non-empty string`);
|
||||
if (typeof e.flags !== "string") fail(`${at}.flags must be a string`);
|
||||
if (typeof e.detail !== "string" || !e.detail.trim()) fail(`${at}.detail must be a non-empty string`);
|
||||
try { new RegExp(e.regex, e.flags); } catch (err) { fail(`${at}.regex does not compile: ${err.message}`); }
|
||||
});
|
||||
return data;
|
||||
}
|
||||
|
||||
let cache = null;
|
||||
function loadLists(filePath) {
|
||||
if (cache && !filePath) return cache;
|
||||
const p = filePath || process.env.UNSLOP_LISTS || path.join(__dirname, "lists.json");
|
||||
let raw;
|
||||
try {
|
||||
raw = fs.readFileSync(p, "utf8");
|
||||
} catch (e) {
|
||||
throw new Error(`lists.json invalid: cannot read ${p}: ${e.message}`);
|
||||
}
|
||||
if (!raw.trim()) throw new Error(`lists.json invalid: empty file: ${p}`);
|
||||
let data;
|
||||
try {
|
||||
data = JSON.parse(raw);
|
||||
} catch (e) {
|
||||
throw new Error(`lists.json invalid: unparseable JSON: ${e.message}`);
|
||||
}
|
||||
const validated = validateLists(data);
|
||||
if (!filePath) cache = validated;
|
||||
return validated;
|
||||
}
|
||||
|
||||
// ── checker ───────────────────────────────────────────────────────────────────
|
||||
|
||||
const escapeRegex = (s) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
||||
|
||||
function checkText(raw) {
|
||||
const lists = loadLists();
|
||||
const text = stripCode(String(raw));
|
||||
// Phrase haystack: lowercased, then curly apostrophes normalized to ASCII.
|
||||
// This must be a SEPARATE string from `text`: punct counting reads the
|
||||
// original, so curly quotes still fire the punct rule (review F2 ordering).
|
||||
const phraseHay = text.toLowerCase().replace(/[\u2018\u2019]/g, "'");
|
||||
const findings = [];
|
||||
|
||||
for (const w of lists.words) {
|
||||
const re = new RegExp("\\b" + escapeRegex(w.value) + "\\b", "gi");
|
||||
const count = (text.match(re) || []).length;
|
||||
if (count > 0) findings.push({ rule: "word", detail: `banned word "${w.value}" x${count}`, count });
|
||||
}
|
||||
|
||||
for (const p of lists.phrases) {
|
||||
const count = phraseHay.split(p.value).length - 1;
|
||||
if (count > 0) findings.push({ rule: "phrase", detail: `phrase "${p.value}" x${count}`, count });
|
||||
}
|
||||
|
||||
// Density-gated punctuation. The deliberate divergence from SYSTEM.md's
|
||||
// outright em-dash ban is documented in lists.json punctPolicy, not only here.
|
||||
for (const pc of lists.punct) {
|
||||
let count = 0;
|
||||
for (const ch of pc.chars) count += text.split(ch).length - 1;
|
||||
if (count < lists.thresholds.punctMinCount) continue;
|
||||
if (count / Math.max(text.length, 1) * 1000 < lists.thresholds.punctDensityPer1k) continue;
|
||||
findings.push({ rule: "punct", detail: `${pc.label} x${count} (density-gated)`, count });
|
||||
}
|
||||
|
||||
for (const pt of lists.patterns) {
|
||||
const flags = pt.flags.includes("g") ? pt.flags : pt.flags + "g";
|
||||
const m = text.match(new RegExp(pt.regex, flags));
|
||||
const count = m ? m.length : 0;
|
||||
if (count > 0) findings.push({ rule: "pattern", detail: `"${pt.detail}" x${count}`, count });
|
||||
}
|
||||
|
||||
return { clean: findings.length === 0, findings, charsChecked: text.length };
|
||||
}
|
||||
|
||||
module.exports = { checkText, stripCode, loadLists, validateLists };
|
||||
|
||||
if (require.main === module) {
|
||||
let input;
|
||||
try {
|
||||
input = process.argv[2] ? fs.readFileSync(process.argv[2], "utf8") : fs.readFileSync(0, "utf8");
|
||||
} catch (e) {
|
||||
// An unreadable input must not exit 1: that is the violations code, and a
|
||||
// wrapper keying on rc alone would report slop-free for a file it never
|
||||
// read (review S3).
|
||||
console.error(`unslop-check: cannot read input: ${e.message}`);
|
||||
process.exit(2);
|
||||
}
|
||||
let result;
|
||||
try {
|
||||
result = checkText(input);
|
||||
} catch (e) {
|
||||
if (String(e.message).startsWith("lists.json invalid")) {
|
||||
console.error(`unslop-check: ${e.message}`);
|
||||
process.exit(2);
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
console.log(JSON.stringify(result, null, 2));
|
||||
process.exit(result.clean ? 0 : 1);
|
||||
}
|
||||
Reference in New Issue
Block a user