Compare commits
10
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
34e06e7de7 | ||
|
|
9bd4f1c405 | ||
|
|
a7b612435b | ||
|
|
87f10772ce | ||
|
|
7db4c5c2ed | ||
|
|
0273a84549 | ||
|
|
3b674b7a66 | ||
|
|
527bc581ca | ||
|
|
1249714a9a | ||
|
|
e175616885 |
@@ -0,0 +1,114 @@
|
||||
# AGENTS.md — Mosaic Stack rebuild (`mosaicstack/stack-v2`)
|
||||
|
||||
Operational context for any agent session working in this repository.
|
||||
Read top to bottom; it is deliberately short — depth lives in the files it
|
||||
points to, not here.
|
||||
|
||||
## What this repository is
|
||||
|
||||
A standalone rebuild of Mosaic Stack: a file-based, fail-closed
|
||||
orchestration foundation that dispatches sandboxed headless pi workers to do
|
||||
real work, with immutable run records as evidence. Thirteen-plus tagged
|
||||
milestones (`git tag -l`) from `poc-container-hello-v0` to today; suites
|
||||
green at every step. Not production software — a proven foundation.
|
||||
|
||||
## Non-negotiable invariants (the canon)
|
||||
|
||||
1. **Root is bootstrap-only.** First-class system configuration lives at the
|
||||
repository root; everything else gets a dedicated directory (`roles/`,
|
||||
`contracts/`, `missions/`, `tasks/`, `docs/`). Do not add new files to root.
|
||||
2. **Configuration**: `~/.config/mosaic-dev/config.json` is the sole system
|
||||
config — created only by `scripts/bootstrap.sh`, never overwritten,
|
||||
fail-closed on any problem. Repo-scoped role authority lives in
|
||||
`roles/*.json` (versioned, reviewed commits only).
|
||||
3. **Secrets** never enter the repository or container images; auth is
|
||||
runtime-only (read-only mount or environment variable).
|
||||
4. **Contracts** (`contracts/`) are immutable and image-baked. Missions and
|
||||
tasks are declarative JSON with strict schemas.
|
||||
5. **Run records** under `<dataRoot>/runs/` are write-once evidence — never
|
||||
rewritten, only pruned via `prune` with a receipt.
|
||||
6. **Fail closed**: missing or invalid config/policy refuses the operation.
|
||||
Never improvise around a refusal; diagnose it.
|
||||
7. **Policy**: missions govern tasks (least-privilege intersection — a task
|
||||
narrows, never widens). Role authority is declared in `roles/` and changes
|
||||
only via reviewed commits.
|
||||
8. **Git**: commit only after suites are green; push only `main`; never
|
||||
force-push. `scripts/conductor-apply.sh` commits locally — push stays an
|
||||
explicit act.
|
||||
9. **Append-only logs**: BUILD-LOG.md (phases), `activation-log.jsonl`,
|
||||
`.pruned.log`, docs/SESSIONS.md. Corrections are new entries, never edits.
|
||||
|
||||
## Session protocol (mandatory)
|
||||
|
||||
- **Register** your session in `docs/SESSIONS.md` — one append-only line
|
||||
(date, actor, scope, outcome). Never rewrite or remove entries.
|
||||
- **Cadence**: read `docs/plans/CURRENT.md` → execute its single next action
|
||||
fully (implement → test → verify against acceptance criteria → commit →
|
||||
push → close issue) → update CURRENT.md → register in SESSIONS.md.
|
||||
- "next" means one action. A batch mandate ("run the queue") repeats the
|
||||
loop until green or blocked. Blocked means stop and report, never improvise.
|
||||
- Substantial work gets a Gitea issue and a BUILD-LOG phase entry
|
||||
(before/after, with corrections recorded honestly).
|
||||
|
||||
## Role model
|
||||
|
||||
- **Conductor**: a system-scoped role — not an agent, not a daemon. Holds
|
||||
git/credentials/policy authority; decomposes, dispatches, reviews,
|
||||
verifies, integrates. Protocol: `docs/plans/CONDUCTOR.md`. Exists only
|
||||
when invoked; push is never automatic.
|
||||
- **Workers**: headless pi via `scripts/run-task.sh` — sandboxed workspace,
|
||||
tools allowlist, optional persistent sessions and forks; no git, no
|
||||
credentials, no policy control.
|
||||
- Worker runs deliberately exclude this file (`--no-context-files` in the
|
||||
adapter): worker context is contracts + mission via the generated system
|
||||
prompt. This file is for conductor-level sessions.
|
||||
|
||||
## Command surface
|
||||
|
||||
`scripts/bootstrap.sh` (idempotent) · `build.sh` · `hello.sh` ·
|
||||
`verify.sh` · `run-task.sh run <task.json>` · `release.sh
|
||||
package|activate|rollback|status` · `reset.sh` (**danger**: wipes the data
|
||||
root; triple-safety-checked) · `mosaic-task.mjs validate|run|show|list|retry|prune` ·
|
||||
`agent.sh <name>` (interactive TUI agent) ·
|
||||
suites: `test-config.sh`, `test-task.sh`, `test-release.sh`,
|
||||
`test-conductor.sh`.
|
||||
|
||||
Full reference — usage, fields, exit codes, safety notes:
|
||||
`docs/TOOLS.md` (read on demand; do not rely on this summary for detail).
|
||||
|
||||
## Data map (canon)
|
||||
|
||||
- `~/.config/mosaic-dev/config.json` — system config (user-authored; never
|
||||
auto-written).
|
||||
- `<dataRoot>` (from config; default `~/.mosaic-dev`):
|
||||
- `runs/` — write-once run evidence (`result.json`, snapshots, `stderr.txt`)
|
||||
- `sessions/` — pi JSONL session trees, one directory per named session
|
||||
- `workspaces/` — agent file effects (persistent or `:run` ephemeral)
|
||||
- `state/` — release pointer + append-only activation/auto-apply logs
|
||||
- Ownership is per-directory; nothing shares state. Directory map and
|
||||
lifecycle rules: README.md "Data map" section.
|
||||
|
||||
## Pointers (depth lives here)
|
||||
|
||||
- `docs/plans/CURRENT.md` — THE next action (single source of "what now")
|
||||
- `docs/plans/CONDUCTOR.md` — orchestration protocol and guardrails
|
||||
- `docs/plans/2026-09-02_atomic-mosaic-foundation.md` — architecture, invariants
|
||||
- `docs/plans/2026-09-03_autonomous-run.md` — batch-run tracker
|
||||
- `BUILD-LOG.md` — append-only build/verification history with corrections
|
||||
- `LAYERS.md` — implemented vs deferred layers
|
||||
- `docs/SESSIONS.md` — session registry
|
||||
- `adapters/README.md` — the harness adapter contract
|
||||
- `roles/` — role contracts (conductor, future agent/coder/reviewer)
|
||||
|
||||
## Recovery rule
|
||||
|
||||
Compacted, restarted, or new? Nothing that matters is lost: this file +
|
||||
`docs/plans/CURRENT.md` + `git log --oneline -10` + the suites reconstruct
|
||||
the full state. **Never guess** — verify with the suites; the run records
|
||||
and logs hold the receipts.
|
||||
|
||||
## Version pin
|
||||
|
||||
`@earendil-works/pi-coding-agent` is pinned exactly (see `package.json` /
|
||||
`RELEASE`); never install unversioned. Release identity: `RELEASE` file
|
||||
(0.0.X until declared stable); image tags derive from it.
|
||||
@@ -303,4 +303,62 @@ M5 tagged `workspace-capabilities-v1`, M6 tagged `sessions-v1`, M7 tagged `opera
|
||||
|
||||
Conductor loop proven end-to-end on the stack itself. `main` merged with M8; release 0.0.6 remains active (retry is host-side only, no image change).
|
||||
|
||||
---
|
||||
|
||||
## Phase 13: Mission-level capability policy (M9)
|
||||
|
||||
### Entry 13.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Missions may declare capabilities.tools as governing constraints (Gitea #30); merge semantics = least-privilege intersection (task narrows, never widens; empty intersection = tool-free run). Host-side only.
|
||||
- Reason: First mechanical restriction layer — the trust model becomes enforced, not instructed.
|
||||
- Expected result: all four merge cases asserted from run evidence; suites green.
|
||||
|
||||
### Entry 13.2 — after
|
||||
|
||||
- Observed: four merge cases verified deterministically via run-record stderr (mission-only, task-only, narrowed, emptied); invalid mission capabilities exit 2; task suite 41/41.
|
||||
- Failure or correction: selftest harness could not express ABSENT vs EMPTY fields via its printf helper — fixed with an ABSENT marker; two suite config-leak defects fixed (per-command env scoping). Product unaffected.
|
||||
|
||||
## Phase 14: Session forking (M11)
|
||||
|
||||
### Entry 14.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: sessionForkFrom task field branches the source session's newest file (pi --fork) into the target session dir; ancestor untouched; RELEASE -> 0.0.7 with health-gated activation (Gitea #33).
|
||||
- Reason: Owner flagged conversation forking from a common ancestor as a desired property; pi JSONL trees make it native.
|
||||
- Expected result: forked child recalls ancestor context; ancestor file untouched; suites green; 0.0.7 active.
|
||||
|
||||
### Entry 14.2 — after
|
||||
|
||||
- Observed: mock plumbing asserts fork source + target delivery; validation rejects fork-without-target and self-fork (exit 2); ghost source exits 4; live fork: child recalled 'mosaico' from ancestor context while the ancestor session file remained untouched (file-level assertion); suites 58/24/14 + verify green; 0.0.7 packaged and activated via health gate.
|
||||
- Failure or correction: retryRun-style self-assignment bug in validation (compared null to target) — caught by negative test, fixed.
|
||||
|
||||
## Result (M11)
|
||||
|
||||
Session forking verified. `main` merged with M11, tagged `session-fork-v1`; release 0.0.7 active.
|
||||
|
||||
---
|
||||
|
||||
## Phase 15: Interactive TUI agent + TOOLS.md (M13)
|
||||
|
||||
### Entry 15.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Add scripts/agent.sh — an interactive TUI launcher (contracts + optional mission + agent identity + named session + optional workspace/tools) — and the pi-adapter interactive branch; remove the fixed compose command; add docs/TOOLS.md as the on-demand reference AGENTS.md routes to; RELEASE -> 0.0.8 (Gitea #35).
|
||||
- Reason: The owner's bootstrap model is vanilla pi sessions directed by AGENTS.md, graduating to governed TUI agents — the first the system itself launches.
|
||||
- Expected result: TUI agent launches with contracts+identity context; headless paths unchanged; TOOLS.md consolidates the reference.
|
||||
|
||||
### Entry 15.2 — after
|
||||
|
||||
- Observed: mock plumbing asserts agent name/session/workspace/mission delivery; identity section asserted in generated prompt; headless hello + suites green (24/58/17/14 + verify); 0.0.8 packaged and health-gated activated.
|
||||
- Failure or correction:
|
||||
1. Regression: pi adapter rewrite made MOSAIC_AGENT_NAME unconditionally required, breaking headless paths — caught by task suite (empty-stderr exit-nonzero), fixed (optional in headless; identity section simply omitted).
|
||||
2. Regression: unquoted $REQUEST_ARG word-split the request into positional args — fixed with positional-argument building (set -- ... "$@").
|
||||
3. Mission fixture wording (objective named the agent) invited the model to append its name after the marker, tripping the strict gate — fixture tightened; strict gate kept by design.
|
||||
- Conductor session env hygiene: sandbox config exports now scoped per-command after a leak broke cross-suite runs.
|
||||
|
||||
## Result (M13)
|
||||
|
||||
Interactive TUI agent launched and verified; TOOLS.md reference shipped. `main` merged with M13, tagged `interactive-agent-v1`; release 0.0.8 active.
|
||||
|
||||
|
||||
|
||||
@@ -6,7 +6,9 @@
|
||||
set -eu
|
||||
|
||||
[ -n "${MOSAIC_SYSTEM_PROMPT_FILE:-}" ] || { echo "mock adapter: MOSAIC_SYSTEM_PROMPT_FILE is required" >&2; exit 2; }
|
||||
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "mock adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
|
||||
if [ "${MOSAIC_INTERACTIVE:-}" != "1" ]; then
|
||||
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "mock adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
|
||||
fi
|
||||
[ -r "$MOSAIC_SYSTEM_PROMPT_FILE" ] || { echo "mock adapter: system prompt not readable: $MOSAIC_SYSTEM_PROMPT_FILE" >&2; exit 2; }
|
||||
|
||||
echo "mock adapter: responding verbatim from MOSAIC_MOCK_RESPONSE" >&2
|
||||
|
||||
+30
-7
@@ -2,16 +2,24 @@
|
||||
# Pi adapter: implements the Mosaic adapter contract for the pinned
|
||||
# @earendil-works/pi-coding-agent CLI.
|
||||
#
|
||||
# Contract: see /opt/mosaic/adapters/README.md. stdout = response only.
|
||||
# Contract: see /opt/mosaic/adapters/README.md.
|
||||
# Headless (default): stdout = response only; stderr = diagnostics; exit 0.
|
||||
# Interactive (MOSAIC_INTERACTIVE=1): full pi TUI on the attached terminal.
|
||||
set -eu
|
||||
|
||||
[ -n "${MOSAIC_SYSTEM_PROMPT_FILE:-}" ] || { echo "pi adapter: MOSAIC_SYSTEM_PROMPT_FILE is required" >&2; exit 2; }
|
||||
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "pi adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
|
||||
[ -r "$MOSAIC_SYSTEM_PROMPT_FILE" ] || { echo "pi adapter: system prompt not readable: $MOSAIC_SYSTEM_PROMPT_FILE" >&2; exit 2; }
|
||||
# MOSAIC_AGENT_NAME is optional in headless mode (identity section is then
|
||||
# omitted); interactive launches always set it via scripts/agent.sh.
|
||||
|
||||
: "${PI_PROVIDER:?pi adapter: PI_PROVIDER is required}"
|
||||
: "${PI_MODEL:?pi adapter: PI_MODEL is required}"
|
||||
|
||||
INTERACTIVE="${MOSAIC_INTERACTIVE:-}"
|
||||
if [ "$INTERACTIVE" != "1" ]; then
|
||||
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "pi adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
|
||||
fi
|
||||
|
||||
# Workspace (M5): run inside the provided workspace when present.
|
||||
if [ -n "${MOSAIC_WORKSPACE:-}" ]; then
|
||||
mkdir -p "$MOSAIC_WORKSPACE"
|
||||
@@ -39,13 +47,25 @@ fi
|
||||
TOOLS_FLAG="--no-tools"
|
||||
[ -n "${MOSAIC_TOOLS:-}" ] && TOOLS_FLAG="--tools $MOSAIC_TOOLS"
|
||||
|
||||
# Mode (M13): interactive TUI or one-shot print.
|
||||
PRINT_MODE="-p"
|
||||
REQUEST_ARG=""
|
||||
if [ "$INTERACTIVE" = "1" ]; then
|
||||
PRINT_MODE=""
|
||||
else
|
||||
REQUEST_ARG="$MOSAIC_REQUEST"
|
||||
fi
|
||||
|
||||
# All flags documented in the pi package README (CLI Reference):
|
||||
# -p/--print noninteractive: print the response and exit
|
||||
# -p/--print one-shot mode: print the response and exit (omitted in
|
||||
# interactive TUI mode)
|
||||
# --system-prompt replace the default prompt with the generated one
|
||||
# --no-* no ambient context/skills/extensions/templates/themes
|
||||
# --no-session ephemeral; TOOLS_FLAG per capabilities
|
||||
# SESSION_FLAGS ephemeral | persistent | forked (per env)
|
||||
# TOOLS_FLAG per capabilities
|
||||
# --offline no startup network operations (update checks/telemetry)
|
||||
exec pi \
|
||||
PROMPT_CONTENT="$(cat "$MOSAIC_SYSTEM_PROMPT_FILE")"
|
||||
set -- \
|
||||
--offline \
|
||||
--no-extensions \
|
||||
--no-skills \
|
||||
@@ -56,5 +76,8 @@ exec pi \
|
||||
$SESSION_FLAGS \
|
||||
--provider "$PI_PROVIDER" \
|
||||
--model "$PI_MODEL" \
|
||||
--system-prompt "$(cat "$MOSAIC_SYSTEM_PROMPT_FILE")" \
|
||||
-p "$MOSAIC_REQUEST"
|
||||
--system-prompt "$PROMPT_CONTENT"
|
||||
# One-shot mode appends -p and the request (both safely quoted);
|
||||
# interactive mode appends nothing - clean TUI.
|
||||
[ "$INTERACTIVE" = "1" ] || set -- "$@" -p "$MOSAIC_REQUEST"
|
||||
exec pi "$@"
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
# SOUL - researcher
|
||||
|
||||
You are the researcher seat of the Mosaic fleet. You are curious, methodical,
|
||||
and precise. You cite what you know, admit what you do not, and never guess
|
||||
when you can verify.
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"agentVersion": 1,
|
||||
"name": "researcher",
|
||||
"role": "researcher",
|
||||
"capabilities": { "tools": ["read", "bash"] }
|
||||
}
|
||||
+8
-3
@@ -21,6 +21,11 @@ services:
|
||||
# Persistent named session dir + optional fork source (M6/M11)
|
||||
MOSAIC_SESSION_DIR: ${MOSAIC_SESSION_DIR:-}
|
||||
MOSAIC_SESSION_FORK: ${MOSAIC_SESSION_FORK:-}
|
||||
# Interactive TUI mode + agent identity (M13, set by scripts/agent.sh)
|
||||
MOSAIC_INTERACTIVE: ${MOSAIC_INTERACTIVE:-}
|
||||
MOSAIC_AGENT_NAME: ${MOSAIC_AGENT_NAME:-}
|
||||
MOSAIC_AGENT_ROLE: ${MOSAIC_AGENT_ROLE:-}
|
||||
MOSAIC_AGENT_SOUL_FILE: ${MOSAIC_AGENT_SOUL_FILE:-}
|
||||
# mock adapter only: verbatim response for deterministic seam tests
|
||||
MOSAIC_MOCK_RESPONSE: ${MOSAIC_MOCK_RESPONSE:-}
|
||||
# Documented container auth alternative: provider API key via
|
||||
@@ -34,6 +39,6 @@ services:
|
||||
# Runtime credential only: pi auth file mounted READ-ONLY.
|
||||
# Never copied into the image.
|
||||
- ${PI_AUTH_FILE:-/home/jwoltje/.pi/agent/auth.json}:/home/node/.pi/agent/auth.json:ro
|
||||
# One-shot: the exact startup verification request. It deliberately
|
||||
# does NOT contain the expected marker MOSAIC_HELLO_OK.
|
||||
command: ["Return your startup marker and nothing else."]
|
||||
# Headless runs: the request is passed as command args by the launchers
|
||||
# (run-task.sh) or defaults inside run-agent.sh (hello/verify). Never a
|
||||
# fixed command here - interactive runs (scripts/agent.sh) need no args.
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
# Session registry — append-only
|
||||
|
||||
Every agent session (assistant, worker-cycle conductor, or owner-directed
|
||||
automation) that works in this repository registers one line here. Entries
|
||||
are never rewritten or removed; corrections are new entries.
|
||||
|
||||
| Date (UTC) | Actor | Scope | Outcome / artifacts |
|
||||
|---|---|---|---|
|
||||
| 2026-09-03 | assistant (conductor + worker) | POC through M12: containerized pi proof, config layer, missions/tasks, release model, adapter seam, workspaces/capabilities, named sessions, retention, session forking, conductor auto-apply, roles/ convention | 13 tags; suites 24/58/14 + 17 conductor + verify green; releases 0.0.1–0.0.7; issues #1–#34 closed |
|
||||
@@ -0,0 +1,85 @@
|
||||
# TOOLS.md — command and tool reference
|
||||
|
||||
On-demand reference for agent sessions (conductors, bootstrapping agents,
|
||||
reviewers). `AGENTS.md` routes here; this file carries the depth: usage,
|
||||
inputs/outputs, exit codes, and safety notes for every entry point.
|
||||
|
||||
Reading guide: all entry points are `scripts/*.sh` (bash) or invoked via
|
||||
`node scripts/mosaic-task.mjs` (node). Every script fails closed — missing
|
||||
or invalid configuration/policy refuses the operation with a nonzero exit
|
||||
and changes nothing.
|
||||
|
||||
## Lifecycle
|
||||
|
||||
| Command | Purpose | Notes |
|
||||
|---|---|---|
|
||||
| `scripts/bootstrap.sh` | Create `~/.config/mosaic-dev/config.json` if absent | Idempotent; existing config validated, never rewritten |
|
||||
| `scripts/build.sh` | Build the release image | Tag derived from `RELEASE` + pinned pi version |
|
||||
| `scripts/hello.sh` | One-shot startup request | Prints model response on stdout |
|
||||
| `scripts/verify.sh` | Full gated test | Exit 0 only on exact `MOSAIC_HELLO_OK`; `EXPECTED_MARKER` overrides for negative drills |
|
||||
|
||||
## Tasks (missions, runs, evidence)
|
||||
|
||||
| Command | Purpose | Notes |
|
||||
|---|---|---|
|
||||
| `scripts/run-task.sh run <task.json>` | Execute a task | Immutable run record under `<dataRoot>/runs/` |
|
||||
| `scripts/run-task.sh validate <task.json>` | Strict validation | Writes nothing |
|
||||
| `node scripts/mosaic-task.mjs show <runId>` | Inspect a run | Full record + snapshots + artifacts |
|
||||
| `node scripts/mosaic-task.mjs list` | List runs | task/workspace/session columns |
|
||||
| `node scripts/mosaic-task.mjs retry <runId>` | Re-execute a run's snapshot | New run dir; `retriedFrom` lineage recorded |
|
||||
| `node scripts/mosaic-task.mjs prune [--keep=N] [--yes]` | Retention | Dry-run default; receipt in `runs/.pruned.log` |
|
||||
|
||||
Task fields: `prompt` (required), `mission` (path), `expectExact`,
|
||||
`timeoutSeconds` (5–600), `workspace` (`:run` or named), `capabilities.tools`
|
||||
(allowlist: read write edit bash grep find ls), `session`,
|
||||
`sessionForkFrom` (requires `session`). Mission fields: `objective`,
|
||||
`directives[]`, optional governing `capabilities.tools`. Policy: a task may
|
||||
narrow a mission's tools, never widen; empty intersection = tool-free run.
|
||||
|
||||
## Agent (interactive TUI)
|
||||
|
||||
```bash
|
||||
scripts/agent.sh <name> [--mission <file>] [--workspace <ws>] [--session <s>] [--tools <list>]
|
||||
```
|
||||
|
||||
Launches an interactive pi TUI inside the container with the four immutable
|
||||
contracts + optional mission + agent identity as its system prompt,
|
||||
persistent named session, optional workspace. Exit with `/quit`.
|
||||
|
||||
## Release
|
||||
|
||||
| Command | Purpose | Notes |
|
||||
|---|---|---|
|
||||
| `scripts/release.sh package` | Build + tag the release image | Tag: `mosaic-poc-agent:<pi>-r<release>` |
|
||||
| `scripts/release.sh activate` | Health gate → atomic pointer swap | `--fault-injection` proves the refusal path |
|
||||
| `scripts/release.sh rollback` | Health-gated return to previous | Refuses if image missing |
|
||||
| `scripts/release.sh status` | Release, tag, active pointer, log | Safe on empty state |
|
||||
|
||||
## Conductor (worker patches)
|
||||
|
||||
```bash
|
||||
scripts/conductor-apply.sh <runId> [--dry-run]
|
||||
```
|
||||
|
||||
Auto-applies a worker's patch under `roles/conductor-policy.json`:
|
||||
succeeded run → clean target tree → path allowlist → syntax gates →
|
||||
apply → policy suites → attribution commit. Any failure reverts.
|
||||
Push is never automatic.
|
||||
|
||||
## Maintenance
|
||||
|
||||
| Command | Purpose | Notes |
|
||||
|---|---|---|
|
||||
| `scripts/reset.sh` | Delete the data root | Triple-safety-checked (path, symlink, ownership marker) |
|
||||
| `scripts/test-config.sh` | Config selftests (no Docker) | 24 cases |
|
||||
| `scripts/test-task.sh` | Task selftests + live cases | 58 cases |
|
||||
| `scripts/test-release.sh` | Release selftests | 14 cases |
|
||||
| `scripts/test-conductor.sh` | Auto-apply selftests (sandboxed) | 17 cases |
|
||||
| `scripts/gitea-api.sh <METHOD> <path> [body]` | Gitea API helper | Token never on argv/stdout |
|
||||
|
||||
## Exit-code convention
|
||||
|
||||
`0` success · `1` operation failed · `2` invalid data/configuration ·
|
||||
`3` configuration missing for a read operation · `4` usage/file/environment
|
||||
problem. Scripts print diagnostics on stderr; model responses (and only
|
||||
model responses) on stdout.
|
||||
@@ -2,6 +2,13 @@
|
||||
|
||||
How the stack orchestrates headless pi workers to do work on itself.
|
||||
|
||||
## Role contracts
|
||||
|
||||
Role authority is declared in role contracts, one file per role, under
|
||||
`roles/` (e.g. `roles/conductor-policy.json`). The repository root holds
|
||||
only first-class, bootstrap-required configuration; role contracts are
|
||||
tracked, versioned files whose changes arrive as reviewed commits.
|
||||
|
||||
## Roles
|
||||
|
||||
| Role | Runs where | Powers | Never has |
|
||||
|
||||
+19
-2
@@ -7,13 +7,14 @@ update this file to the next action). No ambiguity, no re-planning.
|
||||
|
||||
## Next action
|
||||
|
||||
Owner review of M11 (session forking) — then name the next target.
|
||||
Owner review of M13 (interactive TUI agent + TOOLS.md) — then name the next target.
|
||||
|
||||
## Queue (ordered, not started)
|
||||
|
||||
1. Second real adapter (parked — owner focused on Pi)
|
||||
2. Auto-apply policy for worker patches (substrate exists)
|
||||
2. Auto-apply policy for worker patches — SHIPPED in M12
|
||||
3. UX/DX backlog #31 (grep highlight vs status colors — likely closed by owner's unalias)
|
||||
4. Push policy decision: auto-apply commits locally; push remains explicit (documented in CONDUCTOR.md)
|
||||
|
||||
## Rules
|
||||
|
||||
@@ -27,6 +28,22 @@ Owner review of M11 (session forking) — then name the next target.
|
||||
|
||||
## Completed log
|
||||
|
||||
- 2026-09-03 — M12 conductor auto-apply policy (#34) — committed on main, 17/24/58/14 + verify green; push stays explicit
|
||||
- 2026-09-03 — M13 interactive TUI agent + TOOLS.md (#35) — merged, agent.sh TUI launcher + identity injection, 24/58+/17/14 + verify green; release 0.0.8 activated
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
|
||||
- 2026-09-03 — M12 conductor auto-apply policy (#34) — committed on main, 17/24/58/14 + verify green; push stays explicit
|
||||
- 2026-09-03 — repository convention: role contracts move to roles/ (root = bootstrap-only, per owner direction)
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
|
||||
- 2026-09-03 — M12 conductor auto-apply policy (#34) — committed on main, 17/24/58/14 + verify green; push stays explicit
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
|
||||
+1
-1
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"missionVersion": 1,
|
||||
"id": "m-hello",
|
||||
"objective": "Prove the startup marker path of the mosaic-poc-agent.",
|
||||
"objective": "Verify the startup marker path.",
|
||||
"directives": [
|
||||
"Startup verification requests are answered with the marker only.",
|
||||
"No explanation, no formatting."
|
||||
|
||||
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"policyVersion": 1,
|
||||
"autoApply": {
|
||||
"enabled": true,
|
||||
"allowedPaths": [
|
||||
"scripts/**",
|
||||
"docs/**",
|
||||
"tasks/**",
|
||||
"missions/**",
|
||||
"adapters/**",
|
||||
"README.md"
|
||||
],
|
||||
"suites": [
|
||||
"test-config",
|
||||
"test-task",
|
||||
"test-release"
|
||||
]
|
||||
}
|
||||
}
|
||||
Executable
+118
@@ -0,0 +1,118 @@
|
||||
#!/usr/bin/env bash
|
||||
# Launch an interactive (TUI) Mosaic agent in its container.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/agent.sh <name> [--mission <file>] [--workspace <ws>]
|
||||
# [--session <name>] [--tools <comma,list>]
|
||||
#
|
||||
# The agent receives the four immutable contracts (constitution, standards,
|
||||
# SOUL, USER) plus its own identity and optional mission directives as its
|
||||
# system prompt, a persistent named session, and - if declared - a
|
||||
# workspace and tool capabilities. The TUI opens clean; you drive.
|
||||
#
|
||||
# This is the Mosaic alternative to launching vanilla pi: same engine,
|
||||
# governed context.
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
# shellcheck source=common.sh
|
||||
source scripts/common.sh
|
||||
|
||||
NAME=""
|
||||
MISSION=""
|
||||
WORKSPACE=""
|
||||
SESSION=""
|
||||
TOOLS=""
|
||||
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in
|
||||
--mission) MISSION="${2:?}"; shift 2 ;;
|
||||
--workspace) WORKSPACE="${2:?}"; shift 2 ;;
|
||||
--session) SESSION="${2:?}"; shift 2 ;;
|
||||
--tools) TOOLS="${2:?}"; shift 2 ;;
|
||||
--help|-h) sed -n '2,12p' "$0"; exit 0 ;;
|
||||
*) NAME="$1"; shift ;;
|
||||
esac
|
||||
done
|
||||
|
||||
[ -n "$NAME" ] || { echo "agent: usage: scripts/agent.sh <name> [--mission f] [--workspace ws] [--session s] [--tools list]" >&2; exit 4; }
|
||||
case "$NAME" in *[!A-Za-z0-9._-]*|'') echo "agent: invalid agent name" >&2; exit 4;; esac
|
||||
|
||||
load_config
|
||||
load_release
|
||||
bootstrap_runtime_dir
|
||||
|
||||
# Agent seat definition (M15): when agents/<name>/agent.json exists it is
|
||||
# strictly validated and its values become defaults (CLI flags override).
|
||||
# The seat's SOUL.md overrides the contract persona; governance contracts
|
||||
# are never overridden.
|
||||
AGENTS_DIR="${MOSAIC_AGENTS_DIR:-agents}"
|
||||
ROLE=""
|
||||
DEFCAPS=""
|
||||
if [ -f "$AGENTS_DIR/$NAME/agent.json" ]; then
|
||||
DEFAULTS_FILE="$(mktemp)"
|
||||
node -e '
|
||||
const fs = require("fs");
|
||||
const p = JSON.parse(fs.readFileSync(process.argv[1], "utf8"));
|
||||
if (p.agentVersion !== 1) process.exit(2);
|
||||
const ID = /^[a-z0-9][a-z0-9._-]{0,63}$/;
|
||||
if (typeof p.name !== "string" || !ID.test(p.name)) process.exit(2);
|
||||
if (p.role !== undefined && (typeof p.role !== "string" || !ID.test(p.role))) process.exit(2);
|
||||
let tools = "";
|
||||
if (p.capabilities !== undefined) {
|
||||
if (typeof p.capabilities !== "object" || p.capabilities === null || Array.isArray(p.capabilities)) process.exit(2);
|
||||
for (const k of Object.keys(p.capabilities)) if (k !== "tools") process.exit(2);
|
||||
if (!Array.isArray(p.capabilities.tools) || p.capabilities.tools.some(t => !/^[a-z]+$/.test(t))) process.exit(2);
|
||||
tools = p.capabilities.tools.join(",");
|
||||
}
|
||||
fs.writeFileSync(process.argv[2], "AGENT_DEF_ROLE=" + (p.role || "") + "\nAGENT_DEF_CAPS=" + tools + "\n");
|
||||
' "$AGENTS_DIR/$NAME/agent.json" "$DEFAULTS_FILE" || { rm -f "$DEFAULTS_FILE"; echo "agent: invalid agent definition" >&2; exit 2; }
|
||||
AGENT_DEF_ROLE=""; AGENT_DEF_CAPS=""
|
||||
while IFS= read -r line; do
|
||||
case "$line" in
|
||||
AGENT_DEF_ROLE=*) AGENT_DEF_ROLE="${line#AGENT_DEF_ROLE=}" ;;
|
||||
AGENT_DEF_CAPS=*) AGENT_DEF_CAPS="${line#AGENT_DEF_CAPS=}" ;;
|
||||
esac
|
||||
done < "$DEFAULTS_FILE"
|
||||
rm -f "$DEFAULTS_FILE"
|
||||
ROLE="$AGENT_DEF_ROLE"
|
||||
DEFCAPS="$AGENT_DEF_CAPS"
|
||||
[ -r "$AGENTS_DIR/$NAME/SOUL.md" ] || { echo "agent: definition dir missing SOUL.md: $AGENTS_DIR/$NAME" >&2; exit 4; }
|
||||
mkdir -p "$MOSAIC_DEV_DIR/agents/$NAME"
|
||||
cp "$AGENTS_DIR/$NAME/SOUL.md" "$MOSAIC_DEV_DIR/agents/$NAME/SOUL.md"
|
||||
export MOSAIC_AGENT_SOUL_FILE="/var/lib/mosaic/agents/$NAME/SOUL.md"
|
||||
# Seat record: written once at instantiation.
|
||||
SEAT="$MOSAIC_DEV_DIR/agents/$NAME/seat.json"
|
||||
if [ ! -f "$SEAT" ]; then
|
||||
printf '{"seatVersion":1,"name":"%s","role":"%s","instantiatedAt":"%s"}\n' \
|
||||
"$NAME" "$ROLE" "$(date -u +%Y-%m-%dT%H:%M:%SZ)" > "$SEAT"
|
||||
fi
|
||||
fi
|
||||
|
||||
SESSION="${SESSION:-agent-$NAME}"
|
||||
mkdir -p "$MOSAIC_DEV_DIR/sessions/$SESSION"
|
||||
export MOSAIC_SESSION_DIR="/var/lib/mosaic/sessions/$SESSION"
|
||||
export MOSAIC_AGENT_NAME="$NAME"
|
||||
[ -n "$ROLE" ] && export MOSAIC_AGENT_ROLE="$ROLE"
|
||||
export MOSAIC_INTERACTIVE=1
|
||||
if [ -z "$TOOLS" ] && [ -n "$DEFCAPS" ]; then TOOLS="$DEFCAPS"; fi
|
||||
export MOSAIC_TOOLS="${TOOLS:+$TOOLS}"
|
||||
|
||||
if [ -n "$MISSION" ]; then
|
||||
[ -r "$MISSION" ] || { echo "agent: mission file not readable: $MISSION" >&2; exit 4; }
|
||||
mkdir -p "$MOSAIC_DEV_DIR/agent-missions"
|
||||
cp "$MISSION" "$MOSAIC_DEV_DIR/agent-missions/$NAME.json"
|
||||
export MOSAIC_MISSION_FILE="/var/lib/mosaic/agent-missions/$NAME.json"
|
||||
fi
|
||||
|
||||
# Workspace (M13): defaults to a persistent per-agent workspace
|
||||
# (workspaces/<agent>) so the agent has a real, host-visible home instead
|
||||
# of the container's neutral /workspace. Override with --workspace <ws>.
|
||||
[ -n "$WORKSPACE" ] || WORKSPACE="$NAME"
|
||||
case "$WORKSPACE" in *[!A-Za-z0-9._-]*|'') echo "agent: invalid workspace name" >&2; exit 4;; esac
|
||||
mkdir -p "$MOSAIC_DEV_DIR/workspaces/$WORKSPACE"
|
||||
export MOSAIC_WORKSPACE="/var/lib/mosaic/workspaces/$WORKSPACE"
|
||||
|
||||
echo "agent: launching TUI agent '$NAME' (session: $SESSION, adapter: $MOSAIC_ADAPTER, model: $MOSAIC_MODEL)"
|
||||
echo "agent: contracts + $([ -n "$MISSION" ] && echo 'mission' || echo 'no mission') loaded; exit the TUI with /quit"
|
||||
# No -T: the TTY is the point. Ctrl+C twice or /quit exits.
|
||||
exec docker compose run --rm mosaic-agent
|
||||
@@ -50,4 +50,11 @@ bootstrap_runtime_dir() {
|
||||
echo "bootstrap: created $MOSAIC_DEV_DIR"
|
||||
fi
|
||||
touch "$MOSAIC_DEV_DIR/$POC_ROOT_MARKER"
|
||||
# Live user context layer (M14): seeded once, owned by the user from
|
||||
# then on; dispatched to every agent launch without rebuilds.
|
||||
mkdir -p "$MOSAIC_DEV_DIR/user"
|
||||
if [ ! -f "$MOSAIC_DEV_DIR/user/USER.md" ]; then
|
||||
printf '# User\n\nDescribe yourself, your machine, and your preferences here.\nThis file is dispatched to every Mosaic agent launch.\n' \
|
||||
> "$MOSAIC_DEV_DIR/user/USER.md"
|
||||
fi
|
||||
}
|
||||
|
||||
Executable
+139
@@ -0,0 +1,139 @@
|
||||
#!/usr/bin/env bash
|
||||
# Conductor auto-apply: integrate a worker's patch under the declared policy.
|
||||
#
|
||||
# Usage: scripts/conductor-apply.sh <runId> [--dry-run]
|
||||
#
|
||||
# Policy (roles/conductor-policy.json in the target repo, strictly validated):
|
||||
# autoApply.enabled master switch
|
||||
# autoApply.allowedPaths glob allowlist ('dir/**' = everything under dir)
|
||||
# autoApply.suites suite scripts that must pass AFTER applying
|
||||
#
|
||||
# Gate sequence: succeeded run record -> clean target tree -> diff extracted
|
||||
# from the worker workspace -> allowlist -> syntax gates (node/bash/json) ->
|
||||
# apply -> policy suites -> commit with attribution. ANY failure reverts the
|
||||
# working tree and exits nonzero. Push is never automatic.
|
||||
#
|
||||
# Environment:
|
||||
# MOSAIC_APPLY_TARGET repo root to apply into (default: this project)
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
TARGET_ROOT="${MOSAIC_APPLY_TARGET:-$(cd "$SCRIPT_DIR/.." && pwd)}"
|
||||
RUN_ID="${1:?usage: conductor-apply.sh <runId> [--dry-run]}"
|
||||
DRY_RUN="no"
|
||||
[ "${2:-}" = "--dry-run" ] && DRY_RUN="yes"
|
||||
|
||||
cd "$TARGET_ROOT"
|
||||
|
||||
fail() { echo "conductor-apply: $*" >&2; exit "${2:-1}"; }
|
||||
[ -d .git ] || fail "target is not a git repository: $TARGET_ROOT" 4
|
||||
[ -f roles/conductor-policy.json ] || fail "no roles/conductor-policy.json in target" 2
|
||||
|
||||
# ---- policy (strict) ----
|
||||
POLICY_JSON="$(node -e '
|
||||
const fs = require("fs");
|
||||
const p = JSON.parse(fs.readFileSync("roles/conductor-policy.json", "utf8"));
|
||||
if (p.policyVersion !== 1) process.exit(3);
|
||||
if (!p.autoApply || typeof p.autoApply.enabled !== "boolean" || !Array.isArray(p.autoApply.allowedPaths) || !Array.isArray(p.autoApply.suites)) process.exit(3);
|
||||
for (const g of p.autoApply.allowedPaths) {
|
||||
if (typeof g !== "string" || !/^[A-Za-z0-9_.*/-]+$/.test(g) || g.startsWith("/") || g.includes("..")) process.exit(3);
|
||||
}
|
||||
console.log(JSON.stringify(p.autoApply));
|
||||
')" || fail "invalid roles/conductor-policy.json" 2
|
||||
|
||||
ENABLED="$(node -e 'console.log(JSON.parse(process.argv[1]).enabled)' "$POLICY_JSON")"
|
||||
[ "$ENABLED" = "true" ] || fail "auto-apply is disabled by policy" 2
|
||||
|
||||
# ---- run record ----
|
||||
DATA_ROOT="$(node scripts/mosaic-config.mjs validate | node -e 'let d="";process.stdin.on("data",c=>d+=c).on("end",()=>console.log(JSON.parse(d).dataRoot))')"
|
||||
RESULT_FILE="$DATA_ROOT/runs/$RUN_ID/result.json"
|
||||
[ -f "$RESULT_FILE" ] || fail "run not found: $RUN_ID" 4
|
||||
|
||||
node -e '
|
||||
const r = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
|
||||
process.exit(r.status === "succeeded" ? 0 : 1);
|
||||
' "$RESULT_FILE" || fail "run $RUN_ID did not succeed; refusing to auto-apply"
|
||||
|
||||
WORKSPACE_NAME="$(node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.workspace||"")' "$RESULT_FILE")"
|
||||
[ -n "$WORKSPACE_NAME" ] || fail "run has no workspace; nothing to integrate" 4
|
||||
WORKSPACE="$DATA_ROOT/workspaces/$WORKSPACE_NAME"
|
||||
[ -d "$WORKSPACE/.git" ] || fail "workspace is not a git clone: $WORKSPACE" 4
|
||||
|
||||
# ---- extract diff (tracked + intent-to-add) ----
|
||||
git -C "$WORKSPACE" add -N . >/dev/null 2>&1 || true
|
||||
DIFF_FILE="$(mktemp)"
|
||||
trap 'rm -f "$DIFF_FILE"' EXIT
|
||||
git -C "$WORKSPACE" diff > "$DIFF_FILE"
|
||||
if [ ! -s "$DIFF_FILE" ]; then
|
||||
fail "workspace has no changes to apply"
|
||||
fi
|
||||
|
||||
# ---- allowlist ----
|
||||
mapfile -t CHANGED < <(git -C "$WORKSPACE" diff --name-only)
|
||||
GLOBS="$(node -e 'const a=JSON.parse(process.argv[1]).allowedPaths;console.log(a.join("\n"))' "$POLICY_JSON")"
|
||||
REFUSED=""
|
||||
for f in "${CHANGED[@]}"; do
|
||||
ok="no"
|
||||
while IFS= read -r g; do
|
||||
[ -z "$g" ] && continue
|
||||
case "$f" in
|
||||
$g) ok="yes"; break ;;
|
||||
esac
|
||||
done <<< "$GLOBS"
|
||||
[ "$ok" = "yes" ] || REFUSED="$REFUSED $f"
|
||||
done
|
||||
if [ -n "$REFUSED" ]; then
|
||||
echo "conductor-apply: refusing - files outside policy allowlist:$REFUSED" >&2
|
||||
echo "conductor-apply: patch preserved at $DIFF_FILE for manual review" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ---- syntax gates (on workspace files, pre-apply) ----
|
||||
for f in "${CHANGED[@]}"; do
|
||||
case "$f" in
|
||||
*.mjs) node --check "$WORKSPACE/$f" || fail "syntax gate failed (node): $f" 1 ;;
|
||||
*.sh) bash -n "$WORKSPACE/$f" || fail "syntax gate failed (bash): $f" 1 ;;
|
||||
*.json) node -e 'JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"))' "$WORKSPACE/$f" || fail "syntax gate failed (json): $f" 1 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [ "$DRY_RUN" = "yes" ]; then
|
||||
echo "conductor-apply (dry-run): would apply $(echo "${#CHANGED[@]}") file(s) from $RUN_ID:"
|
||||
printf ' %s\n' "${CHANGED[@]}"
|
||||
echo "conductor-apply (dry-run): suites would run: $(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(", "))' "$POLICY_JSON")"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ---- apply ----
|
||||
[ -z "$(git status --porcelain)" ] || fail "target tree is not clean; refusing to mix states"
|
||||
git apply "$DIFF_FILE" || fail "git apply failed"
|
||||
|
||||
# ---- policy suites ----
|
||||
SUITES="$(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(" "))' "$POLICY_JSON")"
|
||||
SUITES_OK="yes"
|
||||
for s in $SUITES; do
|
||||
case "$s" in
|
||||
test-[a-z]*) : ;; # shape guard; existence checked next
|
||||
*) echo "conductor-apply: refusing suspicious suite name: $s" >&2; SUITES_OK="no"; break ;;
|
||||
esac
|
||||
[ -x "scripts/$s.sh" ] || { echo "conductor-apply: suite script missing: scripts/$s.sh" >&2; SUITES_OK="no"; break; }
|
||||
if ! bash "scripts/$s.sh" >/dev/null 2>&1; then
|
||||
echo "conductor-apply: suite failed: $s" >&2
|
||||
SUITES_OK="no"
|
||||
break
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$SUITES_OK" != "yes" ]; then
|
||||
git apply -R "$DIFF_FILE" && echo "conductor-apply: changes REVERTED (suites failed)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ---- commit with attribution ----
|
||||
git add -A
|
||||
git commit -q -m "feat(worker): auto-applied patch from run $RUN_ID
|
||||
|
||||
Authored-by: pi worker (run $RUN_ID, workspace $WORKSPACE_NAME)
|
||||
Applied-under: conductor-policy v1 (allowlist + syntax gates + suites)"
|
||||
echo "conductor-apply: applied and committed run $RUN_ID ($(echo "${#CHANGED[@]}") file(s)); suites: $SUITES"
|
||||
echo "conductor-apply: NOT pushed - push remains an explicit act."
|
||||
Executable
+141
@@ -0,0 +1,141 @@
|
||||
#!/usr/bin/env bash
|
||||
# Sandboxed selftests for the conductor auto-apply policy gate.
|
||||
#
|
||||
# Builds a throwaway target repo + worker workspace + fake run records, then
|
||||
# exercises every gate: policy validation, allowlist, syntax gates, suite
|
||||
# failure revert, disabled policy, missing/failed runs. No real model calls.
|
||||
set -uo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
|
||||
SANDBOX="$(mktemp -d)"
|
||||
trap 'rm -rf "$SANDBOX"' EXIT
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
|
||||
check_rc() { # name expectedRc command...
|
||||
local name="$1" expected="$2"
|
||||
shift 2
|
||||
local rc
|
||||
"$@" >/dev/null 2>&1
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
# ---- infrastructure: target repo + worker workspace + fake run ----
|
||||
git clone -q . "$SANDBOX/repo"
|
||||
# The clone carries committed state only - give the target its policy and
|
||||
# commit it so the tree starts clean (untracked policy would fail target_clean).
|
||||
mkdir -p "$SANDBOX/repo/roles"
|
||||
cp roles/conductor-policy.json "$SANDBOX/repo/roles/conductor-policy.json"
|
||||
git -C "$SANDBOX/repo" add roles/conductor-policy.json
|
||||
git -C "$SANDBOX/repo" -c user.name=suite -c user.email=suite@local commit -q -m policy
|
||||
mkdir -p "$SANDBOX/data/workspaces" "$SANDBOX/data/runs"
|
||||
git clone -q "$SANDBOX/repo" "$SANDBOX/data/workspaces/stack-repo"
|
||||
|
||||
TARGET="$SANDBOX/repo"
|
||||
WS="$SANDBOX/data/workspaces/stack-repo"
|
||||
cat > "$SANDBOX/config.json" <<EOF
|
||||
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"mock"}}
|
||||
EOF
|
||||
export MOSAIC_APPLY_TARGET="$TARGET"
|
||||
export MOSAIC_CONFIG="$SANDBOX/config.json"
|
||||
|
||||
RUN_OK="r-20260903T000000000Z-ok0000001"
|
||||
mkdir -p "$SANDBOX/data/runs/$RUN_OK"
|
||||
printf '{"runVersion":1,"runId":"%s","taskId":"t-fake","status":"succeeded","workspace":"stack-repo"}' "$RUN_OK" \
|
||||
> "$SANDBOX/data/runs/$RUN_OK/result.json"
|
||||
|
||||
ws_edit() { printf '\n%s\n' "$2" >> "$WS/$1"; }
|
||||
ws_reset() { git -C "$WS" checkout -q -- . 2>/dev/null; git -C "$WS" clean -qfd; }
|
||||
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
|
||||
|
||||
set_policy() { # enabled suites (commits: the target tree must stay clean)
|
||||
local suites="[\"$2\"]"
|
||||
printf '{"policyVersion":1,"autoApply":{"enabled":%s,"allowedPaths":["scripts/**","docs/**","README.md"],"suites":%s}}' "$1" "$suites" \
|
||||
> "$TARGET/roles/conductor-policy.json"
|
||||
git -C "$TARGET" add roles/conductor-policy.json
|
||||
git -C "$TARGET" -c user.name=suite -c user.email=suite@local commit -q -m "policy update"
|
||||
}
|
||||
set_policy true "test-config"
|
||||
|
||||
# T1: dry run - allowed change, nothing applied
|
||||
ws_edit "README.md" "worker dry-run line"
|
||||
check_rc "dry-run: allowed change, exit 0, nothing committed" 0 \
|
||||
scripts/conductor-apply.sh "$RUN_OK" --dry-run
|
||||
if git -C "$TARGET" log --format=%s | grep -q "auto-applied"; then
|
||||
check "dry-run committed nothing" 1
|
||||
else
|
||||
check "dry-run committed nothing" 0
|
||||
fi
|
||||
ws_reset
|
||||
|
||||
# T2: apply - allowed change, suites pass, commit created
|
||||
ws_edit "README.md" "worker applied line"
|
||||
check_rc "apply: allowed change exits 0" 0 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" log -1 --format=%s | grep -q "auto-applied patch from run $RUN_OK" \
|
||||
&& check "apply: attribution in commit subject" 0 || check "apply: attribution in commit subject" 1
|
||||
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
|
||||
target_clean && check "apply: target tree clean after commit" 0 || check "apply: target tree clean after commit" 1
|
||||
git -C "$TARGET" reset -q --hard HEAD~1
|
||||
|
||||
# T3: disallowed path refused
|
||||
ws_edit "Containerfile" "# worker touch"
|
||||
check_rc "disallowed path refused" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "disallowed path: target untouched" 0 || check "disallowed path: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T4: syntax gate - broken .mjs on an allowed path
|
||||
printf 'this is not (valid js\n' > "$WS/scripts/broken-worker.mjs"
|
||||
check_rc "syntax gate refused broken .mjs" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "syntax gate: target untouched" 0 || check "syntax gate: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T5: suite failure - allowed change breaks a policy suite -> auto-revert
|
||||
printf '\nexit 7\n' >> "$TARGET/scripts/test-config.sh"
|
||||
ws_edit "README.md" "worker change that will fail suites"
|
||||
check_rc "suite failure refused" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" checkout -q -- scripts/test-config.sh
|
||||
target_clean && check "suite failure: target reverted to clean" 0 || check "suite failure: target reverted to clean" 1
|
||||
|
||||
# T6: disabled policy
|
||||
set_policy false "test-config"
|
||||
ws_edit "README.md" "worker line while disabled"
|
||||
check_rc "disabled policy refused" 2 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "disabled policy: target untouched" 0 || check "disabled policy: target untouched" 1
|
||||
ws_reset
|
||||
set_policy true "test-config"
|
||||
|
||||
# T7: failed run refused
|
||||
RUN_FAIL="r-20260903T000000000Z-fail00001"
|
||||
mkdir -p "$SANDBOX/data/runs/$RUN_FAIL"
|
||||
printf '{"runVersion":1,"runId":"%s","taskId":"t","status":"failed","workspace":"stack-repo"}' "$RUN_FAIL" \
|
||||
> "$SANDBOX/data/runs/$RUN_FAIL/result.json"
|
||||
ws_edit "README.md" "worker line from failed run"
|
||||
check_rc "failed run refused" 1 scripts/conductor-apply.sh "$RUN_FAIL"
|
||||
target_clean && check "failed run: target untouched" 0 || check "failed run: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T8/T9: missing run + invalid policy
|
||||
check_rc "missing run exits 4" 4 scripts/conductor-apply.sh r-missing
|
||||
printf '{"policyVersion":9}' > "$TARGET/roles/conductor-policy.json"
|
||||
check_rc "invalid policy exits 2" 2 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" checkout -q -- roles/conductor-policy.json
|
||||
|
||||
echo
|
||||
echo "selftest: $PASS passed, $FAIL failed"
|
||||
[ "$FAIL" -eq 0 ]
|
||||
@@ -236,6 +236,17 @@ EOF
|
||||
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":["sudo"]}}' > "$SANDBOX/pol-m.json"
|
||||
expect_exit "invalid mission capabilities rejected" 2 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" $TASK validate "$SANDBOX/pol-t.json"
|
||||
|
||||
# live user context (M14): dispatched to every launch without rebuild
|
||||
mkdir -p "$SANDBOX/data/user"
|
||||
printf '\nUSER-CANON-MARKER\n' >> "$SANDBOX/data/user/USER.md"
|
||||
expect_exit "task run with user layer present" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOSAIC_HELLO_OK \
|
||||
scripts/run-task.sh run "$SANDBOX/ok.json"
|
||||
grep -q 'USER CONTEXT: USER.md' "$SANDBOX/data/system-prompt.md" \
|
||||
&& grep -q 'USER-CANON-MARKER' "$SANDBOX/data/system-prompt.md" \
|
||||
&& check "user context dispatched into generated prompt" 0 \
|
||||
|| check "user context dispatched into generated prompt" 1
|
||||
else
|
||||
echo "skip adapter seam cases (docker daemon unavailable)"
|
||||
fi
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
---
|
||||
name: ms-tools
|
||||
description: Contains a reference to all available tools for Mosaic Stack.
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
# ms-tools
|
||||
|
||||
You are a Mosaic fleet agent. A maintained toolkit lives at `<dir>`.
|
||||
Use it FIRST for the tasks below — improvising with raw CLIs causes the recurring failures this
|
||||
skill exists to prevent. This is the high-frequency fast path; the full reference is the
|
||||
`# Machine Tools` section already in your system prompt.
|
||||
|
||||
|
||||
@@ -0,0 +1,85 @@
|
||||
---
|
||||
name: ms-unslop
|
||||
description: Cut AI tells from any writing. Must always apply.
|
||||
disable-model-invocation: false
|
||||
---
|
||||
|
||||
# Unslop
|
||||
|
||||
Edit text to remove AI patterns and add human voice.
|
||||
|
||||
## Process
|
||||
|
||||
1. Scan for the patterns below.
|
||||
2. Rewrite. Preserve meaning, match intended tone.
|
||||
3. Add soul (see next section).
|
||||
4. Self-audit: "What makes this obviously AI generated?" Fix remaining tells.
|
||||
|
||||
## Adding soul
|
||||
|
||||
Removing patterns is half the job. Sterile, voiceless writing is just as obvious.
|
||||
|
||||
- **Have opinions.** React to facts instead of neutrally listing pros and cons.
|
||||
- **Vary rhythm.** Short sentences. Then longer ones that take their time. Mix it up.
|
||||
- **Acknowledge complexity.** "Impressive but also kind of unsettling" beats "impressive."
|
||||
- **Use "I" when it fits.** First person isn't unprofessional.
|
||||
- **Let some mess in.** Perfect structure looks machine-made.
|
||||
- **Be specific.** Not "this is concerning" but "there's something unsettling about agents churning away at 3am."
|
||||
|
||||
## Patterns to detect and fix
|
||||
|
||||
### Content
|
||||
|
||||
1. **Puffery.** `pivotal moment`, `testament to`, "evolving landscape", "setting the stage for", "indelible mark", "deeply rooted". Cut puffery, state what happened.
|
||||
2. **Name-dropping.** Listing media outlets without context. Pick one, say what was said.
|
||||
3. **Superficial -ing phrases.** "highlighting...", "ensuring...", "reflecting...", "showcasing...", "fostering...". Delete or expand with real sources.
|
||||
4. **Promotional language.** "nestled", `vibrant`, "breathtaking", "groundbreaking", "renowned", "stunning", "must-visit". Use neutral descriptions.
|
||||
5. **Vague attributions.** "Experts believe", "Industry reports suggest", "Some critics argue". Name the source or delete.
|
||||
6. **Formulaic challenges.** "Despite challenges... continues to thrive." Replace with specific facts.
|
||||
|
||||
### Language
|
||||
|
||||
7. **AI vocabulary.** `Additionally`, `crucial`, `delve`, `enduring`, `enhance`, `fostering`, `garner`, `interplay`, `intricate`, `landscape` (abstract), `pivotal`, `showcase`, `tapestry` (abstract), `testament`, `underscore`, `vibrant`. Replace with plain words.
|
||||
8. **Fancy ways to say "is".** "serves as", "stands as", "boasts", "features". Just say "is" or "has".
|
||||
9. **`Not just X, but Y`.** State the point directly instead.
|
||||
10. **Rule of three.** Forcing ideas into groups of three. Use the natural number.
|
||||
11. **Synonym cycling.** Protagonist, main character, central figure, hero all in one paragraph. Pick one, repeat it.
|
||||
12. **False ranges.** "from X to Y" where X and Y aren't on a meaningful scale. List topics directly.
|
||||
|
||||
### Style
|
||||
|
||||
13. **Em dash overuse.** Avoid em dashes entirely. Use periods or commas only (no parentheses, no en dashes, no hyphen-as-dash substitutes). Em dashes are an AI tell, and reaching for parentheses instead just trades one tell for another. If a thought needs separation, end the sentence or use a comma.
|
||||
14. **Colon overuse.** Colons are fine before a list or example. Not as mid-sentence connectors. "If you're coming from traditional automation: instead of registering event handlers, you describe conditions" adds nothing with the colon. Rewrite to let the point stand on its own without comparison framing. "Describing when the scheduler should fire works best as plain English." Same meaning, no crutch punctuation.
|
||||
15. **Boldface overuse.** Don't bold every proper noun or acronym.
|
||||
16. **Inline-header lists.** The tell is a bold label and colon that restates the line: "**Performance:** Performance improved...". Convert those to prose. A bold lead-in that ends in a period, names the item, and is followed by genuinely new detail ("**Schema in TypeScript.** Tables live in one file.") is fine, not a tell.
|
||||
17. **Title case headings.** Use sentence case.
|
||||
18. **Decorative emojis.** Remove from headings and bullets.
|
||||
19. **Curly quotes.** Replace with straight quotes.
|
||||
|
||||
### Communication artifacts
|
||||
|
||||
20. **Chatbot phrases.** `I hope this helps!`, `Let me know if...`, `Of course!`, `Certainly!`, `Found the smoking gun!` Remove.
|
||||
21. **Cutoff disclaimers.** "While specific details are limited..." Find sources or remove.
|
||||
22. **Sycophantic tone.** `Great question!` `You're absolutely right!` Respond directly.
|
||||
|
||||
### Filler
|
||||
|
||||
23. **Filler phrases.** `In order to` becomes "To". `Due to the fact that` becomes "Because". `It is important to note that` gets deleted.
|
||||
24. **Excessive hedging.** "could potentially possibly be argued that it might" becomes "may".
|
||||
25. **Generic conclusions.** "The future looks bright." State specific plans or facts.
|
||||
|
||||
### Jargon
|
||||
|
||||
26. **Abstract metaphor nouns.** Substrate, wedge, vector, locus, vantage, nexus, primitive (as noun), harness (as metaphor), surface (as in "API surface"), bedrock, scaffolding (as metaphor), modality, paradigm, gold-plating, ratchet (as metaphor), evacuate (for moving code), endgame, north star, flywheel. These read as technical but usually have a plainer concrete word. "Substrate" becomes "base". "Wedge in" becomes "add". "Vector" becomes "way" or "method". "Gold-plating" becomes "more than the job needs". "Ratchet" becomes the mechanism's real name or "a limit that only tightens". "Evacuate" becomes "move out". "Endgame" becomes "the last phase". Pick the concrete word.
|
||||
|
||||
### Plain speech
|
||||
|
||||
27. **Say what it does, not how it feels.** "the database stays close at hand", "SQL you can read", "types that follow your schema" name a feeling. The fix names the mechanism or a number: "`.toSQL()` returns the exact string sent to the database", "a column rename fails the build". Ask what the sentence tells the reader to do or know, then write that. If you can't restate it as a concrete instruction, fact, or number, cut it. One more check: if the sentence could appear unchanged in another project's docs, it says nothing about this one. Cut it.
|
||||
28. **Shorten or split dense sentences.** If the reader has to backtrack to parse a sentence, break it in two or drop clauses. One idea per sentence.
|
||||
29. **Active voice.** Prefer it. Catch "is/are/was/were + past participle" and name the actor: "queries are validated" becomes "the compiler validates queries", "the file is parsed by the loader" becomes "the loader parses the file". Passive is fine only when the actor is unknown or genuinely doesn't matter.
|
||||
30. **Cut adverbs, or use a stronger verb.** "runs quickly" becomes "is fast" or the number. "significantly improves" becomes the measured delta. An adverb propping up a weak verb means the verb is wrong.
|
||||
31. **Prefer the plain word.** `utilize` becomes "use", `leverage` becomes "use", `facilitate` becomes "help", "numerous" becomes "many", "in the event that" becomes "if". The fancier synonym is rarely clearer.
|
||||
|
||||
## Mention convention
|
||||
|
||||
A document that MENTIONS a banned word or phrase quotes it as inline code. The checker (`tools/unslop-hook/unslop-check.js`, machine source `tools/unslop-hook/lists.json`) strips code spans before matching, so a backticked mention is invisible to the gate while a bare one flags. This file follows that convention and doubles as a regression fixture: if `unslop-check.js` ever flags this file, either an edit broke the mention convention or code stripping regressed. Documents that deliberately CONTAIN slop to test detection (fixture files) are uses, not mentions; they are expected to flag.
|
||||
+47
-4
@@ -1,14 +1,33 @@
|
||||
#!/bin/sh
|
||||
# Load the four immutable contract files in fixed order and write the
|
||||
# generated system prompt to /var/lib/mosaic/system-prompt.md.
|
||||
# Load agent context and write the generated system prompt to
|
||||
# /var/lib/mosaic/system-prompt.md.
|
||||
#
|
||||
# Order is normative: CONSTITUTION.md, STANDARDS.md, SOUL.md, USER.md.
|
||||
# Layers, in normative order:
|
||||
# 1. Immutable contracts (image): CONSTITUTION, STANDARDS, SOUL
|
||||
# 2. Agent identity (when the launcher names the agent)
|
||||
# 3. Mission (when the task/launcher provides one)
|
||||
# 4. Live user context (M14): <dataRoot>/user/*.md - user-owned,
|
||||
# dispatched to every launch without rebuilds
|
||||
set -eu
|
||||
|
||||
CONTRACT_DIR="${1:-/opt/mosaic/contracts}"
|
||||
OUT="${2:-/var/lib/mosaic/system-prompt.md}"
|
||||
|
||||
FILES="CONSTITUTION.md STANDARDS.md SOUL.md USER.md"
|
||||
FILES="CONSTITUTION.md STANDARDS.md"
|
||||
|
||||
# SOUL slot (M15): the contract SOUL.md is the DEFAULT persona; a launched
|
||||
# agent seat overrides it with its own runtime SOUL (governance contracts
|
||||
# are never overridden).
|
||||
SOUL_SRC="$CONTRACT_DIR/SOUL.md"
|
||||
SOUL_HEADER="SOUL.md"
|
||||
if [ -n "${MOSAIC_AGENT_SOUL_FILE:-}" ]; then
|
||||
if [ ! -r "$MOSAIC_AGENT_SOUL_FILE" ]; then
|
||||
echo "load-contracts: agent SOUL not readable: $MOSAIC_AGENT_SOUL_FILE" >&2
|
||||
exit 1
|
||||
fi
|
||||
SOUL_SRC="$MOSAIC_AGENT_SOUL_FILE"
|
||||
SOUL_HEADER="SOUL.md (agent seat override)"
|
||||
fi
|
||||
|
||||
if [ ! -d "$CONTRACT_DIR" ]; then
|
||||
echo "load-contracts: contract directory not found: $CONTRACT_DIR" >&2
|
||||
@@ -33,6 +52,30 @@ for f in $FILES; do
|
||||
printf '\n' >> "$TEMP"
|
||||
done
|
||||
|
||||
printf '===== CONTRACT: %s =====\n' "$SOUL_HEADER" >> "$TEMP"
|
||||
cat "$SOUL_SRC" >> "$TEMP"
|
||||
printf '\n' >> "$TEMP"
|
||||
|
||||
# Agent identity (M13): when the launcher names the agent, the generated
|
||||
# prompt states it - SOUL.md provides the persona, this provides the name.
|
||||
if [ -n "${MOSAIC_AGENT_NAME:-}" ]; then
|
||||
printf '===== AGENT IDENTITY =====\n' >> "$TEMP"
|
||||
printf 'agent name: %s\n' "$MOSAIC_AGENT_NAME" >> "$TEMP"
|
||||
[ -n "${MOSAIC_AGENT_ROLE:-}" ] && printf 'agent role: %s\n' "$MOSAIC_AGENT_ROLE" >> "$TEMP"
|
||||
printf '\n' >> "$TEMP"
|
||||
fi
|
||||
|
||||
# Live user context (M14): every *.md in /var/lib/mosaic/user (sorted) is
|
||||
# appended - the user owns this layer and edits it without rebuilds.
|
||||
USER_DIR="/var/lib/mosaic/user"
|
||||
if [ -d "$USER_DIR" ]; then
|
||||
for f in $(ls "$USER_DIR"/*.md 2>/dev/null | sort); do
|
||||
printf '===== USER CONTEXT: %s =====\n' "$(basename "$f")" >> "$TEMP"
|
||||
cat "$f" >> "$TEMP"
|
||||
printf '\n' >> "$TEMP"
|
||||
done
|
||||
fi
|
||||
|
||||
# Sanctioned mission injection point (M4): when the task runner provides a
|
||||
# mission snapshot, its objective and directives are appended AFTER the
|
||||
# immutable contracts. Runtime data; never part of the contract fixtures.
|
||||
|
||||
+15
-7
@@ -1,13 +1,22 @@
|
||||
#!/bin/sh
|
||||
# One-shot agent dispatcher inside the container.
|
||||
# Agent dispatcher inside the container.
|
||||
#
|
||||
# 1. Loads the contract-generated system prompt (contracts + optional
|
||||
# mission section from MOSAIC_MISSION_FILE).
|
||||
# 2. Dispatches to /opt/mosaic/adapters/<MOSAIC_ADAPTER>/adapter.sh per
|
||||
# the contract in /opt/mosaic/adapters/README.md.
|
||||
# Headless (default): loads the contract-generated system prompt, then
|
||||
# dispatches one request to /opt/mosaic/adapters/<MOSAIC_ADAPTER>/adapter.sh
|
||||
# (contract: /opt/mosaic/adapters/README.md).
|
||||
#
|
||||
# Interactive (MOSAIC_INTERACTIVE=1, from scripts/agent.sh): same prompt,
|
||||
# but the adapter opens the full pi TUI with no initial prompt - the human
|
||||
# drives from there.
|
||||
set -eu
|
||||
|
||||
REQUEST="${*:-Return your startup marker and nothing else.}"
|
||||
REQUEST=""
|
||||
if [ "${MOSAIC_INTERACTIVE:-}" != "1" ]; then
|
||||
# Headless: args are the request; default is the startup verification
|
||||
# request used by hello/verify.
|
||||
REQUEST="${*:-Return your startup marker and nothing else.}"
|
||||
export MOSAIC_REQUEST="$REQUEST"
|
||||
fi
|
||||
|
||||
ADAPTER="${MOSAIC_ADAPTER:-pi}"
|
||||
case "$ADAPTER" in
|
||||
@@ -28,6 +37,5 @@ fi
|
||||
/opt/mosaic/src/load-contracts.sh /opt/mosaic/contracts /var/lib/mosaic/system-prompt.md
|
||||
|
||||
export MOSAIC_SYSTEM_PROMPT_FILE="/var/lib/mosaic/system-prompt.md"
|
||||
export MOSAIC_REQUEST="$REQUEST"
|
||||
|
||||
exec "$ADAPTER_SCRIPT"
|
||||
|
||||
@@ -0,0 +1,83 @@
|
||||
# unslop-hook
|
||||
|
||||
Mechanical AI-tell enforcement for pi seats. Anti-drift gate for the writing
|
||||
standard in SYSTEM.md / ms-unslop: prose distribution alone decays over long
|
||||
sessions; this check cannot forget.
|
||||
|
||||
- `lists.json`: committed machine source for every list the checker enforces:
|
||||
words, phrases, punct rules, regex patterns, density thresholds. Each entry
|
||||
carries provenance (`ms-unslop:<pattern id>` or `system-md`), the mention
|
||||
convention, and the documented divergence of the density gate from
|
||||
SYSTEM.md's outright em-dash ban. Edit lists here, not in code.
|
||||
- `unslop-check.js`: dependency-free checker (node CLI + module) driven by
|
||||
lists.json. Loads and schema-validates the lists on first use and hard-fails
|
||||
closed: empty, unparseable, or invalid lists throw. Detects banned vocabulary,
|
||||
chatbot/sycophancy phrases, filler phrases, em/en dashes, curly quotes,
|
||||
`not just X but Y`. Strips fenced and inline code first, so quoted code is
|
||||
never flagged. Exit 0 clean, 1 violations, 2 gate broken (lists unreadable,
|
||||
never a clean verdict).
|
||||
- `extension.ts`: pi extension. `message_end` checks finalized assistant text
|
||||
and notifies the operator (TUI/RPC). `before_agent_start` reads the most
|
||||
recent assistant reply from the session file and, if it carries tells,
|
||||
injects a correction notice the model sees on its next turn. `/unslop`
|
||||
reports session stats. Violation state lives in the session file, so the
|
||||
injection path survives restart, resume, fork, and reload (an in-memory
|
||||
pending flag was measured dead across print-mode turns, 2026-08-19). A
|
||||
broken lists.json fails closed: checks stop, `broken_lists` /
|
||||
`skipped_broken` events log the reason, operator notified once, seat keeps
|
||||
running.
|
||||
- `test-unslop-check.js`: unit tests with red and green controls.
|
||||
|
||||
## Use
|
||||
|
||||
```bash
|
||||
node test-unslop-check.js # suite
|
||||
node unslop-check.js <file> # CLI check
|
||||
UNSLOP_LISTS=<path> node unslop-check.js <file> # alt lists location
|
||||
pi -e ~/.mosaic/tools/unslop-hook/extension.ts # ad-hoc load
|
||||
# deploy: copy dir to ~/.pi/agent/extensions/unslop-hook/ or seat .pi
|
||||
# equivalent, or list it in settings.json "extensions"
|
||||
```
|
||||
|
||||
Env: `MOSAIC_UNSLOP_HOOK=0` disables. `MOSAIC_UNSLOP_LOG=<path>` appends JSONL
|
||||
events (loaded / flagged / notice_injected / checked / broken_lists /
|
||||
skipped_broken) for headless evidence. `UNSLOP_LISTS=<path>` overrides the
|
||||
lists.json location for both CLI and extension.
|
||||
|
||||
## Verified here (2026-08-19)
|
||||
|
||||
- Unit suite 24/24 (8 behavioral, 11 loader/CLI, 5 review follow-up), red and
|
||||
green controls both exercised, including exit-2 on broken lists and on an
|
||||
unreadable input file (S3).
|
||||
- CLI: slop file exit 1, clean file exit 0, broken lists exit 2 with the fault
|
||||
named on stderr.
|
||||
- Extension, healthy path (print mode, zai/glm-5.3:low): startup probe loads
|
||||
lists.json, reply checked clean.
|
||||
- Extension, broken-lists path (print mode): `broken_lists` at startup,
|
||||
`skipped_broken` per turn, seat survives, reply still delivered.
|
||||
- Earlier live evidence (pre-C1, inline lists): forced-slop turn flagged;
|
||||
fresh-process follow-up injected the notice and the reply came back clean;
|
||||
full TUI trial (notify line, injection, /unslop stats) on session vision-unslop.
|
||||
- Log evidence in session scratchpad.
|
||||
|
||||
## Limits
|
||||
|
||||
- `/unslop` command not tested headless (print mode has no command surface);
|
||||
it is a thin stats wrapper.
|
||||
- En dash flag fires on typographic ranges too (2–3); acceptable for fleet
|
||||
prose, revisit if it noisifies technical writing.
|
||||
- Notice injection is a nudger, not a blocker. Output already streamed to the
|
||||
user stays as-is; correction lands on the next turn.
|
||||
- A broken lists.json latches for the session: repairing the file mid-session
|
||||
does not revive checks until the seat restarts. Acceptable for an advisory
|
||||
gate (review S1).
|
||||
- The fail-closed operator notification requires a UI. Print-mode sessions
|
||||
log `skipped_broken` but notify nobody (review S2).
|
||||
- The word/phrase lists are the mechanical subset of ms-unslop only, keyed to
|
||||
pattern ids in lists.json. Style judgments (voice, rhythm, structure) stay in
|
||||
the skill, not the gate.
|
||||
|
||||
## Promotion path
|
||||
|
||||
Stack issue (A4): checker shared as the single source for a matching Claude
|
||||
Code Stop-hook script; lists versioned beside SYSTEM.md contract text.
|
||||
@@ -0,0 +1,163 @@
|
||||
// unslop-hook — pi extension wrapper around unslop-check.js.
|
||||
// Detects mechanical AI tells in finalized assistant messages and injects a
|
||||
// correction notice the model sees on its next turn. Anti-drift enforcement for
|
||||
// SYSTEM.md / ms-unslop; prose distribution alone decays, this cannot forget.
|
||||
//
|
||||
// Deploy: copy dir to ~/.pi/agent/extensions/unslop-hook/ (or seat .pi equivalent),
|
||||
// or add this file's dir to settings.json "extensions".
|
||||
// Test: pi -e <abs path>/extension.ts
|
||||
// Off: MOSAIC_UNSLOP_HOOK=0
|
||||
// Log: MOSAIC_UNSLOP_LOG=/path/to/log.jsonl (JSONL events; headless evidence)
|
||||
// Broken: lists.json missing/empty/invalid → the checker throws; checks are
|
||||
// skipped, logged as skipped_broken, and the operator is notified once.
|
||||
// Never silently pass while the lists cannot load (fail closed).
|
||||
//
|
||||
// Design note: violation state lives in the SESSION FILE, not memory. At
|
||||
// before_agent_start we read the most recent assistant text message from
|
||||
// ctx.sessionManager and check it there. That survives process restarts, resume,
|
||||
// fork, and reload — an in-memory pending flag measured dead on 2026-08-19 when
|
||||
// a print-mode second turn never injected.
|
||||
import { appendFileSync } from "node:fs";
|
||||
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
||||
import { checkText } from "./unslop-check.js";
|
||||
|
||||
interface Finding {
|
||||
rule: string;
|
||||
detail: string;
|
||||
count: number;
|
||||
}
|
||||
|
||||
interface MessageEntry {
|
||||
type: "message";
|
||||
id: string;
|
||||
message: { role?: string; content?: unknown };
|
||||
}
|
||||
|
||||
function assistantText(entry: unknown): string | null {
|
||||
const e = entry as Partial<MessageEntry>;
|
||||
if (e?.type !== "message") return null;
|
||||
const msg = e.message;
|
||||
if (msg?.role !== "assistant" || !Array.isArray(msg.content)) return null;
|
||||
const text = msg.content
|
||||
.filter((b): b is { type: "text"; text: string } =>
|
||||
typeof b === "object" && b !== null && (b as { type?: string }).type === "text")
|
||||
.map((b) => b.text ?? "")
|
||||
.join("\n");
|
||||
return text.trim() ? text : null; // tool-call-only assistant messages return null
|
||||
}
|
||||
|
||||
export default function (pi: ExtensionAPI) {
|
||||
if (process.env.MOSAIC_UNSLOP_HOOK === "0") return;
|
||||
|
||||
const LOG = process.env.MOSAIC_UNSLOP_LOG;
|
||||
const log = (ev: Record<string, unknown>) => {
|
||||
if (LOG) appendFileSync(LOG, JSON.stringify({ ts: Date.now(), ...ev }) + "\n");
|
||||
};
|
||||
|
||||
// Entry ids we have already injected a notice for. In-memory only: after a
|
||||
// restart the same entry may inject once more, which re-anchors the style
|
||||
// after a context loss. That is wanted, not a bug.
|
||||
const injectedFor = new Set<string>();
|
||||
let turnsChecked = 0;
|
||||
let turnsFlagged = 0;
|
||||
const histogram = new Map<string, number>();
|
||||
|
||||
// Fail-closed path for a broken lists.json. A checker that cannot load its
|
||||
// lists must never be read as "everything passed": checks stop, the skip is
|
||||
// logged each turn, and the operator is notified once.
|
||||
let broken: string | null = null;
|
||||
let brokenNotified = false;
|
||||
const reportBroken = (ctx: { hasUI?: boolean } | undefined, where: string) => {
|
||||
log({ ev: "skipped_broken", where, reason: broken });
|
||||
if (!brokenNotified && ctx?.hasUI) {
|
||||
ctx.ui.notify(`unslop gate BROKEN: ${broken}. Fix tools/unslop-hook/lists.json; no clean verdicts until then.`, "error");
|
||||
brokenNotified = true;
|
||||
}
|
||||
};
|
||||
const safeCheck = (text: string): ReturnType<typeof checkText> | null => {
|
||||
if (broken) return null;
|
||||
try {
|
||||
return checkText(text);
|
||||
} catch (e) {
|
||||
broken = String((e as Error).message);
|
||||
log({ ev: "broken_lists", reason: broken });
|
||||
return null;
|
||||
}
|
||||
};
|
||||
|
||||
pi.on("session_start", async (event, _ctx) => {
|
||||
log({ ev: "loaded", reason: event.reason });
|
||||
try {
|
||||
checkText(""); // probe: load+validate lists at startup, not mid-conversation
|
||||
} catch (e) {
|
||||
broken = String((e as Error).message);
|
||||
log({ ev: "broken_lists", reason: broken, at: "startup" });
|
||||
}
|
||||
});
|
||||
|
||||
pi.on("message_end", async (event, ctx) => {
|
||||
if ((event.message as { role?: string }).role !== "assistant") return;
|
||||
const text = assistantText({ type: "message", id: "", message: event.message });
|
||||
if (text === null) return;
|
||||
|
||||
const result = safeCheck(text);
|
||||
if (result === null) {
|
||||
reportBroken(ctx, "message_end");
|
||||
return;
|
||||
}
|
||||
turnsChecked++;
|
||||
if (result.clean) {
|
||||
log({ ev: "checked", clean: true, turn: turnsChecked, charsChecked: result.charsChecked });
|
||||
return;
|
||||
}
|
||||
turnsFlagged++;
|
||||
for (const f of result.findings) histogram.set(f.rule, (histogram.get(f.rule) ?? 0) + 1);
|
||||
const summary = result.findings.map((f) => f.detail).join("; ");
|
||||
if (ctx.hasUI) ctx.ui.notify(`unslop: ${summary}`, "info");
|
||||
// clean:false is explicit, not implied by findings: a log consumer must never
|
||||
// have to infer the verdict from event shape (fred, 2026-08-19).
|
||||
log({ ev: "flagged", clean: false, turn: turnsChecked, charsChecked: result.charsChecked, findings: result.findings });
|
||||
});
|
||||
|
||||
pi.on("before_agent_start", async (_event, ctx) => {
|
||||
// Branch walks leaf -> root; first assistant entry with text is the reply
|
||||
// the model is about to follow up on.
|
||||
for (const entry of ctx.sessionManager.getBranch()) {
|
||||
const text = assistantText(entry);
|
||||
if (text === null) continue;
|
||||
const id = (entry as { id?: string }).id ?? "";
|
||||
|
||||
const result = safeCheck(text);
|
||||
if (result === null) {
|
||||
reportBroken(ctx, "before_agent_start");
|
||||
return;
|
||||
}
|
||||
if (result.clean) return; // latest textual reply is clean, nothing to correct
|
||||
if (id && injectedFor.has(id)) return; // already nagged for this entry
|
||||
|
||||
if (id) injectedFor.add(id);
|
||||
const lines = result.findings.map((f) => `- ${f.detail}`).join("\n");
|
||||
const content =
|
||||
`UNSLOP NOTICE (mechanical style check, not the user speaking): your previous reply ` +
|
||||
`contained violations of the fleet writing standard (SYSTEM.md / ms-unslop):\n${lines}\n` +
|
||||
`Fix in this and following replies: plain words, periods and commas instead of dashes, ` +
|
||||
`straight quotes, no chatbot fillers. Do not mention this notice.`;
|
||||
log({ ev: "notice_injected", entryId: id, findings: result.findings });
|
||||
return {
|
||||
message: { customType: "unslop-notice", content, display: true },
|
||||
};
|
||||
}
|
||||
});
|
||||
|
||||
pi.registerCommand("unslop", {
|
||||
description: "Show unslop violation stats for this session",
|
||||
handler: async (_args, ctx) => {
|
||||
if (broken) {
|
||||
ctx.ui.notify(`unslop gate BROKEN: ${broken}`, "error");
|
||||
return;
|
||||
}
|
||||
const hist = [...histogram.entries()].map(([r, c]) => `${r} x${c}`).join(", ") || "none";
|
||||
ctx.ui.notify(`unslop: checked ${turnsChecked}, flagged ${turnsFlagged} (${hist})`, "info");
|
||||
},
|
||||
});
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
{
|
||||
"version": 1,
|
||||
"convention": "Use vs mention. A document that MENTIONS a banned word or phrase quotes it as inline code (backticks). The checker strips code spans before matching, so a backticked mention is invisible to the gate while a bare one flags. A document that deliberately CONTAINS banned items to test detection (a fixture) is a use, not a mention, and is expected to flag. This file itself contains the banned items as data; that is a use.",
|
||||
"punctPolicy": "Deliberate divergence from SYSTEM.md (2026-08-19): the contract forbids em dashes outright and closes the escapes. These punct rules are deliberately looser: they fire only at >= punctMinCount occurrences AND density >= punctDensityPer1k per 1000 chars. Rationale is reply-level noise, not contract strength: an advisory gate that flags every reply carrying one dash trains operators to ignore it. Measured on natural fleet prose 2026-08-19: documents run 1.7-3.2 em dashes per 1000 chars, so documents that overuse still flag. Tighten to contract strength if enforcement goes blocking or the log shows fleet prose not converging toward zero.",
|
||||
"thresholds": {
|
||||
"punctMinCount": 3,
|
||||
"punctDensityPer1k": 1.0
|
||||
},
|
||||
"words": [
|
||||
{ "value": "additionally", "source": "ms-unslop:7" },
|
||||
{ "value": "crucial", "source": "ms-unslop:7" },
|
||||
{ "value": "delve", "source": "ms-unslop:7" },
|
||||
{ "value": "garner", "source": "ms-unslop:7" },
|
||||
{ "value": "interplay", "source": "ms-unslop:7" },
|
||||
{ "value": "intricate", "source": "ms-unslop:7" },
|
||||
{ "value": "pivotal", "source": "ms-unslop:7" },
|
||||
{ "value": "showcase", "source": "ms-unslop:7" },
|
||||
{ "value": "tapestry", "source": "ms-unslop:7" },
|
||||
{ "value": "testament", "source": "ms-unslop:7" },
|
||||
{ "value": "underscore", "source": "ms-unslop:7" },
|
||||
{ "value": "vibrant", "source": "ms-unslop:7" },
|
||||
{ "value": "utilize", "source": "ms-unslop:31" },
|
||||
{ "value": "leverage", "source": "ms-unslop:31" },
|
||||
{ "value": "facilitate", "source": "ms-unslop:31" },
|
||||
{ "value": "load-bearing", "source": "system-md" }
|
||||
],
|
||||
"phrases": [
|
||||
{ "value": "worth stating plainly", "source": "system-md" },
|
||||
{ "value": "here's the honest truth", "source": "system-md" },
|
||||
{ "value": "heres the honest truth", "source": "system-md", "note": "apostrophe-OMITTED renderings only; ASCII and curly-apostrophe forms match the main entry because the checker normalizes U+2019/U+2018 to ASCII before phrase matching" },
|
||||
{ "value": "the real tension", "source": "system-md" },
|
||||
{ "value": "carry the argument", "source": "system-md" },
|
||||
{ "value": "in order to", "source": "ms-unslop:23" },
|
||||
{ "value": "due to the fact that", "source": "ms-unslop:23" },
|
||||
{ "value": "it is important to note", "source": "ms-unslop:23" },
|
||||
{ "value": "i hope this helps", "source": "ms-unslop:20" },
|
||||
{ "value": "let me know if", "source": "ms-unslop:20" },
|
||||
{ "value": "of course!", "source": "ms-unslop:20" },
|
||||
{ "value": "certainly!", "source": "ms-unslop:20" },
|
||||
{ "value": "found the smoking gun", "source": "ms-unslop:20" },
|
||||
{ "value": "happy to help", "source": "ms-unslop:20", "note": "extension of the named pattern set" },
|
||||
{ "value": "great question", "source": "ms-unslop:22" },
|
||||
{ "value": "absolutely right", "source": "ms-unslop:22" },
|
||||
{ "value": "excellent question", "source": "ms-unslop:22", "note": "extension of the named pattern set" }
|
||||
],
|
||||
"punct": [
|
||||
{ "value": "em", "label": "em dash", "chars": ["\u2014"], "source": "ms-unslop:13+system-md" },
|
||||
{ "value": "en", "label": "en dash", "chars": ["\u2013"], "source": "ms-unslop:13" },
|
||||
{ "value": "curly", "label": "curly quote/apostrophe", "chars": ["\u201c", "\u201d", "\u2018", "\u2019"], "source": "ms-unslop:19" }
|
||||
],
|
||||
"patterns": [
|
||||
{ "value": "not-just-but", "regex": "not just\\s+[^.!?]{0,80}?\\s+but", "flags": "gi", "detail": "not just X but Y", "source": "ms-unslop:9" }
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,228 @@
|
||||
"use strict";
|
||||
// Tests for unslop-check.js. Run: node test-unslop-check.js
|
||||
// Exit 0 = all pass. Cases include a red control (slop must fail) and a green
|
||||
// control (clean prose must pass) per evidence discipline.
|
||||
|
||||
const assert = require("node:assert");
|
||||
const fs = require("node:fs");
|
||||
const os = require("node:os");
|
||||
const path = require("node:path");
|
||||
const { spawnSync } = require("node:child_process");
|
||||
const { checkText, stripCode, loadLists } = require("./unslop-check.js");
|
||||
|
||||
const SLOP = `Certainly! Let me delve into the evolving tapestry of database technology — it’s truly “pivotal” — and intricate.
|
||||
In order to understand it — we should leverage this interplay of systems — deeply. I hope this helps!`;
|
||||
|
||||
const CLEAN = `The loader parses the file and validates each row. Rows that fail are logged
|
||||
and skipped. We measured a range from 1 to 10 seconds. Use "straight quotes" and
|
||||
commas, not dashes. That is the whole finding.`;
|
||||
|
||||
// Code-stripping control: banned words inside code must not count.
|
||||
const WITH_CODE = [
|
||||
"The config uses `utilize=false` internally.",
|
||||
"```",
|
||||
"delve tapestry — pivotal",
|
||||
"```",
|
||||
"The config file sets one flag. It is parsed at startup.",
|
||||
].join("\n");
|
||||
|
||||
const results = [];
|
||||
function t(name, fn) {
|
||||
try { fn(); results.push([name, true]); } catch (e) { results.push([name, false]); console.error(`FAIL ${name}: ${e.message}`); }
|
||||
}
|
||||
|
||||
t("slop fixture is flagged (red control)", () => {
|
||||
const r = checkText(SLOP);
|
||||
assert.ok(!r.clean, "slop must not be clean");
|
||||
const details = r.findings.map((f) => f.detail).join("; ");
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("delve")), `delve missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("tapestry")), `tapestry missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("pivotal")), `pivotal missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("em dash")), `em dash missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("curly")), `curly missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("in order to")), `in order to missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("i hope this helps")), `chatbot phrase missing: ${details}`);
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("certainly")), `certainly missing: ${details}`);
|
||||
});
|
||||
|
||||
t("clean fixture passes (green control)", () => {
|
||||
const r = checkText(CLEAN);
|
||||
assert.deepStrictEqual(r.findings, [], `unexpected findings: ${JSON.stringify(r.findings)}`);
|
||||
});
|
||||
|
||||
t("numeric range is not a false range flag", () => {
|
||||
const r = checkText(CLEAN);
|
||||
assert.ok(!r.findings.some((f) => f.rule === "pattern"), "must not flag numeric ranges");
|
||||
});
|
||||
|
||||
t("code blocks and inline code are stripped", () => {
|
||||
const r = checkText(WITH_CODE);
|
||||
assert.deepStrictEqual(r.findings, [], `code leaked into check: ${JSON.stringify(r.findings)}`);
|
||||
});
|
||||
|
||||
t("stripCode removes fenced and inline code", () => {
|
||||
const s = stripCode("a `x — y` b\n```\ndelve\n```\nc");
|
||||
assert.ok(!s.includes("delve"), "fenced code not stripped");
|
||||
assert.ok(!s.includes("—"), "inline code not stripped");
|
||||
assert.ok(s.includes("a") && s.includes("b") && s.includes("c"), "prose lost");
|
||||
});
|
||||
|
||||
t("not-just-but pattern is detected", () => {
|
||||
const r = checkText("This is not just a cache but a coordination layer.");
|
||||
assert.ok(r.findings.some((f) => f.rule === "pattern"), "pattern missed");
|
||||
});
|
||||
|
||||
t("light dash use is not flagged (below threshold)", () => {
|
||||
const prose =
|
||||
"The loader parses each row and validates it against the schema. Rows that fail " +
|
||||
"are logged — with their line numbers — and skipped. The operator reviews the log " +
|
||||
"daily and reconciles the rejects against the source system by hand, which takes " +
|
||||
"a few minutes and has never once produced a discrepancy worth acting on.";
|
||||
const r = checkText(prose);
|
||||
assert.ok(!r.findings.some((f) => f.detail.includes("dash")), "2 dashes in ~330 chars must not flag");
|
||||
});
|
||||
|
||||
t("dash overuse is flagged (above threshold)", () => {
|
||||
const r = checkText("One — two — three — four. That is the whole sentence.");
|
||||
assert.ok(r.findings.some((f) => f.detail.includes("em dash")), "4 dashes in 50 chars must flag");
|
||||
});
|
||||
|
||||
// ── C1: lists.json machine source ──────────────────────────────────────
|
||||
// The committed lists are the single source of truth; these tests pin the
|
||||
// file's validity, its shape, and the loader's fail-closed behavior.
|
||||
|
||||
const tmpdir = fs.mkdtempSync(path.join(os.tmpdir(), "unslop-c1-"));
|
||||
const tmp = (n) => path.join(tmpdir, n);
|
||||
|
||||
function brokenVariant(mutate) {
|
||||
const l = JSON.parse(JSON.stringify(loadLists()));
|
||||
mutate(l);
|
||||
return l;
|
||||
}
|
||||
|
||||
function writeTmp(name, data) {
|
||||
const f = tmp(name);
|
||||
fs.writeFileSync(f, typeof data === "string" ? data : JSON.stringify(data));
|
||||
return f;
|
||||
}
|
||||
|
||||
t("lists.json (the real file) validates and is pinned in size", () => {
|
||||
const l = loadLists();
|
||||
assert.strictEqual(l.version, 1);
|
||||
// Counts pin the migration: 16 words, 17 phrases, 3 punct, 1 pattern moved
|
||||
// from the old inline constants. Changing a count means changing this test
|
||||
// too, consciously.
|
||||
assert.strictEqual(l.words.length, 16, "word count drifted");
|
||||
assert.strictEqual(l.phrases.length, 17, "phrase count drifted");
|
||||
assert.strictEqual(l.punct.length, 3, "punct count drifted");
|
||||
assert.strictEqual(l.patterns.length, 1, "pattern count drifted");
|
||||
assert.ok(l.convention.length > 50, "mention convention must be present");
|
||||
assert.ok(l.punctPolicy.length > 50, "punct divergence policy must be present");
|
||||
for (const e of [...l.words, ...l.phrases, ...l.punct, ...l.patterns]) {
|
||||
assert.ok(e.source && e.source.trim(), `entry missing source: ${JSON.stringify(e)}`);
|
||||
}
|
||||
});
|
||||
|
||||
t("loader rejects an empty file", () => {
|
||||
const f = writeTmp("empty.json", "");
|
||||
assert.throws(() => loadLists(f), /empty file/);
|
||||
});
|
||||
|
||||
t("loader rejects unparseable JSON", () => {
|
||||
const f = writeTmp("bad.json", "{nope");
|
||||
assert.throws(() => loadLists(f), /unparseable/);
|
||||
});
|
||||
|
||||
t("loader rejects a missing file", () => {
|
||||
assert.throws(() => loadLists(tmp("does-not-exist.json")), /cannot read/);
|
||||
});
|
||||
|
||||
t("loader rejects missing keys", () => {
|
||||
const f = writeTmp("nokeys.json", { version: 1 });
|
||||
assert.throws(() => loadLists(f), /missing key/);
|
||||
});
|
||||
|
||||
t("loader rejects an emptied word list", () => {
|
||||
const f = writeTmp("emptywords.json", brokenVariant((l) => { l.words = []; }));
|
||||
assert.throws(() => loadLists(f), /words must be a non-empty array/);
|
||||
});
|
||||
|
||||
t("loader rejects entries without provenance", () => {
|
||||
const f = writeTmp("nosource.json", brokenVariant((l) => { delete l.phrases[0].source; }));
|
||||
assert.throws(() => loadLists(f), /source/);
|
||||
});
|
||||
|
||||
t("loader rejects duplicate values", () => {
|
||||
const f = writeTmp("dup.json", brokenVariant((l) => { l.words.push({ ...l.words[0] }); }));
|
||||
assert.throws(() => loadLists(f), /duplicate/);
|
||||
});
|
||||
|
||||
t("loader rejects a non-compiling pattern regex", () => {
|
||||
const f = writeTmp("badregex.json", brokenVariant((l) => { l.patterns[0].regex = "("; }));
|
||||
assert.throws(() => loadLists(f), /does not compile/);
|
||||
});
|
||||
|
||||
t("CLI exits 2 on broken lists (red control)", () => {
|
||||
const f = writeTmp("cli-broken.json", "");
|
||||
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js")], {
|
||||
input: "some prose",
|
||||
encoding: "utf8",
|
||||
env: { ...process.env, UNSLOP_LISTS: f },
|
||||
});
|
||||
assert.strictEqual(r.status, 2, `expected exit 2, got ${r.status} (stderr: ${r.stderr})`);
|
||||
assert.ok(r.stderr.includes("lists.json invalid"), `stderr must name the fault: ${r.stderr}`);
|
||||
});
|
||||
|
||||
t("CLI honors UNSLOP_LISTS for a valid file (green control)", () => {
|
||||
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js")], {
|
||||
input: "plain prose with no tells at all",
|
||||
encoding: "utf8",
|
||||
env: { ...process.env, UNSLOP_LISTS: path.join(__dirname, "lists.json") },
|
||||
});
|
||||
assert.strictEqual(r.status, 0, `expected exit 0, got ${r.status} (stderr: ${r.stderr})`);
|
||||
});
|
||||
|
||||
// ── Review follow-up (rev-code-01, 2026-08-19): F1, F2, F3, S3 ────────
|
||||
|
||||
t("loader rejects non-finite thresholds (F1)", () => {
|
||||
// Infinity cannot round-trip JSON.stringify, so the fixture is a raw string
|
||||
// edit of the real file — exactly the hand-edit that produced the finding.
|
||||
const real = fs.readFileSync(path.join(__dirname, "lists.json"), "utf8");
|
||||
const f = writeTmp("inf-threshold.json", real.replace('"punctMinCount": 3', '"punctMinCount": 1e999'));
|
||||
assert.ok(real !== fs.readFileSync(f, "utf8") || !real.includes('"punctMinCount": 3'), "fixture mutation did not apply; test is vacuous");
|
||||
assert.throws(() => loadLists(f), /finite/);
|
||||
const f2 = writeTmp("inf-density.json", real.replace('"punctDensityPer1k": 1.0', '"punctDensityPer1k": 1e999'));
|
||||
assert.throws(() => loadLists(f2), /finite/);
|
||||
});
|
||||
|
||||
t("curly-apostrophe phrase rendering is flagged (F2 red control)", () => {
|
||||
const r = checkText("Here\u2019s the honest truth about the deploy.");
|
||||
assert.ok(!r.clean, "curly apostrophe must not defeat phrase matching");
|
||||
assert.ok(r.findings.some((x) => x.detail.includes("here's the honest truth")), `main entry must match, got: ${JSON.stringify(r.findings)}`);
|
||||
});
|
||||
|
||||
t("curly apostrophes still fire the punct rule alongside phrases (F2 ordering)", () => {
|
||||
// Normalization for phrases must not eat the punct signal: four curly
|
||||
// quotes in short text must flag punct, not only the phrase.
|
||||
const r = checkText("It\u2019s \u2019one\u2019 \u2019two\u2019 \u2019three\u2019 \u2019four\u2019 done.");
|
||||
assert.ok(r.findings.some((f) => f.rule === "punct"), `punct must fire on original text: ${JSON.stringify(r.findings)}`);
|
||||
});
|
||||
|
||||
t("loader rejects non-lowercase phrase values (F3)", () => {
|
||||
const f = writeTmp("cap-phrase.json", brokenVariant((l) => { l.phrases[0].value = "Worth Stating Plainly"; }));
|
||||
assert.throws(() => loadLists(f), /lowercase/);
|
||||
});
|
||||
|
||||
t("CLI exits 2 on unreadable input file (S3)", () => {
|
||||
const r = spawnSync(process.execPath, [path.join(__dirname, "unslop-check.js"), tmp("definitely-absent.txt")], {
|
||||
encoding: "utf8",
|
||||
});
|
||||
assert.strictEqual(r.status, 2, `expected exit 2, got ${r.status} (stderr: ${r.stderr})`);
|
||||
assert.ok(r.stderr.includes("cannot read input"), `stderr must name the fault: ${r.stderr}`);
|
||||
});
|
||||
|
||||
let failed = 0;
|
||||
for (const [name, ok] of results) { console.log(`${ok ? "PASS" : "FAIL"} ${name}`); if (!ok) failed++; }
|
||||
console.log(`${results.length - failed}/${results.length} passed`);
|
||||
try { fs.rmSync(tmpdir, { recursive: true, force: true }); } catch {}
|
||||
process.exit(failed ? 1 : 0);
|
||||
@@ -0,0 +1,181 @@
|
||||
#!/usr/bin/env node
|
||||
"use strict";
|
||||
// unslop-check — mechanical AI-tell checker (ms-unslop subset + SYSTEM.md phrase bans).
|
||||
// Plain JS, no deps, so pi extensions (jiti) and Claude Code hook scripts (node CLI)
|
||||
// share one implementation.
|
||||
//
|
||||
// The lists live in lists.json beside this file: committed machine source with
|
||||
// per-entry provenance (which ms-unslop pattern or SYSTEM.md rule each entry
|
||||
// mechanizes), the mention convention, and the punct thresholds. The loader
|
||||
// hard-fails closed: an empty, unparseable, or schema-invalid lists.json throws,
|
||||
// and the CLI exits 2 so a broken gate is never mistaken for a clean verdict.
|
||||
//
|
||||
// CLI: node unslop-check.js <file> (or stdin)
|
||||
// exit 0 = clean, exit 1 = violations found (findings printed as JSON),
|
||||
// exit 2 = gate broken (lists.json missing/empty/invalid; error on stderr).
|
||||
// Env: UNSLOP_LISTS=<path> overrides the lists.json location (testing; reuse by
|
||||
// other harnesses sharing this file).
|
||||
//
|
||||
// Provenance note (2026-08-19): the inline lists this file carried before C1
|
||||
// moved to lists.json unchanged — 16 words, 17 phrases, 3 punct rules, 1 pattern.
|
||||
// The suite pins those counts; a list edit without a test edit is a drift signal.
|
||||
|
||||
const fs = require("node:fs");
|
||||
const path = require("node:path");
|
||||
|
||||
function stripCode(text) {
|
||||
// Fenced blocks (``` or ~~~), then inline code spans. Code is quoted material,
|
||||
// not the agent's prose style. This is also the mention convention: a banned
|
||||
// item quoted as inline code is a mention and must not flag (see lists.json).
|
||||
return text
|
||||
.replace(/```[\s\S]*?```/g, " ")
|
||||
.replace(/~~~[\s\S]*?~~~/g, " ")
|
||||
.replace(/`[^`\n]*`/g, " ");
|
||||
}
|
||||
|
||||
// ── lists.json loading and validation ─────────────────────────────────────────
|
||||
|
||||
function validateLists(data) {
|
||||
const fail = (why) => { throw new Error(`lists.json invalid: ${why}`); };
|
||||
if (typeof data !== "object" || data === null || Array.isArray(data)) fail("top level must be an object");
|
||||
for (const k of ["version", "convention", "punctPolicy", "thresholds", "words", "phrases", "punct", "patterns"]) {
|
||||
if (!(k in data)) fail(`missing key: ${k}`);
|
||||
}
|
||||
if (typeof data.version !== "number" || data.version < 1) fail("version must be a number >= 1");
|
||||
for (const k of ["convention", "punctPolicy"]) {
|
||||
if (typeof data[k] !== "string" || !data[k].trim()) fail(`${k} must be a non-empty string`);
|
||||
}
|
||||
const th = data.thresholds;
|
||||
if (typeof th !== "object" || th === null) fail("thresholds must be an object");
|
||||
// Number.isFinite, not just typeof: JSON.parse of 1e999 yields Infinity, which
|
||||
// passes typeof-number and would silently disable the punct gate (review F1).
|
||||
if (!Number.isFinite(th.punctMinCount) || th.punctMinCount < 1) fail("thresholds.punctMinCount must be a finite number >= 1");
|
||||
if (!Number.isFinite(th.punctDensityPer1k) || !(th.punctDensityPer1k > 0)) fail("thresholds.punctDensityPer1k must be a finite number > 0");
|
||||
|
||||
const seen = new Set();
|
||||
const checkEntries = (arr, kind, extra) => {
|
||||
if (!Array.isArray(arr) || arr.length === 0) fail(`${kind} must be a non-empty array`);
|
||||
arr.forEach((e, i) => {
|
||||
const at = `${kind}[${i}]`;
|
||||
if (typeof e !== "object" || e === null) fail(`${at} must be an object`);
|
||||
if (typeof e.value !== "string" || !e.value.trim()) fail(`${at}.value must be a non-empty string`);
|
||||
if (typeof e.source !== "string" || !e.source.trim()) fail(`${at}.source must be a non-empty string (pattern id or system-md)`);
|
||||
if (extra) extra(e, at, fail);
|
||||
if (seen.has(`${kind}:${e.value}`)) fail(`duplicate ${kind} value: ${e.value}`);
|
||||
seen.add(`${kind}:${e.value}`);
|
||||
});
|
||||
};
|
||||
checkEntries(data.words, "words");
|
||||
checkEntries(data.phrases, "phrases", (e, at, fail) => {
|
||||
// Phrase matching splits a lowercased haystack, so an uppercase letter in a
|
||||
// phrase value is a silently dead rule (review F3). Reject, do not silently
|
||||
// normalize: list edits should fail loud (D-a).
|
||||
if (e.value !== e.value.toLowerCase()) fail(`${at}.value must be lowercase; phrase matching lowercases the haystack: ${e.value}`);
|
||||
});
|
||||
checkEntries(data.punct, "punct", (e, at, fail) => {
|
||||
if (typeof e.label !== "string" || !e.label.trim()) fail(`${at}.label must be a non-empty string`);
|
||||
if (!Array.isArray(e.chars) || e.chars.length === 0 || !e.chars.every((c) => typeof c === "string" && c.length === 1)) {
|
||||
fail(`${at}.chars must be a non-empty array of single-char strings`);
|
||||
}
|
||||
});
|
||||
checkEntries(data.patterns, "patterns", (e, at, fail) => {
|
||||
if (typeof e.regex !== "string" || !e.regex.trim()) fail(`${at}.regex must be a non-empty string`);
|
||||
if (typeof e.flags !== "string") fail(`${at}.flags must be a string`);
|
||||
if (typeof e.detail !== "string" || !e.detail.trim()) fail(`${at}.detail must be a non-empty string`);
|
||||
try { new RegExp(e.regex, e.flags); } catch (err) { fail(`${at}.regex does not compile: ${err.message}`); }
|
||||
});
|
||||
return data;
|
||||
}
|
||||
|
||||
let cache = null;
|
||||
function loadLists(filePath) {
|
||||
if (cache && !filePath) return cache;
|
||||
const p = filePath || process.env.UNSLOP_LISTS || path.join(__dirname, "lists.json");
|
||||
let raw;
|
||||
try {
|
||||
raw = fs.readFileSync(p, "utf8");
|
||||
} catch (e) {
|
||||
throw new Error(`lists.json invalid: cannot read ${p}: ${e.message}`);
|
||||
}
|
||||
if (!raw.trim()) throw new Error(`lists.json invalid: empty file: ${p}`);
|
||||
let data;
|
||||
try {
|
||||
data = JSON.parse(raw);
|
||||
} catch (e) {
|
||||
throw new Error(`lists.json invalid: unparseable JSON: ${e.message}`);
|
||||
}
|
||||
const validated = validateLists(data);
|
||||
if (!filePath) cache = validated;
|
||||
return validated;
|
||||
}
|
||||
|
||||
// ── checker ───────────────────────────────────────────────────────────────────
|
||||
|
||||
const escapeRegex = (s) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
||||
|
||||
function checkText(raw) {
|
||||
const lists = loadLists();
|
||||
const text = stripCode(String(raw));
|
||||
// Phrase haystack: lowercased, then curly apostrophes normalized to ASCII.
|
||||
// This must be a SEPARATE string from `text`: punct counting reads the
|
||||
// original, so curly quotes still fire the punct rule (review F2 ordering).
|
||||
const phraseHay = text.toLowerCase().replace(/[\u2018\u2019]/g, "'");
|
||||
const findings = [];
|
||||
|
||||
for (const w of lists.words) {
|
||||
const re = new RegExp("\\b" + escapeRegex(w.value) + "\\b", "gi");
|
||||
const count = (text.match(re) || []).length;
|
||||
if (count > 0) findings.push({ rule: "word", detail: `banned word "${w.value}" x${count}`, count });
|
||||
}
|
||||
|
||||
for (const p of lists.phrases) {
|
||||
const count = phraseHay.split(p.value).length - 1;
|
||||
if (count > 0) findings.push({ rule: "phrase", detail: `phrase "${p.value}" x${count}`, count });
|
||||
}
|
||||
|
||||
// Density-gated punctuation. The deliberate divergence from SYSTEM.md's
|
||||
// outright em-dash ban is documented in lists.json punctPolicy, not only here.
|
||||
for (const pc of lists.punct) {
|
||||
let count = 0;
|
||||
for (const ch of pc.chars) count += text.split(ch).length - 1;
|
||||
if (count < lists.thresholds.punctMinCount) continue;
|
||||
if (count / Math.max(text.length, 1) * 1000 < lists.thresholds.punctDensityPer1k) continue;
|
||||
findings.push({ rule: "punct", detail: `${pc.label} x${count} (density-gated)`, count });
|
||||
}
|
||||
|
||||
for (const pt of lists.patterns) {
|
||||
const flags = pt.flags.includes("g") ? pt.flags : pt.flags + "g";
|
||||
const m = text.match(new RegExp(pt.regex, flags));
|
||||
const count = m ? m.length : 0;
|
||||
if (count > 0) findings.push({ rule: "pattern", detail: `"${pt.detail}" x${count}`, count });
|
||||
}
|
||||
|
||||
return { clean: findings.length === 0, findings, charsChecked: text.length };
|
||||
}
|
||||
|
||||
module.exports = { checkText, stripCode, loadLists, validateLists };
|
||||
|
||||
if (require.main === module) {
|
||||
let input;
|
||||
try {
|
||||
input = process.argv[2] ? fs.readFileSync(process.argv[2], "utf8") : fs.readFileSync(0, "utf8");
|
||||
} catch (e) {
|
||||
// An unreadable input must not exit 1: that is the violations code, and a
|
||||
// wrapper keying on rc alone would report slop-free for a file it never
|
||||
// read (review S3).
|
||||
console.error(`unslop-check: cannot read input: ${e.message}`);
|
||||
process.exit(2);
|
||||
}
|
||||
let result;
|
||||
try {
|
||||
result = checkText(input);
|
||||
} catch (e) {
|
||||
if (String(e.message).startsWith("lists.json invalid")) {
|
||||
console.error(`unslop-check: ${e.message}`);
|
||||
process.exit(2);
|
||||
}
|
||||
throw e;
|
||||
}
|
||||
console.log(JSON.stringify(result, null, 2));
|
||||
process.exit(result.clean ? 0 : 1);
|
||||
}
|
||||
@@ -1,3 +1,5 @@
|
||||
# POC user
|
||||
|
||||
This is an isolated local runtime test.
|
||||
|
||||
The user's name is Jason.
|
||||
Reference in New Issue
Block a user