Compare commits
18
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
e175616885 | ||
|
|
6955717612 | ||
|
|
88d9cf750f | ||
|
|
c038706eed | ||
|
|
8622c9d826 | ||
|
|
88eef507b0 | ||
|
|
439bea6915 | ||
|
|
fad8a4718c | ||
|
|
2ff49adff4 | ||
|
|
44c476ebbf | ||
|
|
cde480eb60 | ||
|
|
5808248707 | ||
|
|
afd5827db8 | ||
|
|
d9cc990376 | ||
|
|
83c4e9851e | ||
|
|
22508170a2 | ||
|
|
90a67d050e | ||
|
|
24bdef75fa |
+116
@@ -221,4 +221,120 @@ Release model and safe updates verified by drills. `main` merged with M3 and tag
|
||||
|
||||
Adapter seam verified; harness boundary is now additive by construction. `main` merged with M4 and tagged `adapter-seam-v1`.
|
||||
|
||||
---
|
||||
|
||||
## Phase 9: Workspaces + capability envelope (M5)
|
||||
|
||||
### Entry 9.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Optional task workspace (absent / ":run" ephemeral / named persistent) and capabilities.tools allowlist (pi built-ins); runner plumbing via MOSAIC_WORKSPACE/MOSAIC_TOOLS; pi adapter maps to cwd + --tools; mock adapter logs delivered MOSAIC_* vars for deterministic assertions (Gitea #20, #21).
|
||||
- Reason: Agents that only answer text cannot do work; the workspace+tools pair is the smallest real capability step, bounded by the container.
|
||||
- Expected result: plumbing asserted via run-record stderr; live pi writes a host-visible file.
|
||||
|
||||
### Entry 9.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: build; mock plumbing run; validate negatives; full task suite; live workspace demo.
|
||||
- Observed result: task suite 32/32; MOSAIC_WORKSPACE/MOSAIC_TOOLS asserted in run record; dataRoot/workspaces/<name> created host-side; live pi used bash to write proof.txt into the demo workspace (host-visible).
|
||||
- Failure or correction: (1) batch edit dropped SUPPORTED_TOOLS const (runtime ReferenceError, exit 1 instead of 2) — restored; (2) mock env dump used `export` which this dash prints as `export K='v'` — switched to `env`; (3) selftest fed a mismatching mock response to the plain-task case — test bug, fixed.
|
||||
|
||||
## Phase 10: Named sessions (M6)
|
||||
|
||||
### Entry 10.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Optional task.session name -> persistent session dir dataRoot/sessions/<name> via pi --session-dir; resume most recent with -c when present; isolation per name; teach/recall demo fixtures (Gitea #22, #23).
|
||||
- Reason: L1 persistence is the prerequisite for any multi-step agent work.
|
||||
- Expected result: session dir populated after first run; second run recalls taught context.
|
||||
|
||||
### Entry 10.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: build; mock plumbing run; live teach/recall E2E; task suite.
|
||||
- Observed result: task suite 32/32; teach run replied REMEMBERED and session JSONL persisted host-side; recall run resumed (-c) and answered exactly 'mosaico'; single continued session file (no duplicate sessions).
|
||||
- Failure or correction: none. Design note: ephemeral (--no-session) remains the default when no session is declared.
|
||||
|
||||
## Phase 11: Operator ergonomics (M7)
|
||||
|
||||
### Entry 11.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: mosaic-task.mjs show <runId> (full record + snapshots + artifacts, traversal-safe), list with workspace/session columns, RELEASE -> 0.0.6, package + health-gated activate, docs (Gitea #24).
|
||||
- Reason: Run records are only as valuable as they are inspectable; release activation closes the loop on container-content changes.
|
||||
- Expected result: show works for real/missing/traversal ids; suites green; 0.0.6 active.
|
||||
|
||||
### Entry 11.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: build; show on real/missing/traversal ids; full sweep; release package + activate.
|
||||
- Observed result: config 24/24, task 32/32, release 14/14, verify PASS; 0.0.6 packaged and activated via exact-marker health gate.
|
||||
- Failure or correction: showRun initially rejected valid run ids (lowercase-only regex vs uppercase timestamp) and crashed on missing ids (uncaught readdir) — both fixed and covered.
|
||||
|
||||
## Autonomous run result
|
||||
|
||||
M5 tagged `workspace-capabilities-v1`, M6 tagged `sessions-v1`, M7 tagged `operator-ergonomics-v1`; release 0.0.6 active. Tracker: docs/plans/2026-09-03_autonomous-run.md.
|
||||
|
||||
---
|
||||
|
||||
## Phase 12: Conductor loop — self-orchestration (M8)
|
||||
|
||||
### Entry 12.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Stand up the poor-man orchestration loop per docs/plans/CONDUCTOR.md: conductor (host, holds git) mirrors the repo into a worker workspace, dispatches a headless pi worker (session worker-1, tools read/write/edit/bash) to implement retry <runId>, reviews the diff, integrates, verifies (Gitea #25, #26, #27).
|
||||
- Reason: The owner asked for circular task processing with agent workers; the stack now has every primitive needed — this proves it on the stack itself.
|
||||
- Expected result: worker-authored retry merged with suites green and a live retry verified.
|
||||
|
||||
### Entry 12.2 — after
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Commands run: repo mirror clone; worker dispatch (tasks/worker-retry.json); diff review; apply; live retry; refinement dispatch (tasks/worker-retry-refine.json); second review; reverse+reapply combined patch; conductor interpolation fix; live retry; full sweep.
|
||||
- Observed result:
|
||||
- Worker round 1: implemented retry correctly per spec in 2m28s; diff reviewed clean.
|
||||
- Live retry exposed a spec gap (direct invocation lacks launcher env exports).
|
||||
- Worker round 2 (same session, 59s): made spawnEnv self-sufficient, but used PI_* names where compose interpolates MOSAIC_*.
|
||||
- Conductor hotfix: 3-line rename to MOSAIC_PROVIDER/MOSAIC_MODEL/MOSAIC_DATA_ROOT (too trivial for a worker round).
|
||||
- Final: live retry succeeded (replied REMEMBERED, new run recorded); all suites green.
|
||||
- Failure or correction: three rounds total — one spec gap (conductor), one naming mismatch (worker), one trivial rename (conductor). Each was caught by mechanical verification (run record stderr), never by hope.
|
||||
- Attribution: feature authored by headless pi worker (glm-5.3-flash) in sessions worker-1; conductor reviewed, integrated, and hotfixed.
|
||||
|
||||
## Result (M8)
|
||||
|
||||
Conductor loop proven end-to-end on the stack itself. `main` merged with M8; release 0.0.6 remains active (retry is host-side only, no image change).
|
||||
|
||||
---
|
||||
|
||||
## Phase 13: Mission-level capability policy (M9)
|
||||
|
||||
### Entry 13.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: Missions may declare capabilities.tools as governing constraints (Gitea #30); merge semantics = least-privilege intersection (task narrows, never widens; empty intersection = tool-free run). Host-side only.
|
||||
- Reason: First mechanical restriction layer — the trust model becomes enforced, not instructed.
|
||||
- Expected result: all four merge cases asserted from run evidence; suites green.
|
||||
|
||||
### Entry 13.2 — after
|
||||
|
||||
- Observed: four merge cases verified deterministically via run-record stderr (mission-only, task-only, narrowed, emptied); invalid mission capabilities exit 2; task suite 41/41.
|
||||
- Failure or correction: selftest harness could not express ABSENT vs EMPTY fields via its printf helper — fixed with an ABSENT marker; two suite config-leak defects fixed (per-command env scoping). Product unaffected.
|
||||
|
||||
## Phase 14: Session forking (M11)
|
||||
|
||||
### Entry 14.1 — before
|
||||
|
||||
- Timestamp: 2026-09-03
|
||||
- Intended action: sessionForkFrom task field branches the source session's newest file (pi --fork) into the target session dir; ancestor untouched; RELEASE -> 0.0.7 with health-gated activation (Gitea #33).
|
||||
- Reason: Owner flagged conversation forking from a common ancestor as a desired property; pi JSONL trees make it native.
|
||||
- Expected result: forked child recalls ancestor context; ancestor file untouched; suites green; 0.0.7 active.
|
||||
|
||||
### Entry 14.2 — after
|
||||
|
||||
- Observed: mock plumbing asserts fork source + target delivery; validation rejects fork-without-target and self-fork (exit 2); ghost source exits 4; live fork: child recalled 'mosaico' from ancestor context while the ancestor session file remained untouched (file-level assertion); suites 58/24/14 + verify green; 0.0.7 packaged and activated via health gate.
|
||||
- Failure or correction: retryRun-style self-assignment bug in validation (compared null to target) — caught by negative test, fixed.
|
||||
|
||||
## Result (M11)
|
||||
|
||||
Session forking verified. `main` merged with M11, tagged `session-fork-v1`; release 0.0.7 active.
|
||||
|
||||
|
||||
|
||||
@@ -119,6 +119,31 @@ adapters/<name>/adapter.sh env in: MOSAIC_SYSTEM_PROMPT_FILE, MOSAIC_REQUEST,
|
||||
|
||||
See `adapters/README.md` for the full contract.
|
||||
|
||||
## Workspaces, capabilities, sessions (M5/M6)
|
||||
|
||||
Optional task fields extend what an agent can do — all defaulting to the previous behavior:
|
||||
|
||||
```json
|
||||
{
|
||||
"workspace": "demo", // ":run" ephemeral, or persistent dataRoot/workspaces/<name>
|
||||
"capabilities": { "tools": ["bash", "read"] }, // pi tool allowlist; absent = no tools
|
||||
"session": "demo" // persistent session at dataRoot/sessions/<name>
|
||||
}
|
||||
```
|
||||
|
||||
- The adapter runs inside the workspace; files it writes are host-visible (`dataRoot/workspaces/<name>`).
|
||||
- Sessions persist via pi's documented `--session-dir`; a follow-up run in the same session resumes the conversation (`-c`) and can recall prior context. Distinct names never share state. Ephemeral (`--no-session`) remains the default when no session is declared.
|
||||
- Selection authority: config for adapter/provider/model; the task file for workspace/capabilities/session.
|
||||
|
||||
Inspect anything:
|
||||
|
||||
```bash
|
||||
node scripts/mosaic-task.mjs list # runs with task/workspace/session columns
|
||||
node scripts/mosaic-task.mjs show <runId> # full record + snapshots + artifacts
|
||||
```
|
||||
|
||||
Demo fixtures: `tasks/workspace-demo.json`, `tasks/session-demo-1.json` + `tasks/session-demo-2.json`.
|
||||
|
||||
See `docs/plans/2026-09-02_atomic-mosaic-foundation.md` for the full plan.
|
||||
|
||||
Inside the container:
|
||||
|
||||
@@ -18,11 +18,16 @@ if [ -n "${MOSAIC_WORKSPACE:-}" ]; then
|
||||
cd "$MOSAIC_WORKSPACE"
|
||||
fi
|
||||
|
||||
# Session (M6): persistent named session directory; resume the most recent
|
||||
# session in that directory when one exists (pi documented flags).
|
||||
# Default remains ephemeral (--no-session) when no session is declared.
|
||||
# Session (M6/M11): default ephemeral (--no-session). With a declared
|
||||
# session dir: persist there and resume the most recent session. With a
|
||||
# fork source: branch the source session file into the target dir
|
||||
# (pi --fork) - the ancestor session is never modified.
|
||||
SESSION_FLAGS="--no-session"
|
||||
if [ -n "${MOSAIC_SESSION_DIR:-}" ]; then
|
||||
if [ -n "${MOSAIC_SESSION_FORK:-}" ]; then
|
||||
[ -n "${MOSAIC_SESSION_DIR:-}" ] || { echo "pi adapter: session fork requires MOSAIC_SESSION_DIR" >&2; exit 2; }
|
||||
mkdir -p "$MOSAIC_SESSION_DIR"
|
||||
SESSION_FLAGS="--fork $MOSAIC_SESSION_FORK --session-dir $MOSAIC_SESSION_DIR"
|
||||
elif [ -n "${MOSAIC_SESSION_DIR:-}" ]; then
|
||||
mkdir -p "$MOSAIC_SESSION_DIR"
|
||||
SESSION_FLAGS="--session-dir $MOSAIC_SESSION_DIR"
|
||||
if [ -n "$(ls -A "$MOSAIC_SESSION_DIR" 2>/dev/null)" ]; then
|
||||
|
||||
+2
-1
@@ -18,8 +18,9 @@ services:
|
||||
# Workspace + capabilities (set by the task runner; M5)
|
||||
MOSAIC_WORKSPACE: ${MOSAIC_WORKSPACE:-}
|
||||
MOSAIC_TOOLS: ${MOSAIC_TOOLS:-}
|
||||
# Persistent named session dir (set by the task runner; M6)
|
||||
# Persistent named session dir + optional fork source (M6/M11)
|
||||
MOSAIC_SESSION_DIR: ${MOSAIC_SESSION_DIR:-}
|
||||
MOSAIC_SESSION_FORK: ${MOSAIC_SESSION_FORK:-}
|
||||
# mock adapter only: verbatim response for deterministic seam tests
|
||||
MOSAIC_MOCK_RESPONSE: ${MOSAIC_MOCK_RESPONSE:-}
|
||||
# Documented container auth alternative: provider API key via
|
||||
|
||||
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"policyVersion": 1,
|
||||
"autoApply": {
|
||||
"enabled": true,
|
||||
"allowedPaths": [
|
||||
"scripts/**",
|
||||
"docs/**",
|
||||
"tasks/**",
|
||||
"missions/**",
|
||||
"adapters/**",
|
||||
"README.md"
|
||||
],
|
||||
"suites": [
|
||||
"test-config",
|
||||
"test-task",
|
||||
"test-release"
|
||||
]
|
||||
}
|
||||
}
|
||||
@@ -13,8 +13,8 @@ Advance the Mosaic Stack rebuild several verified layers in one batch, focused o
|
||||
|---|---|---|
|
||||
| M5 | Task workspaces + capability envelope (tools allowlist) | DONE |
|
||||
| M6 | Named sessions — persistence and resume (L1) | DONE |
|
||||
| M7 | Operator ergonomics: run inspection commands | DONE (partial by design) |
|
||||
| — | Release 0.0.6 packaged + health-gated activated | DONE |
|
||||
| M7 | Operator ergonomics: run inspection commands | DONE |
|
||||
| — | Releases 0.0.5+0.0.6 packaged; 0.0.6 health-gated activated | DONE |
|
||||
|
||||
Explicitly deferred (do not mistake for forgotten):
|
||||
- Claude/Codex/OpenCode adapters (owner: focus on Pi for now)
|
||||
@@ -60,14 +60,17 @@ Explicitly deferred (do not mistake for forgotten):
|
||||
6. Inspect any run: `node scripts/mosaic-task.mjs show <runId>`
|
||||
7. Gitea: milestones M5/M6/M7 closed; issues referenced by merge commits
|
||||
|
||||
## Results (to be filled at end of run)
|
||||
## Results
|
||||
|
||||
- M5: pending
|
||||
- M6: pending
|
||||
- M7: pending
|
||||
- Release activation: pending
|
||||
- Final suite counts: pending
|
||||
- Corrections encountered: pending
|
||||
- M5 merged on `main` (merge commit `ddb1554`), tagged `workspace-capabilities-v1`
|
||||
- M6 merged on `main` (merge commit `4e2a413`), tagged `sessions-v1`
|
||||
- M7 merged on `main`, tagged `operator-ergonomics-v1`
|
||||
- Release 0.0.6 packaged, health-gated activated, full sweep green
|
||||
- Suites at end of run: config 24/24, task 32/32, release 14/14, verify PASS
|
||||
- Build log: Phases 9 (M5), 10 (M6), 11 (M7) appended with corrections
|
||||
- Corrections encountered: dropped constant from a failed atomic edit batch (SUPPORTED_TOOLS); dash `export` output format vs `env`; three selftest authoring defects; showRun id-regex case sensitivity + missing-run crash. All fixed and covered by tests.
|
||||
- Commits pushed incrementally; nothing left uncommitted
|
||||
- Live proof: workspace file host-visible; session teach/recall ('mosaico') verified
|
||||
|
||||
## Next steps after this run (not started)
|
||||
|
||||
|
||||
@@ -0,0 +1,51 @@
|
||||
# Conductor protocol — poor-man orchestration loop
|
||||
|
||||
How the stack orchestrates headless pi workers to do work on itself.
|
||||
|
||||
## Roles
|
||||
|
||||
| Role | Runs where | Powers | Never has |
|
||||
|---|---|---|---|
|
||||
| **Conductor** | host (assistant or owner) | git (clone/commit/push), task dispatch, review, verification suites, Gitea | nothing new |
|
||||
| **Worker** | container (headless pi via `scripts/run-task.sh`) | read/write/edit/bash inside its workspace; persistent session on request | git credentials, docker socket, host filesystem |
|
||||
|
||||
## The loop
|
||||
|
||||
1. **Decompose**: conductor turns a goal into worker tasks small enough to
|
||||
specify completely in one prompt (file paths, acceptance criteria, style
|
||||
constraints, verification the worker can run itself, e.g. `node --check`).
|
||||
2. **Mirror**: conductor maintains the repo clone at
|
||||
`<dataRoot>/workspaces/stack-repo` (host-side git; workers see it read-write
|
||||
through their workspace mount).
|
||||
3. **Dispatch**: `scripts/run-task.sh run <worker-task.json>` — worker edits the
|
||||
clone. Session name `worker-<n>` keeps continuity across refinement rounds.
|
||||
4. **Extract**: `git -C <workspace> diff > patch` — the worker's entire output
|
||||
is a reviewable diff. Run record (result.json, stderr.txt) is the receipt.
|
||||
5. **Review**: conductor reads the diff line by line. Bad output → refine the
|
||||
prompt, re-dispatch (same session: "your patch had these problems…").
|
||||
6. **Integrate**: conductor applies the patch to the real repo, runs the full
|
||||
suites, commits and pushes. Suites failing → revert apply, back to step 5.
|
||||
7. **Record**: update CURRENT.md, BUILD-LOG, close the Gitea issue.
|
||||
|
||||
## Guardrails
|
||||
|
||||
- Workers never receive credentials; they never run git; they never leave the
|
||||
workspace (container is the boundary; tools allowlist is the gate).
|
||||
- Every worker diff is reviewed by the conductor before integration. No
|
||||
auto-apply. (Auto-apply would be a capability-policy decision for later.)
|
||||
- Verification is mechanical: suites + `node --check` / `bash -n` gates.
|
||||
- Recursive decomposition = "fail → smaller task", never "hope."
|
||||
|
||||
## Worker task template
|
||||
|
||||
```json
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-worker-<name>",
|
||||
"prompt": "<full spec: goal, files, constraints, acceptance, self-checks>",
|
||||
"workspace": "stack-repo",
|
||||
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
|
||||
"session": "worker-1",
|
||||
"timeoutSeconds": 600
|
||||
}
|
||||
```
|
||||
@@ -0,0 +1,47 @@
|
||||
# CURRENT — single source of "what happens next"
|
||||
|
||||
This file always names exactly one next action. Any "continue" / "next" /
|
||||
"proceed" message means: execute the action below, fully (implement → test →
|
||||
verify against its acceptance criteria → commit → push → close the issue →
|
||||
update this file to the next action). No ambiguity, no re-planning.
|
||||
|
||||
## Next action
|
||||
|
||||
Owner review of M11 (session forking) — then name the next target.
|
||||
|
||||
## Queue (ordered, not started)
|
||||
|
||||
1. Second real adapter (parked — owner focused on Pi)
|
||||
2. Auto-apply policy for worker patches (substrate exists)
|
||||
3. UX/DX backlog #31 (grep highlight vs status colors — likely closed by owner's unalias)
|
||||
|
||||
## Rules
|
||||
|
||||
- One action in flight. Update this file at the END of every action.
|
||||
- Blocked? Move the item to "Blocked" below with the reason and stop.
|
||||
- Completed actions move to the log at the bottom (date + issue + result).
|
||||
|
||||
## Blocked
|
||||
|
||||
(none)
|
||||
|
||||
## Completed log
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
|
||||
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
|
||||
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
|
||||
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
|
||||
|
||||
- 2026-09-03 — M5 workspaces + capability envelope (#20, #21) — merged, tests green
|
||||
- 2026-09-03 — M6 named sessions + resume (#22, #23) — merged, teach/recall verified
|
||||
- 2026-09-03 — M7 run inspection + release 0.0.6 (#24) — merged, activated
|
||||
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
|
||||
- 2026-09-03 — retry lineage + relative mission resolution (#28) — merged, 36/24/14 suites + verify green
|
||||
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection, 41/36/14 + verify green
|
||||
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
|
||||
- 2026-09-03 — retry lineage + relative mission resolution (#28) — merged, 36/24/14 suites + verify green
|
||||
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
|
||||
Executable
+139
@@ -0,0 +1,139 @@
|
||||
#!/usr/bin/env bash
|
||||
# Conductor auto-apply: integrate a worker's patch under the declared policy.
|
||||
#
|
||||
# Usage: scripts/conductor-apply.sh <runId> [--dry-run]
|
||||
#
|
||||
# Policy (conductor-policy.json in the target repo, strictly validated):
|
||||
# autoApply.enabled master switch
|
||||
# autoApply.allowedPaths glob allowlist ('dir/**' = everything under dir)
|
||||
# autoApply.suites suite scripts that must pass AFTER applying
|
||||
#
|
||||
# Gate sequence: succeeded run record -> clean target tree -> diff extracted
|
||||
# from the worker workspace -> allowlist -> syntax gates (node/bash/json) ->
|
||||
# apply -> policy suites -> commit with attribution. ANY failure reverts the
|
||||
# working tree and exits nonzero. Push is never automatic.
|
||||
#
|
||||
# Environment:
|
||||
# MOSAIC_APPLY_TARGET repo root to apply into (default: this project)
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
TARGET_ROOT="${MOSAIC_APPLY_TARGET:-$(cd "$SCRIPT_DIR/.." && pwd)}"
|
||||
RUN_ID="${1:?usage: conductor-apply.sh <runId> [--dry-run]}"
|
||||
DRY_RUN="no"
|
||||
[ "${2:-}" = "--dry-run" ] && DRY_RUN="yes"
|
||||
|
||||
cd "$TARGET_ROOT"
|
||||
|
||||
fail() { echo "conductor-apply: $*" >&2; exit "${2:-1}"; }
|
||||
[ -d .git ] || fail "target is not a git repository: $TARGET_ROOT" 4
|
||||
[ -f conductor-policy.json ] || fail "no conductor-policy.json in target" 2
|
||||
|
||||
# ---- policy (strict) ----
|
||||
POLICY_JSON="$(node -e '
|
||||
const fs = require("fs");
|
||||
const p = JSON.parse(fs.readFileSync("conductor-policy.json", "utf8"));
|
||||
if (p.policyVersion !== 1) process.exit(3);
|
||||
if (!p.autoApply || typeof p.autoApply.enabled !== "boolean" || !Array.isArray(p.autoApply.allowedPaths) || !Array.isArray(p.autoApply.suites)) process.exit(3);
|
||||
for (const g of p.autoApply.allowedPaths) {
|
||||
if (typeof g !== "string" || !/^[A-Za-z0-9_.*/-]+$/.test(g) || g.startsWith("/") || g.includes("..")) process.exit(3);
|
||||
}
|
||||
console.log(JSON.stringify(p.autoApply));
|
||||
')" || fail "invalid conductor-policy.json" 2
|
||||
|
||||
ENABLED="$(node -e 'console.log(JSON.parse(process.argv[1]).enabled)' "$POLICY_JSON")"
|
||||
[ "$ENABLED" = "true" ] || fail "auto-apply is disabled by policy" 2
|
||||
|
||||
# ---- run record ----
|
||||
DATA_ROOT="$(node scripts/mosaic-config.mjs validate | node -e 'let d="";process.stdin.on("data",c=>d+=c).on("end",()=>console.log(JSON.parse(d).dataRoot))')"
|
||||
RESULT_FILE="$DATA_ROOT/runs/$RUN_ID/result.json"
|
||||
[ -f "$RESULT_FILE" ] || fail "run not found: $RUN_ID" 4
|
||||
|
||||
node -e '
|
||||
const r = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
|
||||
process.exit(r.status === "succeeded" ? 0 : 1);
|
||||
' "$RESULT_FILE" || fail "run $RUN_ID did not succeed; refusing to auto-apply"
|
||||
|
||||
WORKSPACE_NAME="$(node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.workspace||"")' "$RESULT_FILE")"
|
||||
[ -n "$WORKSPACE_NAME" ] || fail "run has no workspace; nothing to integrate" 4
|
||||
WORKSPACE="$DATA_ROOT/workspaces/$WORKSPACE_NAME"
|
||||
[ -d "$WORKSPACE/.git" ] || fail "workspace is not a git clone: $WORKSPACE" 4
|
||||
|
||||
# ---- extract diff (tracked + intent-to-add) ----
|
||||
git -C "$WORKSPACE" add -N . >/dev/null 2>&1 || true
|
||||
DIFF_FILE="$(mktemp)"
|
||||
trap 'rm -f "$DIFF_FILE"' EXIT
|
||||
git -C "$WORKSPACE" diff > "$DIFF_FILE"
|
||||
if [ ! -s "$DIFF_FILE" ]; then
|
||||
fail "workspace has no changes to apply"
|
||||
fi
|
||||
|
||||
# ---- allowlist ----
|
||||
mapfile -t CHANGED < <(git -C "$WORKSPACE" diff --name-only)
|
||||
GLOBS="$(node -e 'const a=JSON.parse(process.argv[1]).allowedPaths;console.log(a.join("\n"))' "$POLICY_JSON")"
|
||||
REFUSED=""
|
||||
for f in "${CHANGED[@]}"; do
|
||||
ok="no"
|
||||
while IFS= read -r g; do
|
||||
[ -z "$g" ] && continue
|
||||
case "$f" in
|
||||
$g) ok="yes"; break ;;
|
||||
esac
|
||||
done <<< "$GLOBS"
|
||||
[ "$ok" = "yes" ] || REFUSED="$REFUSED $f"
|
||||
done
|
||||
if [ -n "$REFUSED" ]; then
|
||||
echo "conductor-apply: refusing - files outside policy allowlist:$REFUSED" >&2
|
||||
echo "conductor-apply: patch preserved at $DIFF_FILE for manual review" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ---- syntax gates (on workspace files, pre-apply) ----
|
||||
for f in "${CHANGED[@]}"; do
|
||||
case "$f" in
|
||||
*.mjs) node --check "$WORKSPACE/$f" || fail "syntax gate failed (node): $f" 1 ;;
|
||||
*.sh) bash -n "$WORKSPACE/$f" || fail "syntax gate failed (bash): $f" 1 ;;
|
||||
*.json) node -e 'JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"))' "$WORKSPACE/$f" || fail "syntax gate failed (json): $f" 1 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
if [ "$DRY_RUN" = "yes" ]; then
|
||||
echo "conductor-apply (dry-run): would apply $(echo "${#CHANGED[@]}") file(s) from $RUN_ID:"
|
||||
printf ' %s\n' "${CHANGED[@]}"
|
||||
echo "conductor-apply (dry-run): suites would run: $(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(", "))' "$POLICY_JSON")"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# ---- apply ----
|
||||
[ -z "$(git status --porcelain)" ] || fail "target tree is not clean; refusing to mix states"
|
||||
git apply "$DIFF_FILE" || fail "git apply failed"
|
||||
|
||||
# ---- policy suites ----
|
||||
SUITES="$(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(" "))' "$POLICY_JSON")"
|
||||
SUITES_OK="yes"
|
||||
for s in $SUITES; do
|
||||
case "$s" in
|
||||
test-[a-z]*) : ;; # shape guard; existence checked next
|
||||
*) echo "conductor-apply: refusing suspicious suite name: $s" >&2; SUITES_OK="no"; break ;;
|
||||
esac
|
||||
[ -x "scripts/$s.sh" ] || { echo "conductor-apply: suite script missing: scripts/$s.sh" >&2; SUITES_OK="no"; break; }
|
||||
if ! bash "scripts/$s.sh" >/dev/null 2>&1; then
|
||||
echo "conductor-apply: suite failed: $s" >&2
|
||||
SUITES_OK="no"
|
||||
break
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$SUITES_OK" != "yes" ]; then
|
||||
git apply -R "$DIFF_FILE" && echo "conductor-apply: changes REVERTED (suites failed)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# ---- commit with attribution ----
|
||||
git add -A
|
||||
git commit -q -m "feat(worker): auto-applied patch from run $RUN_ID
|
||||
|
||||
Authored-by: pi worker (run $RUN_ID, workspace $WORKSPACE_NAME)
|
||||
Applied-under: conductor-policy v1 (allowlist + syntax gates + suites)"
|
||||
echo "conductor-apply: applied and committed run $RUN_ID ($(echo "${#CHANGED[@]}") file(s)); suites: $SUITES"
|
||||
echo "conductor-apply: NOT pushed - push remains an explicit act."
|
||||
+257
-8
@@ -28,6 +28,7 @@
|
||||
*/
|
||||
|
||||
import fs from "node:fs";
|
||||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import process from "node:process";
|
||||
import { randomBytes } from "node:crypto";
|
||||
@@ -83,7 +84,7 @@ function validateId(value, what) {
|
||||
|
||||
function validateMission(document, file) {
|
||||
if (!isPlainObject(document)) fail(2, "mission must be a JSON object");
|
||||
rejectUnknownKeys(document, ["missionVersion", "id", "objective", "directives"], "mission");
|
||||
rejectUnknownKeys(document, ["missionVersion", "id", "objective", "directives", "capabilities"], "mission");
|
||||
if (document.missionVersion !== 1) {
|
||||
fail(2, `unsupported missionVersion: ${JSON.stringify(document.missionVersion)} (supported: 1)`);
|
||||
}
|
||||
@@ -101,12 +102,34 @@ function validateMission(document, file) {
|
||||
return d;
|
||||
});
|
||||
}
|
||||
return { missionVersion: document.missionVersion, id: document.id, objective: document.objective, directives };
|
||||
|
||||
// Governing capability constraints (M9): same validation as task
|
||||
// capabilities; semantically these BOUND tasks (least-privilege
|
||||
// intersection at run time), never grant beyond them.
|
||||
let capabilities = null;
|
||||
if (document.capabilities !== undefined && document.capabilities !== null) {
|
||||
if (!isPlainObject(document.capabilities)) fail(2, 'mission "capabilities" must be a JSON object');
|
||||
rejectUnknownKeys(document.capabilities, ["tools"], 'mission "capabilities"');
|
||||
if (!Array.isArray(document.capabilities.tools) || document.capabilities.tools.length === 0) {
|
||||
fail(2, 'mission "capabilities.tools" must be a non-empty array of tool names');
|
||||
}
|
||||
const seen = new Set();
|
||||
for (const tool of document.capabilities.tools) {
|
||||
if (!SUPPORTED_TOOLS.includes(tool)) {
|
||||
fail(2, `unsupported mission tool: ${JSON.stringify(tool)} (supported: ${SUPPORTED_TOOLS.join(", ")})`);
|
||||
}
|
||||
if (seen.has(tool)) fail(2, `duplicate tool in mission capabilities.tools: ${tool}`);
|
||||
seen.add(tool);
|
||||
}
|
||||
capabilities = { tools: [...seen] };
|
||||
}
|
||||
|
||||
return { missionVersion: document.missionVersion, id: document.id, objective: document.objective, directives, capabilities };
|
||||
}
|
||||
|
||||
function validateTask(document, file) {
|
||||
if (!isPlainObject(document)) fail(2, "task must be a JSON object");
|
||||
rejectUnknownKeys(document, ["taskVersion", "id", "prompt", "mission", "expectExact", "timeoutSeconds", "workspace", "capabilities", "session"], "task");
|
||||
rejectUnknownKeys(document, ["taskVersion", "id", "prompt", "mission", "expectExact", "timeoutSeconds", "workspace", "capabilities", "session", "sessionForkFrom"], "task");
|
||||
if (document.taskVersion !== 1) {
|
||||
fail(2, `unsupported taskVersion: ${JSON.stringify(document.taskVersion)} (supported: 1)`);
|
||||
}
|
||||
@@ -169,6 +192,23 @@ function validateTask(document, file) {
|
||||
session = document.session;
|
||||
}
|
||||
|
||||
// Session fork (M11): optional source session whose newest session file
|
||||
// is branched (pi --fork) into the target session dir. Requires session.
|
||||
let sessionForkFrom = null;
|
||||
if (document.sessionForkFrom !== undefined && document.sessionForkFrom !== null) {
|
||||
if (typeof document.sessionForkFrom !== "string" || document.sessionForkFrom.length === 0) {
|
||||
fail(2, 'task "sessionForkFrom" must be a non-empty string when present');
|
||||
}
|
||||
validateId(document.sessionForkFrom, "task sessionForkFrom");
|
||||
if (!session) {
|
||||
fail(2, 'task "sessionForkFrom" requires "session" (the fork target) to be set');
|
||||
}
|
||||
if (document.sessionForkFrom === session) {
|
||||
fail(2, 'task "sessionForkFrom" must differ from "session" (cannot fork onto itself)');
|
||||
}
|
||||
sessionForkFrom = document.sessionForkFrom;
|
||||
}
|
||||
|
||||
// Capabilities (M5): optional tools allowlist mapped by adapters to their
|
||||
// native permission flags. Absent = no tools.
|
||||
let tools = null;
|
||||
@@ -201,6 +241,7 @@ function validateTask(document, file) {
|
||||
workspace,
|
||||
tools,
|
||||
session,
|
||||
sessionForkFrom,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -233,7 +274,7 @@ function writeOnce(file, content) {
|
||||
}
|
||||
}
|
||||
|
||||
function runTask(taskFile) {
|
||||
function runTask(taskFile, options = {}) {
|
||||
const resolved = JSON.parse(
|
||||
spawnSync(process.execPath, [path.join(PROJECT_ROOT, "scripts", "mosaic-config.mjs"), "validate"], {
|
||||
cwd: PROJECT_ROOT,
|
||||
@@ -263,6 +304,15 @@ function runTask(taskFile) {
|
||||
// CONTAINER path (dataRoot maps to /var/lib/mosaic in the image).
|
||||
const spawnEnv = { ...process.env };
|
||||
spawnEnv.MOSAIC_ADAPTER = resolved.execution.adapter;
|
||||
// Self-sufficient env: direct invocation (e.g. `retry`) skips the shell
|
||||
// launcher exports, so derive them from the resolved config and release.
|
||||
// (Names here are the compose interpolation consumers, not PI_*.)
|
||||
spawnEnv.MOSAIC_PROVIDER = resolved.execution.provider;
|
||||
spawnEnv.MOSAIC_MODEL = resolved.execution.model;
|
||||
spawnEnv.MOSAIC_DATA_ROOT = resolved.dataRoot;
|
||||
const release = fs.readFileSync(path.join(PROJECT_ROOT, "RELEASE"), "utf8").trim();
|
||||
const piVersion = JSON.parse(fs.readFileSync(path.join(PROJECT_ROOT, "package.json"), "utf8")).dependencies["@earendil-works/pi-coding-agent"];
|
||||
spawnEnv.MOSAIC_IMAGE_TAG = `mosaic-poc-agent:${piVersion}-r${release}`;
|
||||
if (task.missionSnapshot) {
|
||||
const relative = path.relative(resolved.dataRoot, runDir);
|
||||
if (relative.startsWith("..") || path.isAbsolute(relative)) {
|
||||
@@ -281,7 +331,24 @@ function runTask(taskFile) {
|
||||
workspaceContainerPath = `/var/lib/mosaic/workspaces/${task.workspace}`;
|
||||
}
|
||||
if (workspaceContainerPath) spawnEnv.MOSAIC_WORKSPACE = workspaceContainerPath;
|
||||
spawnEnv.MOSAIC_TOOLS = task.tools ? task.tools.join(",") : "";
|
||||
|
||||
// Capability policy (M9): least-privilege intersection. A task may narrow
|
||||
// a mission's tool grant, never widen it. Empty intersection = tool-free.
|
||||
let effectiveTools = task.tools;
|
||||
let policyNote = null;
|
||||
if (task.missionSnapshot?.capabilities) {
|
||||
const missionTools = task.missionSnapshot.capabilities.tools;
|
||||
if (effectiveTools) {
|
||||
effectiveTools = effectiveTools.filter((t) => missionTools.includes(t));
|
||||
if (effectiveTools.length === 0) {
|
||||
policyNote = `capability policy: mission ${task.missionSnapshot.id} and task request no tools in common -> tool-free run`;
|
||||
}
|
||||
} else {
|
||||
effectiveTools = [...missionTools];
|
||||
}
|
||||
}
|
||||
if (policyNote) process.stderr.write(`mosaic-task: ${policyNote}\n`);
|
||||
spawnEnv.MOSAIC_TOOLS = effectiveTools ? effectiveTools.join(",") : "";
|
||||
|
||||
// Session (M6): persistent named session dir, passed as container path.
|
||||
if (task.session) {
|
||||
@@ -289,6 +356,26 @@ function runTask(taskFile) {
|
||||
spawnEnv.MOSAIC_SESSION_DIR = `/var/lib/mosaic/sessions/${task.session}`;
|
||||
}
|
||||
|
||||
// Session fork (M11): resolve the source session's newest file; pi --fork
|
||||
// branches it into the target dir without modifying the ancestor.
|
||||
if (task.sessionForkFrom) {
|
||||
const sourceDir = path.join(resolved.dataRoot, "sessions", task.sessionForkFrom);
|
||||
let sources = [];
|
||||
try {
|
||||
sources = fs.readdirSync(sourceDir).filter((f) => f.endsWith(".jsonl")).sort();
|
||||
} catch {
|
||||
fail(4, `cannot fork: source session dir not found: ${sourceDir}`);
|
||||
}
|
||||
if (sources.length === 0) {
|
||||
fail(4, `cannot fork: source session '${task.sessionForkFrom}' has no session files`);
|
||||
}
|
||||
const relative = path.relative(resolved.dataRoot, path.join(sourceDir, sources[sources.length - 1]));
|
||||
if (relative.startsWith("..") || path.isAbsolute(relative)) {
|
||||
fail(4, `source session is outside the configured dataRoot: ${sourceDir}`);
|
||||
}
|
||||
spawnEnv.MOSAIC_SESSION_FORK = `/var/lib/mosaic/${relative.split(path.sep).join("/")}`;
|
||||
}
|
||||
|
||||
const proc = spawnSync(
|
||||
"docker",
|
||||
["compose", "run", "--rm", "-T", "mosaic-agent", task.prompt],
|
||||
@@ -335,9 +422,11 @@ function runTask(taskFile) {
|
||||
request: task.prompt,
|
||||
response,
|
||||
expectedExact: expected,
|
||||
...(options.retriedFrom ? { retriedFrom: options.retriedFrom } : {}),
|
||||
workspace: task.workspace,
|
||||
tools: task.tools,
|
||||
tools: effectiveTools,
|
||||
session: task.session,
|
||||
sessionForkFrom: task.sessionForkFrom,
|
||||
exitCode: proc.status,
|
||||
signal: proc.signal ?? null,
|
||||
provider: resolved.execution.provider,
|
||||
@@ -365,17 +454,166 @@ function listRuns() {
|
||||
for (const runId of entries) {
|
||||
let status = "unknown";
|
||||
let taskId = "-";
|
||||
let workspace = "-";
|
||||
let session = "-";
|
||||
try {
|
||||
const result = JSON.parse(fs.readFileSync(path.join(root, runId, "result.json"), "utf8"));
|
||||
status = result.status;
|
||||
taskId = result.taskId;
|
||||
workspace = result.workspace ?? "-";
|
||||
session = result.session ?? "-";
|
||||
} catch {
|
||||
// Incomplete run record; report as unknown.
|
||||
}
|
||||
process.stdout.write(`${runId} ${status.padEnd(9)} ${taskId}\n`);
|
||||
process.stdout.write(`${runId} ${status.padEnd(9)} task=${taskId.padEnd(18)} ws=${String(workspace).padEnd(10)} session=${session}\n`);
|
||||
}
|
||||
}
|
||||
|
||||
function showRun(runId) {
|
||||
if (!/^r-[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(runId)) {
|
||||
fail(4, `invalid run id: ${JSON.stringify(runId)} (expected r-<id>)`);
|
||||
}
|
||||
const resolved = loadConfig();
|
||||
const dir = path.join(runsRoot(resolved), runId);
|
||||
if (!fs.existsSync(dir)) {
|
||||
fail(4, `run not found: ${runId} (under ${runsRoot(resolved)})`);
|
||||
}
|
||||
|
||||
const read = (name) => {
|
||||
try {
|
||||
return JSON.parse(fs.readFileSync(path.join(dir, name), "utf8"));
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
};
|
||||
const result = read("result.json");
|
||||
const task = read("task.json");
|
||||
const mission = read("mission.json");
|
||||
|
||||
process.stdout.write(`run: ${runId}\n`);
|
||||
if (result) {
|
||||
process.stdout.write(
|
||||
[
|
||||
`status: ${result.status}${result.reason ? ` (${result.reason})` : ""}`,
|
||||
`task: ${result.taskId}`,
|
||||
result.missionId ? `mission: ${result.missionId}` : null,
|
||||
result.workspace ? `workspace: ${result.workspace}` : null,
|
||||
result.session ? `session: ${result.session}` : null,
|
||||
result.tools ? `tools: ${result.tools.join(", ")}` : null,
|
||||
`adapter: ${"(see config)"} provider=${result.provider} model=${result.model}`,
|
||||
`request: ${JSON.stringify(result.request)}`,
|
||||
`response: ${JSON.stringify(result.response)}`,
|
||||
result.expectedExact !== null && result.expectedExact !== undefined ? `expected: ${JSON.stringify(result.expectedExact)}` : null,
|
||||
result.retriedFrom ? `retriedFrom: ${result.retriedFrom}` : null,
|
||||
`timing: ${result.startedAt} -> ${result.finishedAt} (${result.durationMs} ms)`,
|
||||
`exit: ${result.exitCode}${result.signal ? ` signal=${result.signal}` : ""}`,
|
||||
].filter((line) => line !== null).join("\n") + "\n",
|
||||
);
|
||||
} else {
|
||||
process.stdout.write("result.json: (missing or unreadable)\n");
|
||||
}
|
||||
if (task) process.stdout.write(`task snapshot: ${"task.json"} present\n`);
|
||||
if (mission) process.stdout.write(`mission snapshot: ${mission.id} - ${mission.objective}\n`);
|
||||
process.stdout.write(`artifacts: ${fs.readdirSync(dir).map((f) => `${f}`).join(", ")}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
function retryRun(runId) {
|
||||
if (!/^r-[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(runId)) {
|
||||
fail(4, `invalid run id: ${JSON.stringify(runId)} (expected r-<id>)`);
|
||||
}
|
||||
const resolved = loadConfig();
|
||||
const dir = path.join(runsRoot(resolved), runId);
|
||||
if (!fs.existsSync(dir)) {
|
||||
fail(4, `run not found: ${runId}`);
|
||||
}
|
||||
|
||||
let snapshot;
|
||||
try {
|
||||
snapshot = fs.readFileSync(path.join(dir, "task.json"), "utf8");
|
||||
} catch {
|
||||
fail(4, `run task snapshot unreadable: ${runId}`);
|
||||
}
|
||||
|
||||
// Relative mission paths in a snapshot resolve against the ORIGINAL task
|
||||
// location, which no longer exists here — rewrite them to the run's own
|
||||
// recorded mission.json so retries stay faithful.
|
||||
let snapshotDoc;
|
||||
try {
|
||||
snapshotDoc = JSON.parse(snapshot);
|
||||
} catch {
|
||||
fail(4, `run task snapshot is not valid JSON: ${runId}`);
|
||||
}
|
||||
if (snapshotDoc.mission && !path.isAbsolute(snapshotDoc.mission)) {
|
||||
const recordedMission = path.join(dir, "mission.json");
|
||||
if (!fs.existsSync(recordedMission)) {
|
||||
fail(4, `cannot retry ${runId}: relative mission path but no mission.json snapshot in run dir`);
|
||||
}
|
||||
snapshotDoc.mission = recordedMission;
|
||||
snapshot = `${JSON.stringify(snapshotDoc, null, 2)}\n`;
|
||||
}
|
||||
|
||||
// A retry is a brand-new run: replay the recorded task snapshot through
|
||||
// the ordinary run path; existing run records stay untouched.
|
||||
const tempDir = fs.mkdtempSync(path.join(os.tmpdir(), "mosaic-retry-"));
|
||||
const tempTaskFile = path.join(tempDir, "task.json");
|
||||
fs.writeFileSync(tempTaskFile, snapshot);
|
||||
process.on("exit", () => {
|
||||
try {
|
||||
fs.rmSync(tempDir, { recursive: true, force: true });
|
||||
} catch {
|
||||
// Best-effort cleanup only.
|
||||
}
|
||||
});
|
||||
runTask(tempTaskFile, { retriedFrom: runId });
|
||||
}
|
||||
|
||||
function pruneRuns(args) {
|
||||
const resolved = loadConfig();
|
||||
const root = runsRoot(resolved);
|
||||
let keep = 50;
|
||||
let apply = false;
|
||||
for (const arg of args) {
|
||||
if (arg === "--yes") apply = true;
|
||||
else if (arg === "--keep") fail(4, "prune: --keep requires a value");
|
||||
else if (arg.startsWith("--keep=")) {
|
||||
keep = Number(arg.slice("--keep=".length));
|
||||
if (!Number.isInteger(keep) || keep < 1) fail(4, `--keep must be a positive integer (got ${arg.slice(7)})`);
|
||||
} else fail(4, `unknown prune option: ${arg}`);
|
||||
}
|
||||
|
||||
let entries = [];
|
||||
try {
|
||||
entries = fs.readdirSync(root, { withFileTypes: true })
|
||||
.filter((e) => e.name.startsWith("r-") && e.isDirectory() && !e.isSymbolicLink())
|
||||
.map((e) => e.name)
|
||||
.sort();
|
||||
} catch {
|
||||
// No runs yet.
|
||||
}
|
||||
|
||||
if (entries.length <= keep) {
|
||||
process.stdout.write(`prune: ${entries.length} run(s) present, keep=${keep} -> nothing to prune\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const doomed = entries.slice(0, entries.length - keep); // oldest first
|
||||
if (!apply) {
|
||||
process.stdout.write(`prune (dry-run): would remove ${doomed.length} oldest run(s), keep ${entries.length - doomed.length}:\n`);
|
||||
for (const id of doomed) process.stdout.write(` would remove: ${id}\n`);
|
||||
process.stdout.write("prune: re-run with --yes to apply\n");
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const receipt = path.join(root, ".pruned.log");
|
||||
for (const id of doomed) {
|
||||
fs.rmSync(path.join(root, id), { recursive: true, force: true });
|
||||
fs.appendFileSync(receipt, `${JSON.stringify({ at: new Date().toISOString(), event: "pruned", runId: id })}\n`);
|
||||
}
|
||||
process.stdout.write(`prune: removed ${doomed.length} run(s), kept ${keep}; receipt: ${receipt}\n`);
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
const operation = process.argv[2];
|
||||
const target = process.argv[3];
|
||||
|
||||
@@ -392,9 +630,20 @@ switch (operation) {
|
||||
if (!target) fail(4, "usage: mosaic-task.mjs run <taskFile>");
|
||||
runTask(path.resolve(target));
|
||||
break;
|
||||
case "show":
|
||||
if (!target) fail(4, "usage: mosaic-task.mjs show <runId>");
|
||||
showRun(target);
|
||||
break;
|
||||
case "list":
|
||||
listRuns();
|
||||
process.exit(0);
|
||||
case "prune":
|
||||
pruneRuns(process.argv.slice(3));
|
||||
break;
|
||||
case "retry":
|
||||
if (!target) fail(4, "usage: mosaic-task.mjs retry <runId>");
|
||||
retryRun(target);
|
||||
break;
|
||||
default:
|
||||
fail(4, `unknown operation: ${JSON.stringify(operation ?? "")} (expected validate | run | list)`);
|
||||
fail(4, `unknown operation: ${JSON.stringify(operation ?? "")} (expected validate | run | show | list | retry | prune)`);
|
||||
}
|
||||
|
||||
Executable
+140
@@ -0,0 +1,140 @@
|
||||
#!/usr/bin/env bash
|
||||
# Sandboxed selftests for the conductor auto-apply policy gate.
|
||||
#
|
||||
# Builds a throwaway target repo + worker workspace + fake run records, then
|
||||
# exercises every gate: policy validation, allowlist, syntax gates, suite
|
||||
# failure revert, disabled policy, missing/failed runs. No real model calls.
|
||||
set -uo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
|
||||
SANDBOX="$(mktemp -d)"
|
||||
trap 'rm -rf "$SANDBOX"' EXIT
|
||||
|
||||
PASS=0
|
||||
FAIL=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
|
||||
check_rc() { # name expectedRc command...
|
||||
local name="$1" expected="$2"
|
||||
shift 2
|
||||
local rc
|
||||
"$@" >/dev/null 2>&1
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
# ---- infrastructure: target repo + worker workspace + fake run ----
|
||||
git clone -q . "$SANDBOX/repo"
|
||||
# The clone carries committed state only - give the target its policy and
|
||||
# commit it so the tree starts clean (untracked policy would fail target_clean).
|
||||
cp conductor-policy.json "$SANDBOX/repo/conductor-policy.json"
|
||||
git -C "$SANDBOX/repo" add conductor-policy.json
|
||||
git -C "$SANDBOX/repo" -c user.name=suite -c user.email=suite@local commit -q -m policy
|
||||
mkdir -p "$SANDBOX/data/workspaces" "$SANDBOX/data/runs"
|
||||
git clone -q "$SANDBOX/repo" "$SANDBOX/data/workspaces/stack-repo"
|
||||
|
||||
TARGET="$SANDBOX/repo"
|
||||
WS="$SANDBOX/data/workspaces/stack-repo"
|
||||
cat > "$SANDBOX/config.json" <<EOF
|
||||
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"mock"}}
|
||||
EOF
|
||||
export MOSAIC_APPLY_TARGET="$TARGET"
|
||||
export MOSAIC_CONFIG="$SANDBOX/config.json"
|
||||
|
||||
RUN_OK="r-20260903T000000000Z-ok0000001"
|
||||
mkdir -p "$SANDBOX/data/runs/$RUN_OK"
|
||||
printf '{"runVersion":1,"runId":"%s","taskId":"t-fake","status":"succeeded","workspace":"stack-repo"}' "$RUN_OK" \
|
||||
> "$SANDBOX/data/runs/$RUN_OK/result.json"
|
||||
|
||||
ws_edit() { printf '\n%s\n' "$2" >> "$WS/$1"; }
|
||||
ws_reset() { git -C "$WS" checkout -q -- . 2>/dev/null; git -C "$WS" clean -qfd; }
|
||||
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
|
||||
|
||||
set_policy() { # enabled suites (commits: the target tree must stay clean)
|
||||
local suites="[\"$2\"]"
|
||||
printf '{"policyVersion":1,"autoApply":{"enabled":%s,"allowedPaths":["scripts/**","docs/**","README.md"],"suites":%s}}' "$1" "$suites" \
|
||||
> "$TARGET/conductor-policy.json"
|
||||
git -C "$TARGET" add conductor-policy.json
|
||||
git -C "$TARGET" -c user.name=suite -c user.email=suite@local commit -q -m "policy update"
|
||||
}
|
||||
set_policy true "test-config"
|
||||
|
||||
# T1: dry run - allowed change, nothing applied
|
||||
ws_edit "README.md" "worker dry-run line"
|
||||
check_rc "dry-run: allowed change, exit 0, nothing committed" 0 \
|
||||
scripts/conductor-apply.sh "$RUN_OK" --dry-run
|
||||
if git -C "$TARGET" log --format=%s | grep -q "auto-applied"; then
|
||||
check "dry-run committed nothing" 1
|
||||
else
|
||||
check "dry-run committed nothing" 0
|
||||
fi
|
||||
ws_reset
|
||||
|
||||
# T2: apply - allowed change, suites pass, commit created
|
||||
ws_edit "README.md" "worker applied line"
|
||||
check_rc "apply: allowed change exits 0" 0 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" log -1 --format=%s | grep -q "auto-applied patch from run $RUN_OK" \
|
||||
&& check "apply: attribution in commit subject" 0 || check "apply: attribution in commit subject" 1
|
||||
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
|
||||
target_clean && check "apply: target tree clean after commit" 0 || check "apply: target tree clean after commit" 1
|
||||
git -C "$TARGET" reset -q --hard HEAD~1
|
||||
|
||||
# T3: disallowed path refused
|
||||
ws_edit "Containerfile" "# worker touch"
|
||||
check_rc "disallowed path refused" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "disallowed path: target untouched" 0 || check "disallowed path: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T4: syntax gate - broken .mjs on an allowed path
|
||||
printf 'this is not (valid js\n' > "$WS/scripts/broken-worker.mjs"
|
||||
check_rc "syntax gate refused broken .mjs" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "syntax gate: target untouched" 0 || check "syntax gate: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T5: suite failure - allowed change breaks a policy suite -> auto-revert
|
||||
printf '\nexit 7\n' >> "$TARGET/scripts/test-config.sh"
|
||||
ws_edit "README.md" "worker change that will fail suites"
|
||||
check_rc "suite failure refused" 1 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" checkout -q -- scripts/test-config.sh
|
||||
target_clean && check "suite failure: target reverted to clean" 0 || check "suite failure: target reverted to clean" 1
|
||||
|
||||
# T6: disabled policy
|
||||
set_policy false "test-config"
|
||||
ws_edit "README.md" "worker line while disabled"
|
||||
check_rc "disabled policy refused" 2 scripts/conductor-apply.sh "$RUN_OK"
|
||||
target_clean && check "disabled policy: target untouched" 0 || check "disabled policy: target untouched" 1
|
||||
ws_reset
|
||||
set_policy true "test-config"
|
||||
|
||||
# T7: failed run refused
|
||||
RUN_FAIL="r-20260903T000000000Z-fail00001"
|
||||
mkdir -p "$SANDBOX/data/runs/$RUN_FAIL"
|
||||
printf '{"runVersion":1,"runId":"%s","taskId":"t","status":"failed","workspace":"stack-repo"}' "$RUN_FAIL" \
|
||||
> "$SANDBOX/data/runs/$RUN_FAIL/result.json"
|
||||
ws_edit "README.md" "worker line from failed run"
|
||||
check_rc "failed run refused" 1 scripts/conductor-apply.sh "$RUN_FAIL"
|
||||
target_clean && check "failed run: target untouched" 0 || check "failed run: target untouched" 1
|
||||
ws_reset
|
||||
|
||||
# T8/T9: missing run + invalid policy
|
||||
check_rc "missing run exits 4" 4 scripts/conductor-apply.sh r-missing
|
||||
printf '{"policyVersion":9}' > "$TARGET/conductor-policy.json"
|
||||
check_rc "invalid policy exits 2" 2 scripts/conductor-apply.sh "$RUN_OK"
|
||||
git -C "$TARGET" checkout -q -- conductor-policy.json
|
||||
|
||||
echo
|
||||
echo "selftest: $PASS passed, $FAIL failed"
|
||||
[ "$FAIL" -eq 0 ]
|
||||
+17
-11
@@ -10,10 +10,16 @@ SANDBOX="$(mktemp -d)"
|
||||
trap 'rm -rf "$SANDBOX"' EXIT
|
||||
|
||||
PASS=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
FAIL=0
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
# expect_exit NAME EXPECTED_RC -- command...
|
||||
@@ -25,10 +31,10 @@ expect_exit() {
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS + 1))
|
||||
echo "ok $name (exit $rc)"
|
||||
echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL + 1))
|
||||
echo "FAIL $name (exit $rc, expected $expected)"
|
||||
echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
@@ -62,8 +68,8 @@ check "env exports adapter" $?
|
||||
rm -f "$SANDBOX/config.json"
|
||||
expect_exit "bootstrap creates default when absent" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/config.json" $CONFIG_OP bootstrap
|
||||
[ -f "$SANDBOX/config.json" ] && { PASS=$((PASS+1)); echo "ok bootstrap wrote config file"; } \
|
||||
|| { FAIL=$((FAIL+1)); echo "FAIL bootstrap wrote config file"; }
|
||||
[ -f "$SANDBOX/config.json" ] && { PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} bootstrap wrote config file"; } \
|
||||
|| { FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} bootstrap wrote config file"; }
|
||||
|
||||
SUM_BEFORE=$(sha256sum "$SANDBOX/config.json" | cut -d' ' -f1)
|
||||
MTIME_BEFORE=$(stat -c %Y "$SANDBOX/config.json")
|
||||
@@ -73,9 +79,9 @@ expect_exit "bootstrap is idempotent on existing config" 0 -- \
|
||||
SUM_AFTER=$(sha256sum "$SANDBOX/config.json" | cut -d' ' -f1)
|
||||
MTIME_AFTER=$(stat -c %Y "$SANDBOX/config.json")
|
||||
if [ "$SUM_BEFORE" = "$SUM_AFTER" ] && [ "$MTIME_BEFORE" = "$MTIME_AFTER" ]; then
|
||||
PASS=$((PASS+1)); echo "ok bootstrap did not rewrite existing config"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} bootstrap did not rewrite existing config"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL bootstrap rewrote existing config"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} bootstrap rewrote existing config"
|
||||
fi
|
||||
|
||||
# --- validate ---
|
||||
@@ -140,9 +146,9 @@ cfg valid.json "$(valid_body "$DATA_ROOT")"
|
||||
EVAL_OUT="$(MOSAIC_CONFIG="$SANDBOX/valid.json" $CONFIG_OP env)" || true
|
||||
if eval "$EVAL_OUT" 2>/dev/null && [ "$MOSAIC_DATA_ROOT" = "$DATA_ROOT" ] \
|
||||
&& [ "$MOSAIC_PROVIDER" = "zai" ] && [ "$MOSAIC_MODEL" = "glm-5.3-flash" ]; then
|
||||
PASS=$((PASS+1)); echo "ok env exports resolve correctly"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} env exports resolve correctly"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL env exports resolve correctly"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} env exports resolve correctly"
|
||||
fi
|
||||
|
||||
# --- validation must not modify the file ---
|
||||
@@ -150,9 +156,9 @@ SUM_INVALID_BEFORE=$(sha256sum "$SANDBOX/invalid.json" | cut -d' ' -f1)
|
||||
MOSAIC_CONFIG="$SANDBOX/invalid.json" $CONFIG_OP validate >/dev/null 2>&1
|
||||
SUM_INVALID_AFTER=$(sha256sum "$SANDBOX/invalid.json" | cut -d' ' -f1)
|
||||
if [ "$SUM_INVALID_BEFORE" = "$SUM_INVALID_AFTER" ]; then
|
||||
PASS=$((PASS+1)); echo "ok failed validation modified nothing"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} failed validation modified nothing"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL failed validation modified the file"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} failed validation modified the file"
|
||||
fi
|
||||
|
||||
echo
|
||||
|
||||
@@ -15,6 +15,12 @@ cp RELEASE "$RELEASE_BACKUP"
|
||||
trap 'cp "$RELEASE_BACKUP" RELEASE 2>/dev/null; rm -rf "$SANDBOX" "$RELEASE_BACKUP"' EXIT
|
||||
|
||||
PASS=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
FAIL=0
|
||||
|
||||
expect_exit() {
|
||||
@@ -24,14 +30,14 @@ expect_exit() {
|
||||
"$@" >/dev/null 2>&1
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS+1)); echo "ok $name (exit $rc)"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL $name (exit $rc, expected $expected)"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
# ---------- fast: release identity ----------
|
||||
|
||||
+120
-17
@@ -11,6 +11,12 @@ SANDBOX="$(mktemp -d)"
|
||||
trap 'rm -rf "$SANDBOX"' EXIT
|
||||
|
||||
PASS=0
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
FAIL=0
|
||||
|
||||
expect_exit() {
|
||||
@@ -20,14 +26,14 @@ expect_exit() {
|
||||
"$@" >/dev/null 2>&1
|
||||
rc=$?
|
||||
if [ "$rc" -eq "$expected" ]; then
|
||||
PASS=$((PASS+1)); echo "ok $name (exit $rc)"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL $name (exit $rc, expected $expected)"
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
|
||||
fi
|
||||
}
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
|
||||
}
|
||||
|
||||
latest_reason() {
|
||||
@@ -36,10 +42,6 @@ latest_reason() {
|
||||
[ -n "$latest" ] && node -e 'try{const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.reason??"")}catch{console.log("")}' "$latest/result.json" 2>/dev/null
|
||||
}
|
||||
|
||||
check() {
|
||||
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
|
||||
}
|
||||
|
||||
CONFIG="$SANDBOX/config.json"
|
||||
DATA_ROOT="$SANDBOX/data"
|
||||
mkdir -p "$DATA_ROOT"
|
||||
@@ -102,6 +104,37 @@ $TASK validate "$SANDBOX/ok.json" >/dev/null 2>&1
|
||||
M2=$(stat -c %Y "$SANDBOX/ok.json")
|
||||
[ "$M1" = "$M2" ] && check "validation does not modify the task file" 0 || check "validation does not modify the task file" 1
|
||||
|
||||
# ---------- retention: prune (deterministic, no Docker) ----------
|
||||
mkdir -p "$SANDBOX/data"
|
||||
cat > "$SANDBOX/prune-config.json" <<EOF
|
||||
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m"}}
|
||||
EOF
|
||||
for i in 1 2 3 4 5; do
|
||||
D="$SANDBOX/data/runs/r-20260903T0100_0${i}Z-suite00$i"
|
||||
mkdir -p "$D"
|
||||
printf '{"runVersion":1,"runId":"r-20260903T0100_0%sZ-suite00%s","taskId":"t","status":"succeeded"}' "$i" "$i" > "$D/result.json"
|
||||
done
|
||||
mkdir -p "$SANDBOX/data/sessions/sentinel" "$SANDBOX/data/workspaces/sentinel"
|
||||
|
||||
expect_exit "prune dry-run exits 0" 0 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune
|
||||
[ "$(ls "$SANDBOX/data/runs" | grep -c '^r-')" -eq 5 ] \
|
||||
&& check "dry-run deleted nothing" 0 || check "dry-run deleted nothing" 1
|
||||
|
||||
expect_exit "prune --keep=2 --yes removes oldest" 0 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune --keep=2 --yes
|
||||
[ "$(ls "$SANDBOX/data/runs" | grep -c '^r-')" -eq 2 ] \
|
||||
&& check "kept exactly 2 newest runs" 0 || check "kept exactly 2 newest runs" 1
|
||||
NEWEST="r-20260903T0100_05Z-suite005"
|
||||
[ -d "$SANDBOX/data/runs/$NEWEST" ] \
|
||||
&& check "newest run kept, oldest pruned" 0 || check "newest run kept, oldest pruned" 1
|
||||
[ -f "$SANDBOX/data/runs/.pruned.log" ] \
|
||||
&& [ "$(grep -c 'pruned' "$SANDBOX/data/runs/.pruned.log")" -eq 3 ] \
|
||||
&& check "append-only receipt written (3 entries)" 0 \
|
||||
|| check "append-only receipt written (3 entries)" 1
|
||||
[ -d "$SANDBOX/data/sessions/sentinel" ] && [ -d "$SANDBOX/data/workspaces/sentinel" ] \
|
||||
&& check "sessions/workspaces untouched by prune" 0 \
|
||||
|| check "sessions/workspaces untouched by prune" 1
|
||||
expect_exit "prune with invalid keep exits 4" 4 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune --keep=0 --yes
|
||||
|
||||
# ---------- adapter seam: deterministic mock cases (Docker, no provider) ----------
|
||||
if docker info >/dev/null 2>&1; then
|
||||
good_task "$SANDBOX/ok.json"
|
||||
@@ -142,6 +175,67 @@ EOF
|
||||
&& grep -q 'Seam directive A.' "$SANDBOX/data/system-prompt.md" \
|
||||
&& check "mission section injected into generated prompt" 0 \
|
||||
|| check "mission section injected into generated prompt" 1
|
||||
|
||||
# retry lineage + relative mission path resolution
|
||||
expect_exit "retry of mission run succeeds" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
|
||||
node scripts/mosaic-task.mjs retry "$(ls -dt "$SANDBOX/data/runs"/r-* | head -1 | xargs basename)"
|
||||
RT="$(ls -dt "$SANDBOX/data/runs"/r-* | head -1 | xargs basename)"
|
||||
ORIG="$(ls -dt "$SANDBOX/data/runs"/r-* | sed -n 2p | xargs basename)"
|
||||
node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));process.exit(r.retriedFrom===process.argv[2]?0:1)' \
|
||||
"$SANDBOX/data/runs/$RT/result.json" "$ORIG" \
|
||||
&& check "retriedFrom lineage recorded" 0 || check "retriedFrom lineage recorded" 1
|
||||
grep -q 'MISSION (runtime)' "$SANDBOX/data/system-prompt.md" \
|
||||
&& check "mission section present after retry (relative path resolved)" 0 \
|
||||
|| check "mission section present after retry (relative path resolved)" 1
|
||||
expect_exit "retry of missing run exits 4" 4 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" node scripts/mosaic-task.mjs retry r-missing
|
||||
|
||||
# session fork plumbing (M11): fork source + target dir delivered
|
||||
printf '{"taskVersion":1,"id":"t-forkplumb","prompt":"x","session":"fork-child","sessionForkFrom":"base"}' > "$SANDBOX/forkplumb.json"
|
||||
mkdir -p "$SANDBOX/data/sessions/base"
|
||||
printf '{}' > "$SANDBOX/data/sessions/base/20260903T000000-plumb.jsonl"
|
||||
expect_exit "fork task runs via mock" 0 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
|
||||
scripts/run-task.sh run "$SANDBOX/forkplumb.json"
|
||||
FL="$(ls -dt "$SANDBOX/data/runs"/r-* | head -1)"
|
||||
grep -q '^MOSAIC_SESSION_FORK=/var/lib/mosaic/sessions/base/20260903T000000-plumb.jsonl$' "$FL/stderr.txt" 2>/dev/null \
|
||||
&& grep -q '^MOSAIC_SESSION_DIR=/var/lib/mosaic/sessions/fork-child$' "$FL/stderr.txt" 2>/dev/null \
|
||||
&& check "fork source + target delivered to adapter" 0 \
|
||||
|| check "fork source + target delivered to adapter" 1
|
||||
printf '{"taskVersion":1,"id":"t-f2","prompt":"x","sessionForkFrom":"base"}' > "$SANDBOX/noforktarget.json"
|
||||
expect_exit "fork without session target exits 2" 2 -- $TASK validate "$SANDBOX/noforktarget.json"
|
||||
printf '{"taskVersion":1,"id":"t-f3","prompt":"x","session":"base","sessionForkFrom":"base"}' > "$SANDBOX/selfork.json"
|
||||
expect_exit "self-fork exits 2" 2 -- $TASK validate "$SANDBOX/selfork.json"
|
||||
|
||||
# capability policy (M9): least-privilege intersection
|
||||
POL="$SANDBOX/data/workspaces"; mkdir -p "$POL"
|
||||
pol_run() { # missionTools(ABSENT|json) taskTools(ABSENT|json) -> stderr MOSAIC_TOOLS value
|
||||
if [ "$1" = "ABSENT" ]; then
|
||||
printf '{"missionVersion":1,"id":"m-pol","objective":"o"}' > "$SANDBOX/pol-m.json"
|
||||
else
|
||||
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":[%s]}}' "$1" > "$SANDBOX/pol-m.json"
|
||||
fi
|
||||
if [ "$2" = "ABSENT" ]; then
|
||||
printf '{"taskVersion":1,"id":"t-pol","prompt":"x","mission":"pol-m.json"}' > "$SANDBOX/pol-t.json"
|
||||
else
|
||||
printf '{"taskVersion":1,"id":"t-pol","prompt":"x","mission":"pol-m.json","capabilities":{"tools":[%s]}}' "$2" > "$SANDBOX/pol-t.json"
|
||||
fi
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
|
||||
scripts/run-task.sh run "$SANDBOX/pol-t.json" >/dev/null 2>&1
|
||||
grep -o '^MOSAIC_TOOLS=.*' "$(ls -dt "$SANDBOX/data/runs"/r-* | head -1)/stderr.txt" 2>/dev/null
|
||||
}
|
||||
[ "$(pol_run '"read","bash"' ABSENT)" = 'MOSAIC_TOOLS=read,bash' ] \
|
||||
&& check "policy: mission only -> mission tools" 0 || check "policy: mission only -> mission tools" 1
|
||||
[ "$(pol_run ABSENT '"read","bash"')" = 'MOSAIC_TOOLS=read,bash' ] \
|
||||
&& check "policy: task only -> task tools" 0 || check "policy: task only -> task tools" 1
|
||||
[ "$(pol_run '"read"' '"bash","read"')" = 'MOSAIC_TOOLS=read' ] \
|
||||
&& check "policy: both -> intersection (task narrowed)" 0 || check "policy: both -> intersection (task narrowed)" 1
|
||||
[ "$(pol_run '"grep"' '"bash","read"')" = 'MOSAIC_TOOLS=' ] \
|
||||
&& check "policy: empty intersection -> tool-free" 0 || check "policy: empty intersection -> tool-free" 1
|
||||
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":["sudo"]}}' > "$SANDBOX/pol-m.json"
|
||||
expect_exit "invalid mission capabilities rejected" 2 -- \
|
||||
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" $TASK validate "$SANDBOX/pol-t.json"
|
||||
else
|
||||
echo "skip adapter seam cases (docker daemon unavailable)"
|
||||
fi
|
||||
@@ -194,18 +288,12 @@ dump_latest_run() {
|
||||
fi
|
||||
}
|
||||
|
||||
latest_reason() {
|
||||
local latest
|
||||
latest="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
|
||||
[ -n "$latest" ] && node -e 'try{const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.reason??"")}catch{console.log("")}' "$latest/result.json" 2>/dev/null
|
||||
}
|
||||
|
||||
if docker info >/dev/null 2>&1; then
|
||||
RUNS1=$(ls "$DATA_ROOT/runs" 2>/dev/null | wc -l)
|
||||
if scripts/run-task.sh run "$SANDBOX/ok.json" >/dev/null 2>&1; then
|
||||
PASS=$((PASS+1)); echo "ok live hello task succeeds with exact marker"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} live hello task succeeds with exact marker"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL live hello task succeeds with exact marker" >&2
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} live hello task succeeds with exact marker" >&2
|
||||
dump_latest_run
|
||||
fi
|
||||
|
||||
@@ -222,9 +310,9 @@ process.exit(r.status === "succeeded" && r.response === "MOSAIC_HELLO_OK" && r.e
|
||||
RC=$?
|
||||
WRONG_REASON="$(latest_reason)"
|
||||
if [ "$RC" -eq 1 ] && [ "$WRONG_REASON" = "expect-mismatch" ]; then
|
||||
PASS=$((PASS+1)); echo "ok wrong expectExact fails with exit 1 (reason: expect-mismatch)"
|
||||
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} wrong expectExact fails with exit 1 (reason: expect-mismatch)"
|
||||
else
|
||||
FAIL=$((FAIL+1)); echo "FAIL wrong expectExact (exit $RC, reason: '${WRONG_REASON:-none}')" >&2
|
||||
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} wrong expectExact (exit $RC, reason: '${WRONG_REASON:-none}')" >&2
|
||||
dump_latest_run
|
||||
fi
|
||||
|
||||
@@ -234,6 +322,21 @@ process.exit(r.status === "succeeded" && r.response === "MOSAIC_HELLO_OK" && r.e
|
||||
|
||||
COUNT=$($TASK list | wc -l)
|
||||
[ "$COUNT" -ge 2 ] && check "list shows both runs" 0 || check "list shows both runs" 1
|
||||
|
||||
# session fork (M11): teach in base, fork into child, child recalls;
|
||||
# ancestor file count must be unchanged by the fork
|
||||
printf '{"taskVersion":1,"id":"t-fork-base","prompt":"Remember this code word for later: mosaico. Reply with exactly: REMEMBERED","session":"suite-base","expectExact":"REMEMBERED","timeoutSeconds":180}' > "$SANDBOX/fb.json"
|
||||
expect_exit "fork base: teach succeeds" 0 -- scripts/run-task.sh run "$SANDBOX/fb.json"
|
||||
BASECOUNT=$(ls "$SANDBOX/data/sessions/suite-base" | wc -l)
|
||||
printf '{"taskVersion":1,"id":"t-fork-child","prompt":"What code word did I ask you to remember? Reply with only the code word.","session":"suite-child","sessionForkFrom":"suite-base","timeoutSeconds":180}' > "$SANDBOX/fc.json"
|
||||
expect_exit "fork child recalls ancestor context" 0 -- scripts/run-task.sh run "$SANDBOX/fc.json"
|
||||
FR="$(ls -dt "$DATA_ROOT/runs"/r-* | head -1)"
|
||||
node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));process.exit(String(r.response).toLowerCase().includes("mosaico")?0:1)' "$FR/result.json" \
|
||||
&& check "forked child recalled ancestor code word" 0 || check "forked child recalled ancestor code word" 1
|
||||
[ "$(ls "$SANDBOX/data/sessions/suite-base" | wc -l)" -eq "$BASECOUNT" ] \
|
||||
&& check "ancestor session untouched by fork" 0 || check "ancestor session untouched by fork" 1
|
||||
[ "$(ls "$SANDBOX/data/sessions/suite-child" | wc -l)" -ge 1 ] \
|
||||
&& check "child session dir has its own branch file" 0 || check "child session dir has its own branch file" 1
|
||||
else
|
||||
echo "skip live task cases (docker unavailable)"
|
||||
fi
|
||||
|
||||
+9
-2
@@ -16,6 +16,13 @@ source scripts/common.sh
|
||||
|
||||
EXPECTED="${EXPECTED_MARKER:-MOSAIC_HELLO_OK}"
|
||||
|
||||
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
|
||||
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
|
||||
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
|
||||
else
|
||||
C_OK=""; C_FAIL=""; C_RESET=""
|
||||
fi
|
||||
|
||||
load_config
|
||||
load_release
|
||||
IMAGE="$MOSAIC_IMAGE_TAG"
|
||||
@@ -51,11 +58,11 @@ TRIMMED="$(printf '%s' "$RESPONSE" | sed -e 's/^[[:space:]]*//' -e 's/[[:space:]
|
||||
|
||||
# 4-6. Exact comparison gate.
|
||||
if [ "$TRIMMED" = "$EXPECTED" ]; then
|
||||
echo "PASS: response matches expected marker"
|
||||
echo "${C_OK}PASS${C_RESET}: response matches expected marker"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo "FAIL: response does not match expected marker" >&2
|
||||
echo "${C_FAIL}FAIL${C_RESET}: response does not match expected marker" >&2
|
||||
printf 'expected: %s\n' "$EXPECTED" >&2
|
||||
printf 'actual : %s\n' "$TRIMMED" >&2
|
||||
exit 1
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-worker-retry-refine",
|
||||
"prompt": "Follow-up on your last change to scripts/mosaic-task.mjs (the retry operation). Conductor review found one defect: runTask builds spawnEnv from process.env, so a direct 'node scripts/mosaic-task.mjs retry <runId>' invocation lacks the launcher-exported variables and compose fails with: required variable MOSAIC_IMAGE_TAG is missing.\n\nFIX (in runTask, not retryRun): make spawnEnv self-sufficient —\n1. spawnEnv.PI_PROVIDER = resolved.execution.provider; spawnEnv.PI_MODEL = resolved.execution.model (from the already-loaded resolved config).\n2. spawnEnv.MOSAIC_IMAGE_TAG = 'mosaic-poc-agent:' + <pinned pi version> + '-r' + <RELEASE file content trimmed>, reading RELEASE and package.json the same way scripts/common.sh load_release does.\nKeep the change minimal; do not touch other files; keep code style.\n\nSELF-CHECK: node --check scripts/mosaic-task.mjs must pass.\nWhen finished, reply with exactly: REFINE_DONE",
|
||||
"workspace": "stack-repo",
|
||||
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
|
||||
"session": "worker-1",
|
||||
"timeoutSeconds": 600
|
||||
}
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"taskVersion": 1,
|
||||
"id": "t-worker-retry",
|
||||
"prompt": "You are working alone inside a git repository checkout (your current working directory). Implement a new feature in scripts/mosaic-task.mjs.\n\nFEATURE: add a 'retry <runId>' operation alongside the existing 'validate | run | show | list' operations.\n\nBEHAVIOR: 'retry <runId>' re-executes a previously recorded run. Steps: (1) validate the runId argument with the same pattern showRun uses; (2) load the config exactly like showRun does; (3) locate the run directory under the runs root; if it does not exist, fail with exit code 4 and the message 'run not found: <runId>'; (4) read that run directory's task.json snapshot; if unreadable, fail with exit code 4; (5) write the snapshot to a temp file and execute it through the EXISTING runTask(path) function so a brand-new run is created — do not duplicate any execution logic; (6) exit with whatever exit code the new run produced.\n\nCONSTRAINTS: reuse runTask; do not refactor unrelated code; do not touch any file other than scripts/mosaic-task.mjs; keep the existing code style; update the default usage error message to include retry.\n\nSELF-CHECK (run it yourself with the bash tool): node --check scripts/mosaic-task.mjs must pass.\n\nWhen finished, reply with exactly: RETRY_DONE",
|
||||
"workspace": "stack-repo",
|
||||
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
|
||||
"session": "worker-1",
|
||||
"timeoutSeconds": 600
|
||||
}
|
||||
Reference in New Issue
Block a user