Compare commits

..
Author SHA1 Message Date
jason.woltje e175616885 feat(conductor): auto-apply policy gate for worker patches (#34)
- conductor-policy.json (tracked, strictly validated): enabled switch,
  path allowlist globs, gating suites - the autonomy decision lives in a
  declarative file the owner controls
- scripts/conductor-apply.sh <runId> [--dry-run]: succeeded-run check ->
  clean target tree -> diff from worker workspace -> allowlist -> syntax
  gates (node/bash/json) -> apply -> policy suites -> attribution commit;
  ANY failure reverts the tree; push is never automatic
- scripts/test-conductor.sh: 17 sandbox cases covering every gate incl.
  suite-failure auto-revert and disabled policy
- policy defaults: scripts/docs/tasks/missions/adapters + README; all
  three suites gate

Closes #34
2026-09-03 07:03:29 -05:00
jason.woltje 6955717612 Merge M11: session forking from a common ancestor
Closes #33
2026-09-03 06:43:02 -05:00
jason.woltje 88d9cf750f feat(sessions): sessionForkFrom - branch conversations from a common ancestor (#33)
- task schema: optional sessionForkFrom (source session name); requires
  session target; self-fork rejected
- runner: resolves source newest .jsonl (fail 4 if none/outside dataRoot);
  passes MOSAIC_SESSION_FORK + MOSAIC_SESSION_DIR; result records lineage
- pi adapter: --fork <source> --session-dir <target> when forking;
  ephemeral default unchanged; plain session resume unchanged
- compose passthrough; RELEASE -> 0.0.7 (adapter changed)
- suite +9 cases (58 total): plumbing via mock stderr, validation
  negatives, live fork - child recalls ancestor code word, ancestor
  session file untouched

Closes #33
2026-09-03 06:43:02 -05:00
jason.woltje c038706eed docs(plan): CURRENT.md - M10 shipped, retention next in review 2026-09-03 06:33:54 -05:00
jason.woltje 8622c9d826 Merge M10: run-record retention
Closes #32
2026-09-03 06:33:27 -05:00
jason.woltje 88eef507b0 feat(retention): run-record pruning - keep newest N, dry-run default (#32)
- mosaic-task.mjs prune [--keep=N] [--yes]: default keep 50; without
  --yes lists candidates without deleting
- only r-* directories under the runs root; symlinks skipped;
  sessions/workspaces/state/config untouched (asserted by suite sentinels)
- append-only receipt runs/.pruned.log records every pruned id
- test-task.sh: +8 retention cases (dry-run no-delete, keep-N, newest
  kept, receipt, isolation, invalid keep, empty no-op)

Also: suite hardening - prune section scopes its config per-command
(no export/unset leaking into later sections); duplicated check()
removed; latest_reason hoisted to helpers; status colors now green OK /
red FAIL (terminal-only, NO_COLOR-aware) per owner UX feedback.

Closes #32
2026-09-03 06:33:27 -05:00
jason.woltje 439bea6915 ui(test): green OK/PASS, red FAIL - terminal-only, NO_COLOR-aware
Owner feedback: grep match-highlighting made the word 'policy' red while
status words were plain - counter-indicative. Suites + verify now emit
ANSI colors (green success, red failure) when stdout is a terminal;
piped/machine-parsed output stays plain, honoring NO_COLOR. Word 'ok'
promoted to 'OK' for scannability.

Verified byte-level via forced-pty run; piped output unchanged; suites
41/24/14 + verify green.
2026-09-03 06:23:47 -05:00
jason.woltje fad8a4718c Merge M9: mission-level capability policy
Closes #30
2026-09-03 06:16:01 -05:00
jason.woltje 2ff49adff4 feat(policy): mission-level capability policy - least-privilege intersection (#30)
- mission schema: optional capabilities.tools (same validation as task)
- merge semantics in runTask: neither -> none; mission only -> mission;
  task only -> task; both -> intersection (task narrows, never widens);
  empty intersection -> tool-free run with an explicit stderr note
- result.json records EFFECTIVE tools; task/mission snapshots remain the
  immutable declaration of intent
- adapters unchanged; host-side only (no image change, 0.0.6 still active)
- task suite +5 cases (41 total): all four merge cases asserted from run
  evidence + invalid mission capabilities rejected

Policy decision recorded: missions govern; tasks cannot escalate.

Closes #30
2026-09-03 06:16:01 -05:00
jason.woltje 44c476ebbf fix(ops): show displays retriedFrom lineage (#29)
result.json recorded lineage correctly; the human-facing show command
omitted the field. Found by owner test: show | grep retriedFrom was
empty on a run whose result.json contained it.

Closes #29
2026-09-03 06:09:51 -05:00
jason.woltje cde480eb60 docs(plan): CURRENT.md — retry lineage shipped, M9 queued for decision 2026-09-03 05:31:24 -05:00
jason.woltje 5808248707 Merge retry lineage + relative mission resolution
Closes #28
2026-09-03 05:30:57 -05:00
jason.woltje afd5827db8 fix(retry): lineage tracking + relative mission path resolution (#28)
- retryRun rewrites a snapshot's relative mission path to the run's own
  recorded mission.json (absolute) before execution — retries stay
  faithful to what originally ran
- runTask accepts options.retriedFrom; retry records lineage in
  result.json (additive optional field, no schema break)
- task suite +4 cases: retry succeeds, lineage recorded, mission section
  present after retry (36 total), missing-run retry exits 4

Closes #28
2026-09-03 05:30:57 -05:00
jason.woltje d9cc990376 Merge M8: conductor loop - self-orchestration
Closes #25, closes #26, closes #27
2026-09-02 22:42:52 -05:00
jason.woltje 83c4e9851e feat(orchestration): retry <runId> — authored by headless pi worker (#26, #27)
Collaboration record (conductor loop, docs/plans/CONDUCTOR.md):
- round 1 (worker session worker-1, 2m28s): retry implemented per spec
- conductor live test exposed spec gap: direct invocation lacked
  launcher env exports
- round 2 (same worker session, 59s): spawnEnv made self-sufficient,
  but used PI_* where compose interpolates MOSAIC_*
- conductor hotfix: 3-line rename to MOSAIC_PROVIDER/MOSAIC_MODEL/
  MOSAIC_DATA_ROOT

Final: node scripts/mosaic-task.mjs retry <runId> re-executes a run's
task snapshot as a new run; live retry replied REMEMBERED; all suites
green (24/32/14 + verify).

Known limitation: retrying a run whose task used a RELATIVE mission path
resolves it against the temp dir; lineage tracking deferred.

Closes #25, closes #26, closes #27
2026-09-02 22:42:52 -05:00
jason.woltje 22508170a2 docs(plan): CURRENT.md — single next-action pointer for cadence-driven work 2026-09-02 22:25:56 -05:00
jason.woltje 90a67d050e Merge M7: operator ergonomics + release 0.0.6
Closes #24
2026-09-02 22:14:18 -05:00
jason.woltje 24bdef75fa feat(ops): run inspection, release 0.0.6, docs (#24)
- mosaic-task.mjs show <runId>: full record + snapshots + artifacts;
  uppercase-tolerant id validation; missing/traversal ids exit 4
- list: task/workspace/session columns
- RELEASE -> 0.0.6; README workspaces/capabilities/sessions sections;
  BUILD-LOG Phases 9-11; autonomous-run tracker results filled

Closes #24
2026-09-02 22:14:18 -05:00
18 changed files with 991 additions and 56 deletions
+116
View File
@@ -221,4 +221,120 @@ Release model and safe updates verified by drills. `main` merged with M3 and tag
Adapter seam verified; harness boundary is now additive by construction. `main` merged with M4 and tagged `adapter-seam-v1`.
---
## Phase 9: Workspaces + capability envelope (M5)
### Entry 9.1 — before
- Timestamp: 2026-09-03
- Intended action: Optional task workspace (absent / ":run" ephemeral / named persistent) and capabilities.tools allowlist (pi built-ins); runner plumbing via MOSAIC_WORKSPACE/MOSAIC_TOOLS; pi adapter maps to cwd + --tools; mock adapter logs delivered MOSAIC_* vars for deterministic assertions (Gitea #20, #21).
- Reason: Agents that only answer text cannot do work; the workspace+tools pair is the smallest real capability step, bounded by the container.
- Expected result: plumbing asserted via run-record stderr; live pi writes a host-visible file.
### Entry 9.2 — after
- Timestamp: 2026-09-03
- Commands run: build; mock plumbing run; validate negatives; full task suite; live workspace demo.
- Observed result: task suite 32/32; MOSAIC_WORKSPACE/MOSAIC_TOOLS asserted in run record; dataRoot/workspaces/<name> created host-side; live pi used bash to write proof.txt into the demo workspace (host-visible).
- Failure or correction: (1) batch edit dropped SUPPORTED_TOOLS const (runtime ReferenceError, exit 1 instead of 2) — restored; (2) mock env dump used `export` which this dash prints as `export K='v'` — switched to `env`; (3) selftest fed a mismatching mock response to the plain-task case — test bug, fixed.
## Phase 10: Named sessions (M6)
### Entry 10.1 — before
- Timestamp: 2026-09-03
- Intended action: Optional task.session name -> persistent session dir dataRoot/sessions/<name> via pi --session-dir; resume most recent with -c when present; isolation per name; teach/recall demo fixtures (Gitea #22, #23).
- Reason: L1 persistence is the prerequisite for any multi-step agent work.
- Expected result: session dir populated after first run; second run recalls taught context.
### Entry 10.2 — after
- Timestamp: 2026-09-03
- Commands run: build; mock plumbing run; live teach/recall E2E; task suite.
- Observed result: task suite 32/32; teach run replied REMEMBERED and session JSONL persisted host-side; recall run resumed (-c) and answered exactly 'mosaico'; single continued session file (no duplicate sessions).
- Failure or correction: none. Design note: ephemeral (--no-session) remains the default when no session is declared.
## Phase 11: Operator ergonomics (M7)
### Entry 11.1 — before
- Timestamp: 2026-09-03
- Intended action: mosaic-task.mjs show <runId> (full record + snapshots + artifacts, traversal-safe), list with workspace/session columns, RELEASE -> 0.0.6, package + health-gated activate, docs (Gitea #24).
- Reason: Run records are only as valuable as they are inspectable; release activation closes the loop on container-content changes.
- Expected result: show works for real/missing/traversal ids; suites green; 0.0.6 active.
### Entry 11.2 — after
- Timestamp: 2026-09-03
- Commands run: build; show on real/missing/traversal ids; full sweep; release package + activate.
- Observed result: config 24/24, task 32/32, release 14/14, verify PASS; 0.0.6 packaged and activated via exact-marker health gate.
- Failure or correction: showRun initially rejected valid run ids (lowercase-only regex vs uppercase timestamp) and crashed on missing ids (uncaught readdir) — both fixed and covered.
## Autonomous run result
M5 tagged `workspace-capabilities-v1`, M6 tagged `sessions-v1`, M7 tagged `operator-ergonomics-v1`; release 0.0.6 active. Tracker: docs/plans/2026-09-03_autonomous-run.md.
---
## Phase 12: Conductor loop — self-orchestration (M8)
### Entry 12.1 — before
- Timestamp: 2026-09-03
- Intended action: Stand up the poor-man orchestration loop per docs/plans/CONDUCTOR.md: conductor (host, holds git) mirrors the repo into a worker workspace, dispatches a headless pi worker (session worker-1, tools read/write/edit/bash) to implement retry <runId>, reviews the diff, integrates, verifies (Gitea #25, #26, #27).
- Reason: The owner asked for circular task processing with agent workers; the stack now has every primitive needed — this proves it on the stack itself.
- Expected result: worker-authored retry merged with suites green and a live retry verified.
### Entry 12.2 — after
- Timestamp: 2026-09-03
- Commands run: repo mirror clone; worker dispatch (tasks/worker-retry.json); diff review; apply; live retry; refinement dispatch (tasks/worker-retry-refine.json); second review; reverse+reapply combined patch; conductor interpolation fix; live retry; full sweep.
- Observed result:
- Worker round 1: implemented retry correctly per spec in 2m28s; diff reviewed clean.
- Live retry exposed a spec gap (direct invocation lacks launcher env exports).
- Worker round 2 (same session, 59s): made spawnEnv self-sufficient, but used PI_* names where compose interpolates MOSAIC_*.
- Conductor hotfix: 3-line rename to MOSAIC_PROVIDER/MOSAIC_MODEL/MOSAIC_DATA_ROOT (too trivial for a worker round).
- Final: live retry succeeded (replied REMEMBERED, new run recorded); all suites green.
- Failure or correction: three rounds total — one spec gap (conductor), one naming mismatch (worker), one trivial rename (conductor). Each was caught by mechanical verification (run record stderr), never by hope.
- Attribution: feature authored by headless pi worker (glm-5.3-flash) in sessions worker-1; conductor reviewed, integrated, and hotfixed.
## Result (M8)
Conductor loop proven end-to-end on the stack itself. `main` merged with M8; release 0.0.6 remains active (retry is host-side only, no image change).
---
## Phase 13: Mission-level capability policy (M9)
### Entry 13.1 — before
- Timestamp: 2026-09-03
- Intended action: Missions may declare capabilities.tools as governing constraints (Gitea #30); merge semantics = least-privilege intersection (task narrows, never widens; empty intersection = tool-free run). Host-side only.
- Reason: First mechanical restriction layer — the trust model becomes enforced, not instructed.
- Expected result: all four merge cases asserted from run evidence; suites green.
### Entry 13.2 — after
- Observed: four merge cases verified deterministically via run-record stderr (mission-only, task-only, narrowed, emptied); invalid mission capabilities exit 2; task suite 41/41.
- Failure or correction: selftest harness could not express ABSENT vs EMPTY fields via its printf helper — fixed with an ABSENT marker; two suite config-leak defects fixed (per-command env scoping). Product unaffected.
## Phase 14: Session forking (M11)
### Entry 14.1 — before
- Timestamp: 2026-09-03
- Intended action: sessionForkFrom task field branches the source session's newest file (pi --fork) into the target session dir; ancestor untouched; RELEASE -> 0.0.7 with health-gated activation (Gitea #33).
- Reason: Owner flagged conversation forking from a common ancestor as a desired property; pi JSONL trees make it native.
- Expected result: forked child recalls ancestor context; ancestor file untouched; suites green; 0.0.7 active.
### Entry 14.2 — after
- Observed: mock plumbing asserts fork source + target delivery; validation rejects fork-without-target and self-fork (exit 2); ghost source exits 4; live fork: child recalled 'mosaico' from ancestor context while the ancestor session file remained untouched (file-level assertion); suites 58/24/14 + verify green; 0.0.7 packaged and activated via health gate.
- Failure or correction: retryRun-style self-assignment bug in validation (compared null to target) — caught by negative test, fixed.
## Result (M11)
Session forking verified. `main` merged with M11, tagged `session-fork-v1`; release 0.0.7 active.
+25
View File
@@ -119,6 +119,31 @@ adapters/<name>/adapter.sh env in: MOSAIC_SYSTEM_PROMPT_FILE, MOSAIC_REQUEST,
See `adapters/README.md` for the full contract.
## Workspaces, capabilities, sessions (M5/M6)
Optional task fields extend what an agent can do — all defaulting to the previous behavior:
```json
{
"workspace": "demo", // ":run" ephemeral, or persistent dataRoot/workspaces/<name>
"capabilities": { "tools": ["bash", "read"] }, // pi tool allowlist; absent = no tools
"session": "demo" // persistent session at dataRoot/sessions/<name>
}
```
- The adapter runs inside the workspace; files it writes are host-visible (`dataRoot/workspaces/<name>`).
- Sessions persist via pi's documented `--session-dir`; a follow-up run in the same session resumes the conversation (`-c`) and can recall prior context. Distinct names never share state. Ephemeral (`--no-session`) remains the default when no session is declared.
- Selection authority: config for adapter/provider/model; the task file for workspace/capabilities/session.
Inspect anything:
```bash
node scripts/mosaic-task.mjs list # runs with task/workspace/session columns
node scripts/mosaic-task.mjs show <runId> # full record + snapshots + artifacts
```
Demo fixtures: `tasks/workspace-demo.json`, `tasks/session-demo-1.json` + `tasks/session-demo-2.json`.
See `docs/plans/2026-09-02_atomic-mosaic-foundation.md` for the full plan.
Inside the container:
+1 -1
View File
@@ -1 +1 @@
0.0.5
0.0.7
+9 -4
View File
@@ -18,11 +18,16 @@ if [ -n "${MOSAIC_WORKSPACE:-}" ]; then
cd "$MOSAIC_WORKSPACE"
fi
# Session (M6): persistent named session directory; resume the most recent
# session in that directory when one exists (pi documented flags).
# Default remains ephemeral (--no-session) when no session is declared.
# Session (M6/M11): default ephemeral (--no-session). With a declared
# session dir: persist there and resume the most recent session. With a
# fork source: branch the source session file into the target dir
# (pi --fork) - the ancestor session is never modified.
SESSION_FLAGS="--no-session"
if [ -n "${MOSAIC_SESSION_DIR:-}" ]; then
if [ -n "${MOSAIC_SESSION_FORK:-}" ]; then
[ -n "${MOSAIC_SESSION_DIR:-}" ] || { echo "pi adapter: session fork requires MOSAIC_SESSION_DIR" >&2; exit 2; }
mkdir -p "$MOSAIC_SESSION_DIR"
SESSION_FLAGS="--fork $MOSAIC_SESSION_FORK --session-dir $MOSAIC_SESSION_DIR"
elif [ -n "${MOSAIC_SESSION_DIR:-}" ]; then
mkdir -p "$MOSAIC_SESSION_DIR"
SESSION_FLAGS="--session-dir $MOSAIC_SESSION_DIR"
if [ -n "$(ls -A "$MOSAIC_SESSION_DIR" 2>/dev/null)" ]; then
+2 -1
View File
@@ -18,8 +18,9 @@ services:
# Workspace + capabilities (set by the task runner; M5)
MOSAIC_WORKSPACE: ${MOSAIC_WORKSPACE:-}
MOSAIC_TOOLS: ${MOSAIC_TOOLS:-}
# Persistent named session dir (set by the task runner; M6)
# Persistent named session dir + optional fork source (M6/M11)
MOSAIC_SESSION_DIR: ${MOSAIC_SESSION_DIR:-}
MOSAIC_SESSION_FORK: ${MOSAIC_SESSION_FORK:-}
# mock adapter only: verbatim response for deterministic seam tests
MOSAIC_MOCK_RESPONSE: ${MOSAIC_MOCK_RESPONSE:-}
# Documented container auth alternative: provider API key via
+19
View File
@@ -0,0 +1,19 @@
{
"policyVersion": 1,
"autoApply": {
"enabled": true,
"allowedPaths": [
"scripts/**",
"docs/**",
"tasks/**",
"missions/**",
"adapters/**",
"README.md"
],
"suites": [
"test-config",
"test-task",
"test-release"
]
}
}
+12 -9
View File
@@ -13,8 +13,8 @@ Advance the Mosaic Stack rebuild several verified layers in one batch, focused o
|---|---|---|
| M5 | Task workspaces + capability envelope (tools allowlist) | DONE |
| M6 | Named sessions — persistence and resume (L1) | DONE |
| M7 | Operator ergonomics: run inspection commands | DONE (partial by design) |
| — | Release 0.0.6 packaged + health-gated activated | DONE |
| M7 | Operator ergonomics: run inspection commands | DONE |
| — | Releases 0.0.5+0.0.6 packaged; 0.0.6 health-gated activated | DONE |
Explicitly deferred (do not mistake for forgotten):
- Claude/Codex/OpenCode adapters (owner: focus on Pi for now)
@@ -60,14 +60,17 @@ Explicitly deferred (do not mistake for forgotten):
6. Inspect any run: `node scripts/mosaic-task.mjs show <runId>`
7. Gitea: milestones M5/M6/M7 closed; issues referenced by merge commits
## Results (to be filled at end of run)
## Results
- M5: pending
- M6: pending
- M7: pending
- Release activation: pending
- Final suite counts: pending
- Corrections encountered: pending
- M5 merged on `main` (merge commit `ddb1554`), tagged `workspace-capabilities-v1`
- M6 merged on `main` (merge commit `4e2a413`), tagged `sessions-v1`
- M7 merged on `main`, tagged `operator-ergonomics-v1`
- Release 0.0.6 packaged, health-gated activated, full sweep green
- Suites at end of run: config 24/24, task 32/32, release 14/14, verify PASS
- Build log: Phases 9 (M5), 10 (M6), 11 (M7) appended with corrections
- Corrections encountered: dropped constant from a failed atomic edit batch (SUPPORTED_TOOLS); dash `export` output format vs `env`; three selftest authoring defects; showRun id-regex case sensitivity + missing-run crash. All fixed and covered by tests.
- Commits pushed incrementally; nothing left uncommitted
- Live proof: workspace file host-visible; session teach/recall ('mosaico') verified
## Next steps after this run (not started)
+51
View File
@@ -0,0 +1,51 @@
# Conductor protocol — poor-man orchestration loop
How the stack orchestrates headless pi workers to do work on itself.
## Roles
| Role | Runs where | Powers | Never has |
|---|---|---|---|
| **Conductor** | host (assistant or owner) | git (clone/commit/push), task dispatch, review, verification suites, Gitea | nothing new |
| **Worker** | container (headless pi via `scripts/run-task.sh`) | read/write/edit/bash inside its workspace; persistent session on request | git credentials, docker socket, host filesystem |
## The loop
1. **Decompose**: conductor turns a goal into worker tasks small enough to
specify completely in one prompt (file paths, acceptance criteria, style
constraints, verification the worker can run itself, e.g. `node --check`).
2. **Mirror**: conductor maintains the repo clone at
`<dataRoot>/workspaces/stack-repo` (host-side git; workers see it read-write
through their workspace mount).
3. **Dispatch**: `scripts/run-task.sh run <worker-task.json>` — worker edits the
clone. Session name `worker-<n>` keeps continuity across refinement rounds.
4. **Extract**: `git -C <workspace> diff > patch` — the worker's entire output
is a reviewable diff. Run record (result.json, stderr.txt) is the receipt.
5. **Review**: conductor reads the diff line by line. Bad output → refine the
prompt, re-dispatch (same session: "your patch had these problems…").
6. **Integrate**: conductor applies the patch to the real repo, runs the full
suites, commits and pushes. Suites failing → revert apply, back to step 5.
7. **Record**: update CURRENT.md, BUILD-LOG, close the Gitea issue.
## Guardrails
- Workers never receive credentials; they never run git; they never leave the
workspace (container is the boundary; tools allowlist is the gate).
- Every worker diff is reviewed by the conductor before integration. No
auto-apply. (Auto-apply would be a capability-policy decision for later.)
- Verification is mechanical: suites + `node --check` / `bash -n` gates.
- Recursive decomposition = "fail → smaller task", never "hope."
## Worker task template
```json
{
"taskVersion": 1,
"id": "t-worker-<name>",
"prompt": "<full spec: goal, files, constraints, acceptance, self-checks>",
"workspace": "stack-repo",
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
"session": "worker-1",
"timeoutSeconds": 600
}
```
+47
View File
@@ -0,0 +1,47 @@
# CURRENT — single source of "what happens next"
This file always names exactly one next action. Any "continue" / "next" /
"proceed" message means: execute the action below, fully (implement → test →
verify against its acceptance criteria → commit → push → close the issue →
update this file to the next action). No ambiguity, no re-planning.
## Next action
Owner review of M11 (session forking) — then name the next target.
## Queue (ordered, not started)
1. Second real adapter (parked — owner focused on Pi)
2. Auto-apply policy for worker patches (substrate exists)
3. UX/DX backlog #31 (grep highlight vs status colors — likely closed by owner's unalias)
## Rules
- One action in flight. Update this file at the END of every action.
- Blocked? Move the item to "Blocked" below with the reason and stop.
- Completed actions move to the log at the bottom (date + issue + result).
## Blocked
(none)
## Completed log
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
- 2026-09-03 — M11 session forking (#33) — merged, child recalls ancestor context, ancestor untouched, 58/24/14 + verify green; release 0.0.7 activated
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection
- 2026-09-03 — test UX: green OK/red FAIL status colors (#31 adjacent) — terminal-only, pipe-safe
- 2026-09-03 — M10 run-record retention (#32) — merged, prune keep-N/dry-run/receipt, 49/24/14 + verify green
- 2026-09-03 — M5 workspaces + capability envelope (#20, #21) — merged, tests green
- 2026-09-03 — M6 named sessions + resume (#22, #23) — merged, teach/recall verified
- 2026-09-03 — M7 run inspection + release 0.0.6 (#24) — merged, activated
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
- 2026-09-03 — retry lineage + relative mission resolution (#28) — merged, 36/24/14 suites + verify green
- 2026-09-03 — M9 mission capability policy (#30) — merged, least-privilege intersection, 41/36/14 + verify green
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
- 2026-09-03 — retry lineage + relative mission resolution (#28) — merged, 36/24/14 suites + verify green
- 2026-09-03 — M8 conductor loop + worker-built retry (#25, #26, #27) — merged; worker authored retry in 2 refinement rounds, conductor fixed a 3-line interpolation rename; live retry verified
+139
View File
@@ -0,0 +1,139 @@
#!/usr/bin/env bash
# Conductor auto-apply: integrate a worker's patch under the declared policy.
#
# Usage: scripts/conductor-apply.sh <runId> [--dry-run]
#
# Policy (conductor-policy.json in the target repo, strictly validated):
# autoApply.enabled master switch
# autoApply.allowedPaths glob allowlist ('dir/**' = everything under dir)
# autoApply.suites suite scripts that must pass AFTER applying
#
# Gate sequence: succeeded run record -> clean target tree -> diff extracted
# from the worker workspace -> allowlist -> syntax gates (node/bash/json) ->
# apply -> policy suites -> commit with attribution. ANY failure reverts the
# working tree and exits nonzero. Push is never automatic.
#
# Environment:
# MOSAIC_APPLY_TARGET repo root to apply into (default: this project)
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
TARGET_ROOT="${MOSAIC_APPLY_TARGET:-$(cd "$SCRIPT_DIR/.." && pwd)}"
RUN_ID="${1:?usage: conductor-apply.sh <runId> [--dry-run]}"
DRY_RUN="no"
[ "${2:-}" = "--dry-run" ] && DRY_RUN="yes"
cd "$TARGET_ROOT"
fail() { echo "conductor-apply: $*" >&2; exit "${2:-1}"; }
[ -d .git ] || fail "target is not a git repository: $TARGET_ROOT" 4
[ -f conductor-policy.json ] || fail "no conductor-policy.json in target" 2
# ---- policy (strict) ----
POLICY_JSON="$(node -e '
const fs = require("fs");
const p = JSON.parse(fs.readFileSync("conductor-policy.json", "utf8"));
if (p.policyVersion !== 1) process.exit(3);
if (!p.autoApply || typeof p.autoApply.enabled !== "boolean" || !Array.isArray(p.autoApply.allowedPaths) || !Array.isArray(p.autoApply.suites)) process.exit(3);
for (const g of p.autoApply.allowedPaths) {
if (typeof g !== "string" || !/^[A-Za-z0-9_.*/-]+$/.test(g) || g.startsWith("/") || g.includes("..")) process.exit(3);
}
console.log(JSON.stringify(p.autoApply));
')" || fail "invalid conductor-policy.json" 2
ENABLED="$(node -e 'console.log(JSON.parse(process.argv[1]).enabled)' "$POLICY_JSON")"
[ "$ENABLED" = "true" ] || fail "auto-apply is disabled by policy" 2
# ---- run record ----
DATA_ROOT="$(node scripts/mosaic-config.mjs validate | node -e 'let d="";process.stdin.on("data",c=>d+=c).on("end",()=>console.log(JSON.parse(d).dataRoot))')"
RESULT_FILE="$DATA_ROOT/runs/$RUN_ID/result.json"
[ -f "$RESULT_FILE" ] || fail "run not found: $RUN_ID" 4
node -e '
const r = JSON.parse(require("fs").readFileSync(process.argv[1], "utf8"));
process.exit(r.status === "succeeded" ? 0 : 1);
' "$RESULT_FILE" || fail "run $RUN_ID did not succeed; refusing to auto-apply"
WORKSPACE_NAME="$(node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.workspace||"")' "$RESULT_FILE")"
[ -n "$WORKSPACE_NAME" ] || fail "run has no workspace; nothing to integrate" 4
WORKSPACE="$DATA_ROOT/workspaces/$WORKSPACE_NAME"
[ -d "$WORKSPACE/.git" ] || fail "workspace is not a git clone: $WORKSPACE" 4
# ---- extract diff (tracked + intent-to-add) ----
git -C "$WORKSPACE" add -N . >/dev/null 2>&1 || true
DIFF_FILE="$(mktemp)"
trap 'rm -f "$DIFF_FILE"' EXIT
git -C "$WORKSPACE" diff > "$DIFF_FILE"
if [ ! -s "$DIFF_FILE" ]; then
fail "workspace has no changes to apply"
fi
# ---- allowlist ----
mapfile -t CHANGED < <(git -C "$WORKSPACE" diff --name-only)
GLOBS="$(node -e 'const a=JSON.parse(process.argv[1]).allowedPaths;console.log(a.join("\n"))' "$POLICY_JSON")"
REFUSED=""
for f in "${CHANGED[@]}"; do
ok="no"
while IFS= read -r g; do
[ -z "$g" ] && continue
case "$f" in
$g) ok="yes"; break ;;
esac
done <<< "$GLOBS"
[ "$ok" = "yes" ] || REFUSED="$REFUSED $f"
done
if [ -n "$REFUSED" ]; then
echo "conductor-apply: refusing - files outside policy allowlist:$REFUSED" >&2
echo "conductor-apply: patch preserved at $DIFF_FILE for manual review" >&2
exit 1
fi
# ---- syntax gates (on workspace files, pre-apply) ----
for f in "${CHANGED[@]}"; do
case "$f" in
*.mjs) node --check "$WORKSPACE/$f" || fail "syntax gate failed (node): $f" 1 ;;
*.sh) bash -n "$WORKSPACE/$f" || fail "syntax gate failed (bash): $f" 1 ;;
*.json) node -e 'JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"))' "$WORKSPACE/$f" || fail "syntax gate failed (json): $f" 1 ;;
esac
done
if [ "$DRY_RUN" = "yes" ]; then
echo "conductor-apply (dry-run): would apply $(echo "${#CHANGED[@]}") file(s) from $RUN_ID:"
printf ' %s\n' "${CHANGED[@]}"
echo "conductor-apply (dry-run): suites would run: $(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(", "))' "$POLICY_JSON")"
exit 0
fi
# ---- apply ----
[ -z "$(git status --porcelain)" ] || fail "target tree is not clean; refusing to mix states"
git apply "$DIFF_FILE" || fail "git apply failed"
# ---- policy suites ----
SUITES="$(node -e 'console.log(JSON.parse(process.argv[1]).suites.join(" "))' "$POLICY_JSON")"
SUITES_OK="yes"
for s in $SUITES; do
case "$s" in
test-[a-z]*) : ;; # shape guard; existence checked next
*) echo "conductor-apply: refusing suspicious suite name: $s" >&2; SUITES_OK="no"; break ;;
esac
[ -x "scripts/$s.sh" ] || { echo "conductor-apply: suite script missing: scripts/$s.sh" >&2; SUITES_OK="no"; break; }
if ! bash "scripts/$s.sh" >/dev/null 2>&1; then
echo "conductor-apply: suite failed: $s" >&2
SUITES_OK="no"
break
fi
done
if [ "$SUITES_OK" != "yes" ]; then
git apply -R "$DIFF_FILE" && echo "conductor-apply: changes REVERTED (suites failed)" >&2
exit 1
fi
# ---- commit with attribution ----
git add -A
git commit -q -m "feat(worker): auto-applied patch from run $RUN_ID
Authored-by: pi worker (run $RUN_ID, workspace $WORKSPACE_NAME)
Applied-under: conductor-policy v1 (allowlist + syntax gates + suites)"
echo "conductor-apply: applied and committed run $RUN_ID ($(echo "${#CHANGED[@]}") file(s)); suites: $SUITES"
echo "conductor-apply: NOT pushed - push remains an explicit act."
+257 -8
View File
@@ -28,6 +28,7 @@
*/
import fs from "node:fs";
import os from "node:os";
import path from "node:path";
import process from "node:process";
import { randomBytes } from "node:crypto";
@@ -83,7 +84,7 @@ function validateId(value, what) {
function validateMission(document, file) {
if (!isPlainObject(document)) fail(2, "mission must be a JSON object");
rejectUnknownKeys(document, ["missionVersion", "id", "objective", "directives"], "mission");
rejectUnknownKeys(document, ["missionVersion", "id", "objective", "directives", "capabilities"], "mission");
if (document.missionVersion !== 1) {
fail(2, `unsupported missionVersion: ${JSON.stringify(document.missionVersion)} (supported: 1)`);
}
@@ -101,12 +102,34 @@ function validateMission(document, file) {
return d;
});
}
return { missionVersion: document.missionVersion, id: document.id, objective: document.objective, directives };
// Governing capability constraints (M9): same validation as task
// capabilities; semantically these BOUND tasks (least-privilege
// intersection at run time), never grant beyond them.
let capabilities = null;
if (document.capabilities !== undefined && document.capabilities !== null) {
if (!isPlainObject(document.capabilities)) fail(2, 'mission "capabilities" must be a JSON object');
rejectUnknownKeys(document.capabilities, ["tools"], 'mission "capabilities"');
if (!Array.isArray(document.capabilities.tools) || document.capabilities.tools.length === 0) {
fail(2, 'mission "capabilities.tools" must be a non-empty array of tool names');
}
const seen = new Set();
for (const tool of document.capabilities.tools) {
if (!SUPPORTED_TOOLS.includes(tool)) {
fail(2, `unsupported mission tool: ${JSON.stringify(tool)} (supported: ${SUPPORTED_TOOLS.join(", ")})`);
}
if (seen.has(tool)) fail(2, `duplicate tool in mission capabilities.tools: ${tool}`);
seen.add(tool);
}
capabilities = { tools: [...seen] };
}
return { missionVersion: document.missionVersion, id: document.id, objective: document.objective, directives, capabilities };
}
function validateTask(document, file) {
if (!isPlainObject(document)) fail(2, "task must be a JSON object");
rejectUnknownKeys(document, ["taskVersion", "id", "prompt", "mission", "expectExact", "timeoutSeconds", "workspace", "capabilities", "session"], "task");
rejectUnknownKeys(document, ["taskVersion", "id", "prompt", "mission", "expectExact", "timeoutSeconds", "workspace", "capabilities", "session", "sessionForkFrom"], "task");
if (document.taskVersion !== 1) {
fail(2, `unsupported taskVersion: ${JSON.stringify(document.taskVersion)} (supported: 1)`);
}
@@ -169,6 +192,23 @@ function validateTask(document, file) {
session = document.session;
}
// Session fork (M11): optional source session whose newest session file
// is branched (pi --fork) into the target session dir. Requires session.
let sessionForkFrom = null;
if (document.sessionForkFrom !== undefined && document.sessionForkFrom !== null) {
if (typeof document.sessionForkFrom !== "string" || document.sessionForkFrom.length === 0) {
fail(2, 'task "sessionForkFrom" must be a non-empty string when present');
}
validateId(document.sessionForkFrom, "task sessionForkFrom");
if (!session) {
fail(2, 'task "sessionForkFrom" requires "session" (the fork target) to be set');
}
if (document.sessionForkFrom === session) {
fail(2, 'task "sessionForkFrom" must differ from "session" (cannot fork onto itself)');
}
sessionForkFrom = document.sessionForkFrom;
}
// Capabilities (M5): optional tools allowlist mapped by adapters to their
// native permission flags. Absent = no tools.
let tools = null;
@@ -201,6 +241,7 @@ function validateTask(document, file) {
workspace,
tools,
session,
sessionForkFrom,
};
}
@@ -233,7 +274,7 @@ function writeOnce(file, content) {
}
}
function runTask(taskFile) {
function runTask(taskFile, options = {}) {
const resolved = JSON.parse(
spawnSync(process.execPath, [path.join(PROJECT_ROOT, "scripts", "mosaic-config.mjs"), "validate"], {
cwd: PROJECT_ROOT,
@@ -263,6 +304,15 @@ function runTask(taskFile) {
// CONTAINER path (dataRoot maps to /var/lib/mosaic in the image).
const spawnEnv = { ...process.env };
spawnEnv.MOSAIC_ADAPTER = resolved.execution.adapter;
// Self-sufficient env: direct invocation (e.g. `retry`) skips the shell
// launcher exports, so derive them from the resolved config and release.
// (Names here are the compose interpolation consumers, not PI_*.)
spawnEnv.MOSAIC_PROVIDER = resolved.execution.provider;
spawnEnv.MOSAIC_MODEL = resolved.execution.model;
spawnEnv.MOSAIC_DATA_ROOT = resolved.dataRoot;
const release = fs.readFileSync(path.join(PROJECT_ROOT, "RELEASE"), "utf8").trim();
const piVersion = JSON.parse(fs.readFileSync(path.join(PROJECT_ROOT, "package.json"), "utf8")).dependencies["@earendil-works/pi-coding-agent"];
spawnEnv.MOSAIC_IMAGE_TAG = `mosaic-poc-agent:${piVersion}-r${release}`;
if (task.missionSnapshot) {
const relative = path.relative(resolved.dataRoot, runDir);
if (relative.startsWith("..") || path.isAbsolute(relative)) {
@@ -281,7 +331,24 @@ function runTask(taskFile) {
workspaceContainerPath = `/var/lib/mosaic/workspaces/${task.workspace}`;
}
if (workspaceContainerPath) spawnEnv.MOSAIC_WORKSPACE = workspaceContainerPath;
spawnEnv.MOSAIC_TOOLS = task.tools ? task.tools.join(",") : "";
// Capability policy (M9): least-privilege intersection. A task may narrow
// a mission's tool grant, never widen it. Empty intersection = tool-free.
let effectiveTools = task.tools;
let policyNote = null;
if (task.missionSnapshot?.capabilities) {
const missionTools = task.missionSnapshot.capabilities.tools;
if (effectiveTools) {
effectiveTools = effectiveTools.filter((t) => missionTools.includes(t));
if (effectiveTools.length === 0) {
policyNote = `capability policy: mission ${task.missionSnapshot.id} and task request no tools in common -> tool-free run`;
}
} else {
effectiveTools = [...missionTools];
}
}
if (policyNote) process.stderr.write(`mosaic-task: ${policyNote}\n`);
spawnEnv.MOSAIC_TOOLS = effectiveTools ? effectiveTools.join(",") : "";
// Session (M6): persistent named session dir, passed as container path.
if (task.session) {
@@ -289,6 +356,26 @@ function runTask(taskFile) {
spawnEnv.MOSAIC_SESSION_DIR = `/var/lib/mosaic/sessions/${task.session}`;
}
// Session fork (M11): resolve the source session's newest file; pi --fork
// branches it into the target dir without modifying the ancestor.
if (task.sessionForkFrom) {
const sourceDir = path.join(resolved.dataRoot, "sessions", task.sessionForkFrom);
let sources = [];
try {
sources = fs.readdirSync(sourceDir).filter((f) => f.endsWith(".jsonl")).sort();
} catch {
fail(4, `cannot fork: source session dir not found: ${sourceDir}`);
}
if (sources.length === 0) {
fail(4, `cannot fork: source session '${task.sessionForkFrom}' has no session files`);
}
const relative = path.relative(resolved.dataRoot, path.join(sourceDir, sources[sources.length - 1]));
if (relative.startsWith("..") || path.isAbsolute(relative)) {
fail(4, `source session is outside the configured dataRoot: ${sourceDir}`);
}
spawnEnv.MOSAIC_SESSION_FORK = `/var/lib/mosaic/${relative.split(path.sep).join("/")}`;
}
const proc = spawnSync(
"docker",
["compose", "run", "--rm", "-T", "mosaic-agent", task.prompt],
@@ -335,9 +422,11 @@ function runTask(taskFile) {
request: task.prompt,
response,
expectedExact: expected,
...(options.retriedFrom ? { retriedFrom: options.retriedFrom } : {}),
workspace: task.workspace,
tools: task.tools,
tools: effectiveTools,
session: task.session,
sessionForkFrom: task.sessionForkFrom,
exitCode: proc.status,
signal: proc.signal ?? null,
provider: resolved.execution.provider,
@@ -365,17 +454,166 @@ function listRuns() {
for (const runId of entries) {
let status = "unknown";
let taskId = "-";
let workspace = "-";
let session = "-";
try {
const result = JSON.parse(fs.readFileSync(path.join(root, runId, "result.json"), "utf8"));
status = result.status;
taskId = result.taskId;
workspace = result.workspace ?? "-";
session = result.session ?? "-";
} catch {
// Incomplete run record; report as unknown.
}
process.stdout.write(`${runId} ${status.padEnd(9)} ${taskId}\n`);
process.stdout.write(`${runId} ${status.padEnd(9)} task=${taskId.padEnd(18)} ws=${String(workspace).padEnd(10)} session=${session}\n`);
}
}
function showRun(runId) {
if (!/^r-[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(runId)) {
fail(4, `invalid run id: ${JSON.stringify(runId)} (expected r-<id>)`);
}
const resolved = loadConfig();
const dir = path.join(runsRoot(resolved), runId);
if (!fs.existsSync(dir)) {
fail(4, `run not found: ${runId} (under ${runsRoot(resolved)})`);
}
const read = (name) => {
try {
return JSON.parse(fs.readFileSync(path.join(dir, name), "utf8"));
} catch {
return null;
}
};
const result = read("result.json");
const task = read("task.json");
const mission = read("mission.json");
process.stdout.write(`run: ${runId}\n`);
if (result) {
process.stdout.write(
[
`status: ${result.status}${result.reason ? ` (${result.reason})` : ""}`,
`task: ${result.taskId}`,
result.missionId ? `mission: ${result.missionId}` : null,
result.workspace ? `workspace: ${result.workspace}` : null,
result.session ? `session: ${result.session}` : null,
result.tools ? `tools: ${result.tools.join(", ")}` : null,
`adapter: ${"(see config)"} provider=${result.provider} model=${result.model}`,
`request: ${JSON.stringify(result.request)}`,
`response: ${JSON.stringify(result.response)}`,
result.expectedExact !== null && result.expectedExact !== undefined ? `expected: ${JSON.stringify(result.expectedExact)}` : null,
result.retriedFrom ? `retriedFrom: ${result.retriedFrom}` : null,
`timing: ${result.startedAt} -> ${result.finishedAt} (${result.durationMs} ms)`,
`exit: ${result.exitCode}${result.signal ? ` signal=${result.signal}` : ""}`,
].filter((line) => line !== null).join("\n") + "\n",
);
} else {
process.stdout.write("result.json: (missing or unreadable)\n");
}
if (task) process.stdout.write(`task snapshot: ${"task.json"} present\n`);
if (mission) process.stdout.write(`mission snapshot: ${mission.id} - ${mission.objective}\n`);
process.stdout.write(`artifacts: ${fs.readdirSync(dir).map((f) => `${f}`).join(", ")}\n`);
process.exit(0);
}
function retryRun(runId) {
if (!/^r-[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(runId)) {
fail(4, `invalid run id: ${JSON.stringify(runId)} (expected r-<id>)`);
}
const resolved = loadConfig();
const dir = path.join(runsRoot(resolved), runId);
if (!fs.existsSync(dir)) {
fail(4, `run not found: ${runId}`);
}
let snapshot;
try {
snapshot = fs.readFileSync(path.join(dir, "task.json"), "utf8");
} catch {
fail(4, `run task snapshot unreadable: ${runId}`);
}
// Relative mission paths in a snapshot resolve against the ORIGINAL task
// location, which no longer exists here — rewrite them to the run's own
// recorded mission.json so retries stay faithful.
let snapshotDoc;
try {
snapshotDoc = JSON.parse(snapshot);
} catch {
fail(4, `run task snapshot is not valid JSON: ${runId}`);
}
if (snapshotDoc.mission && !path.isAbsolute(snapshotDoc.mission)) {
const recordedMission = path.join(dir, "mission.json");
if (!fs.existsSync(recordedMission)) {
fail(4, `cannot retry ${runId}: relative mission path but no mission.json snapshot in run dir`);
}
snapshotDoc.mission = recordedMission;
snapshot = `${JSON.stringify(snapshotDoc, null, 2)}\n`;
}
// A retry is a brand-new run: replay the recorded task snapshot through
// the ordinary run path; existing run records stay untouched.
const tempDir = fs.mkdtempSync(path.join(os.tmpdir(), "mosaic-retry-"));
const tempTaskFile = path.join(tempDir, "task.json");
fs.writeFileSync(tempTaskFile, snapshot);
process.on("exit", () => {
try {
fs.rmSync(tempDir, { recursive: true, force: true });
} catch {
// Best-effort cleanup only.
}
});
runTask(tempTaskFile, { retriedFrom: runId });
}
function pruneRuns(args) {
const resolved = loadConfig();
const root = runsRoot(resolved);
let keep = 50;
let apply = false;
for (const arg of args) {
if (arg === "--yes") apply = true;
else if (arg === "--keep") fail(4, "prune: --keep requires a value");
else if (arg.startsWith("--keep=")) {
keep = Number(arg.slice("--keep=".length));
if (!Number.isInteger(keep) || keep < 1) fail(4, `--keep must be a positive integer (got ${arg.slice(7)})`);
} else fail(4, `unknown prune option: ${arg}`);
}
let entries = [];
try {
entries = fs.readdirSync(root, { withFileTypes: true })
.filter((e) => e.name.startsWith("r-") && e.isDirectory() && !e.isSymbolicLink())
.map((e) => e.name)
.sort();
} catch {
// No runs yet.
}
if (entries.length <= keep) {
process.stdout.write(`prune: ${entries.length} run(s) present, keep=${keep} -> nothing to prune\n`);
process.exit(0);
}
const doomed = entries.slice(0, entries.length - keep); // oldest first
if (!apply) {
process.stdout.write(`prune (dry-run): would remove ${doomed.length} oldest run(s), keep ${entries.length - doomed.length}:\n`);
for (const id of doomed) process.stdout.write(` would remove: ${id}\n`);
process.stdout.write("prune: re-run with --yes to apply\n");
process.exit(0);
}
const receipt = path.join(root, ".pruned.log");
for (const id of doomed) {
fs.rmSync(path.join(root, id), { recursive: true, force: true });
fs.appendFileSync(receipt, `${JSON.stringify({ at: new Date().toISOString(), event: "pruned", runId: id })}\n`);
}
process.stdout.write(`prune: removed ${doomed.length} run(s), kept ${keep}; receipt: ${receipt}\n`);
process.exit(0);
}
const operation = process.argv[2];
const target = process.argv[3];
@@ -392,9 +630,20 @@ switch (operation) {
if (!target) fail(4, "usage: mosaic-task.mjs run <taskFile>");
runTask(path.resolve(target));
break;
case "show":
if (!target) fail(4, "usage: mosaic-task.mjs show <runId>");
showRun(target);
break;
case "list":
listRuns();
process.exit(0);
case "prune":
pruneRuns(process.argv.slice(3));
break;
case "retry":
if (!target) fail(4, "usage: mosaic-task.mjs retry <runId>");
retryRun(target);
break;
default:
fail(4, `unknown operation: ${JSON.stringify(operation ?? "")} (expected validate | run | list)`);
fail(4, `unknown operation: ${JSON.stringify(operation ?? "")} (expected validate | run | show | list | retry | prune)`);
}
+140
View File
@@ -0,0 +1,140 @@
#!/usr/bin/env bash
# Sandboxed selftests for the conductor auto-apply policy gate.
#
# Builds a throwaway target repo + worker workspace + fake run records, then
# exercises every gate: policy validation, allowlist, syntax gates, suite
# failure revert, disabled policy, missing/failed runs. No real model calls.
set -uo pipefail
cd "$(dirname "$0")/.."
SANDBOX="$(mktemp -d)"
trap 'rm -rf "$SANDBOX"' EXIT
PASS=0
FAIL=0
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
check_rc() { # name expectedRc command...
local name="$1" expected="$2"
shift 2
local rc
"$@" >/dev/null 2>&1
rc=$?
if [ "$rc" -eq "$expected" ]; then
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
else
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
fi
}
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
}
# ---- infrastructure: target repo + worker workspace + fake run ----
git clone -q . "$SANDBOX/repo"
# The clone carries committed state only - give the target its policy and
# commit it so the tree starts clean (untracked policy would fail target_clean).
cp conductor-policy.json "$SANDBOX/repo/conductor-policy.json"
git -C "$SANDBOX/repo" add conductor-policy.json
git -C "$SANDBOX/repo" -c user.name=suite -c user.email=suite@local commit -q -m policy
mkdir -p "$SANDBOX/data/workspaces" "$SANDBOX/data/runs"
git clone -q "$SANDBOX/repo" "$SANDBOX/data/workspaces/stack-repo"
TARGET="$SANDBOX/repo"
WS="$SANDBOX/data/workspaces/stack-repo"
cat > "$SANDBOX/config.json" <<EOF
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m","adapter":"mock"}}
EOF
export MOSAIC_APPLY_TARGET="$TARGET"
export MOSAIC_CONFIG="$SANDBOX/config.json"
RUN_OK="r-20260903T000000000Z-ok0000001"
mkdir -p "$SANDBOX/data/runs/$RUN_OK"
printf '{"runVersion":1,"runId":"%s","taskId":"t-fake","status":"succeeded","workspace":"stack-repo"}' "$RUN_OK" \
> "$SANDBOX/data/runs/$RUN_OK/result.json"
ws_edit() { printf '\n%s\n' "$2" >> "$WS/$1"; }
ws_reset() { git -C "$WS" checkout -q -- . 2>/dev/null; git -C "$WS" clean -qfd; }
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
set_policy() { # enabled suites (commits: the target tree must stay clean)
local suites="[\"$2\"]"
printf '{"policyVersion":1,"autoApply":{"enabled":%s,"allowedPaths":["scripts/**","docs/**","README.md"],"suites":%s}}' "$1" "$suites" \
> "$TARGET/conductor-policy.json"
git -C "$TARGET" add conductor-policy.json
git -C "$TARGET" -c user.name=suite -c user.email=suite@local commit -q -m "policy update"
}
set_policy true "test-config"
# T1: dry run - allowed change, nothing applied
ws_edit "README.md" "worker dry-run line"
check_rc "dry-run: allowed change, exit 0, nothing committed" 0 \
scripts/conductor-apply.sh "$RUN_OK" --dry-run
if git -C "$TARGET" log --format=%s | grep -q "auto-applied"; then
check "dry-run committed nothing" 1
else
check "dry-run committed nothing" 0
fi
ws_reset
# T2: apply - allowed change, suites pass, commit created
ws_edit "README.md" "worker applied line"
check_rc "apply: allowed change exits 0" 0 scripts/conductor-apply.sh "$RUN_OK"
git -C "$TARGET" log -1 --format=%s | grep -q "auto-applied patch from run $RUN_OK" \
&& check "apply: attribution in commit subject" 0 || check "apply: attribution in commit subject" 1
target_clean() { [ -z "$(git -C "$TARGET" status --porcelain)" ]; }
target_clean && check "apply: target tree clean after commit" 0 || check "apply: target tree clean after commit" 1
git -C "$TARGET" reset -q --hard HEAD~1
# T3: disallowed path refused
ws_edit "Containerfile" "# worker touch"
check_rc "disallowed path refused" 1 scripts/conductor-apply.sh "$RUN_OK"
target_clean && check "disallowed path: target untouched" 0 || check "disallowed path: target untouched" 1
ws_reset
# T4: syntax gate - broken .mjs on an allowed path
printf 'this is not (valid js\n' > "$WS/scripts/broken-worker.mjs"
check_rc "syntax gate refused broken .mjs" 1 scripts/conductor-apply.sh "$RUN_OK"
target_clean && check "syntax gate: target untouched" 0 || check "syntax gate: target untouched" 1
ws_reset
# T5: suite failure - allowed change breaks a policy suite -> auto-revert
printf '\nexit 7\n' >> "$TARGET/scripts/test-config.sh"
ws_edit "README.md" "worker change that will fail suites"
check_rc "suite failure refused" 1 scripts/conductor-apply.sh "$RUN_OK"
git -C "$TARGET" checkout -q -- scripts/test-config.sh
target_clean && check "suite failure: target reverted to clean" 0 || check "suite failure: target reverted to clean" 1
# T6: disabled policy
set_policy false "test-config"
ws_edit "README.md" "worker line while disabled"
check_rc "disabled policy refused" 2 scripts/conductor-apply.sh "$RUN_OK"
target_clean && check "disabled policy: target untouched" 0 || check "disabled policy: target untouched" 1
ws_reset
set_policy true "test-config"
# T7: failed run refused
RUN_FAIL="r-20260903T000000000Z-fail00001"
mkdir -p "$SANDBOX/data/runs/$RUN_FAIL"
printf '{"runVersion":1,"runId":"%s","taskId":"t","status":"failed","workspace":"stack-repo"}' "$RUN_FAIL" \
> "$SANDBOX/data/runs/$RUN_FAIL/result.json"
ws_edit "README.md" "worker line from failed run"
check_rc "failed run refused" 1 scripts/conductor-apply.sh "$RUN_FAIL"
target_clean && check "failed run: target untouched" 0 || check "failed run: target untouched" 1
ws_reset
# T8/T9: missing run + invalid policy
check_rc "missing run exits 4" 4 scripts/conductor-apply.sh r-missing
printf '{"policyVersion":9}' > "$TARGET/conductor-policy.json"
check_rc "invalid policy exits 2" 2 scripts/conductor-apply.sh "$RUN_OK"
git -C "$TARGET" checkout -q -- conductor-policy.json
echo
echo "selftest: $PASS passed, $FAIL failed"
[ "$FAIL" -eq 0 ]
+17 -11
View File
@@ -10,10 +10,16 @@ SANDBOX="$(mktemp -d)"
trap 'rm -rf "$SANDBOX"' EXIT
PASS=0
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
FAIL=0
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
}
# expect_exit NAME EXPECTED_RC -- command...
@@ -25,10 +31,10 @@ expect_exit() {
rc=$?
if [ "$rc" -eq "$expected" ]; then
PASS=$((PASS + 1))
echo "ok $name (exit $rc)"
echo "${C_OK}OK${C_RESET} $name (exit $rc)"
else
FAIL=$((FAIL + 1))
echo "FAIL $name (exit $rc, expected $expected)"
echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
fi
}
@@ -62,8 +68,8 @@ check "env exports adapter" $?
rm -f "$SANDBOX/config.json"
expect_exit "bootstrap creates default when absent" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/config.json" $CONFIG_OP bootstrap
[ -f "$SANDBOX/config.json" ] && { PASS=$((PASS+1)); echo "ok bootstrap wrote config file"; } \
|| { FAIL=$((FAIL+1)); echo "FAIL bootstrap wrote config file"; }
[ -f "$SANDBOX/config.json" ] && { PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} bootstrap wrote config file"; } \
|| { FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} bootstrap wrote config file"; }
SUM_BEFORE=$(sha256sum "$SANDBOX/config.json" | cut -d' ' -f1)
MTIME_BEFORE=$(stat -c %Y "$SANDBOX/config.json")
@@ -73,9 +79,9 @@ expect_exit "bootstrap is idempotent on existing config" 0 -- \
SUM_AFTER=$(sha256sum "$SANDBOX/config.json" | cut -d' ' -f1)
MTIME_AFTER=$(stat -c %Y "$SANDBOX/config.json")
if [ "$SUM_BEFORE" = "$SUM_AFTER" ] && [ "$MTIME_BEFORE" = "$MTIME_AFTER" ]; then
PASS=$((PASS+1)); echo "ok bootstrap did not rewrite existing config"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} bootstrap did not rewrite existing config"
else
FAIL=$((FAIL+1)); echo "FAIL bootstrap rewrote existing config"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} bootstrap rewrote existing config"
fi
# --- validate ---
@@ -140,9 +146,9 @@ cfg valid.json "$(valid_body "$DATA_ROOT")"
EVAL_OUT="$(MOSAIC_CONFIG="$SANDBOX/valid.json" $CONFIG_OP env)" || true
if eval "$EVAL_OUT" 2>/dev/null && [ "$MOSAIC_DATA_ROOT" = "$DATA_ROOT" ] \
&& [ "$MOSAIC_PROVIDER" = "zai" ] && [ "$MOSAIC_MODEL" = "glm-5.3-flash" ]; then
PASS=$((PASS+1)); echo "ok env exports resolve correctly"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} env exports resolve correctly"
else
FAIL=$((FAIL+1)); echo "FAIL env exports resolve correctly"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} env exports resolve correctly"
fi
# --- validation must not modify the file ---
@@ -150,9 +156,9 @@ SUM_INVALID_BEFORE=$(sha256sum "$SANDBOX/invalid.json" | cut -d' ' -f1)
MOSAIC_CONFIG="$SANDBOX/invalid.json" $CONFIG_OP validate >/dev/null 2>&1
SUM_INVALID_AFTER=$(sha256sum "$SANDBOX/invalid.json" | cut -d' ' -f1)
if [ "$SUM_INVALID_BEFORE" = "$SUM_INVALID_AFTER" ]; then
PASS=$((PASS+1)); echo "ok failed validation modified nothing"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} failed validation modified nothing"
else
FAIL=$((FAIL+1)); echo "FAIL failed validation modified the file"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} failed validation modified the file"
fi
echo
+9 -3
View File
@@ -15,6 +15,12 @@ cp RELEASE "$RELEASE_BACKUP"
trap 'cp "$RELEASE_BACKUP" RELEASE 2>/dev/null; rm -rf "$SANDBOX" "$RELEASE_BACKUP"' EXIT
PASS=0
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
FAIL=0
expect_exit() {
@@ -24,14 +30,14 @@ expect_exit() {
"$@" >/dev/null 2>&1
rc=$?
if [ "$rc" -eq "$expected" ]; then
PASS=$((PASS+1)); echo "ok $name (exit $rc)"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
else
FAIL=$((FAIL+1)); echo "FAIL $name (exit $rc, expected $expected)"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
fi
}
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
}
# ---------- fast: release identity ----------
+120 -17
View File
@@ -11,6 +11,12 @@ SANDBOX="$(mktemp -d)"
trap 'rm -rf "$SANDBOX"' EXIT
PASS=0
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
FAIL=0
expect_exit() {
@@ -20,14 +26,14 @@ expect_exit() {
"$@" >/dev/null 2>&1
rc=$?
if [ "$rc" -eq "$expected" ]; then
PASS=$((PASS+1)); echo "ok $name (exit $rc)"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $name (exit $rc)"
else
FAIL=$((FAIL+1)); echo "FAIL $name (exit $rc, expected $expected)"
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $name (exit $rc, expected $expected)"
fi
}
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} $1"; else FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} $1"; fi
}
latest_reason() {
@@ -36,10 +42,6 @@ latest_reason() {
[ -n "$latest" ] && node -e 'try{const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.reason??"")}catch{console.log("")}' "$latest/result.json" 2>/dev/null
}
check() {
if [ "$2" = "0" ]; then PASS=$((PASS+1)); echo "ok $1"; else FAIL=$((FAIL+1)); echo "FAIL $1"; fi
}
CONFIG="$SANDBOX/config.json"
DATA_ROOT="$SANDBOX/data"
mkdir -p "$DATA_ROOT"
@@ -102,6 +104,37 @@ $TASK validate "$SANDBOX/ok.json" >/dev/null 2>&1
M2=$(stat -c %Y "$SANDBOX/ok.json")
[ "$M1" = "$M2" ] && check "validation does not modify the task file" 0 || check "validation does not modify the task file" 1
# ---------- retention: prune (deterministic, no Docker) ----------
mkdir -p "$SANDBOX/data"
cat > "$SANDBOX/prune-config.json" <<EOF
{"configVersion":1,"environment":"development","dataRoot":"$SANDBOX/data","execution":{"backend":"docker","provider":"zai","model":"m"}}
EOF
for i in 1 2 3 4 5; do
D="$SANDBOX/data/runs/r-20260903T0100_0${i}Z-suite00$i"
mkdir -p "$D"
printf '{"runVersion":1,"runId":"r-20260903T0100_0%sZ-suite00%s","taskId":"t","status":"succeeded"}' "$i" "$i" > "$D/result.json"
done
mkdir -p "$SANDBOX/data/sessions/sentinel" "$SANDBOX/data/workspaces/sentinel"
expect_exit "prune dry-run exits 0" 0 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune
[ "$(ls "$SANDBOX/data/runs" | grep -c '^r-')" -eq 5 ] \
&& check "dry-run deleted nothing" 0 || check "dry-run deleted nothing" 1
expect_exit "prune --keep=2 --yes removes oldest" 0 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune --keep=2 --yes
[ "$(ls "$SANDBOX/data/runs" | grep -c '^r-')" -eq 2 ] \
&& check "kept exactly 2 newest runs" 0 || check "kept exactly 2 newest runs" 1
NEWEST="r-20260903T0100_05Z-suite005"
[ -d "$SANDBOX/data/runs/$NEWEST" ] \
&& check "newest run kept, oldest pruned" 0 || check "newest run kept, oldest pruned" 1
[ -f "$SANDBOX/data/runs/.pruned.log" ] \
&& [ "$(grep -c 'pruned' "$SANDBOX/data/runs/.pruned.log")" -eq 3 ] \
&& check "append-only receipt written (3 entries)" 0 \
|| check "append-only receipt written (3 entries)" 1
[ -d "$SANDBOX/data/sessions/sentinel" ] && [ -d "$SANDBOX/data/workspaces/sentinel" ] \
&& check "sessions/workspaces untouched by prune" 0 \
|| check "sessions/workspaces untouched by prune" 1
expect_exit "prune with invalid keep exits 4" 4 -- env MOSAIC_CONFIG="$SANDBOX/prune-config.json" node scripts/mosaic-task.mjs prune --keep=0 --yes
# ---------- adapter seam: deterministic mock cases (Docker, no provider) ----------
if docker info >/dev/null 2>&1; then
good_task "$SANDBOX/ok.json"
@@ -142,6 +175,67 @@ EOF
&& grep -q 'Seam directive A.' "$SANDBOX/data/system-prompt.md" \
&& check "mission section injected into generated prompt" 0 \
|| check "mission section injected into generated prompt" 1
# retry lineage + relative mission path resolution
expect_exit "retry of mission run succeeds" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
node scripts/mosaic-task.mjs retry "$(ls -dt "$SANDBOX/data/runs"/r-* | head -1 | xargs basename)"
RT="$(ls -dt "$SANDBOX/data/runs"/r-* | head -1 | xargs basename)"
ORIG="$(ls -dt "$SANDBOX/data/runs"/r-* | sed -n 2p | xargs basename)"
node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));process.exit(r.retriedFrom===process.argv[2]?0:1)' \
"$SANDBOX/data/runs/$RT/result.json" "$ORIG" \
&& check "retriedFrom lineage recorded" 0 || check "retriedFrom lineage recorded" 1
grep -q 'MISSION (runtime)' "$SANDBOX/data/system-prompt.md" \
&& check "mission section present after retry (relative path resolved)" 0 \
|| check "mission section present after retry (relative path resolved)" 1
expect_exit "retry of missing run exits 4" 4 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" node scripts/mosaic-task.mjs retry r-missing
# session fork plumbing (M11): fork source + target dir delivered
printf '{"taskVersion":1,"id":"t-forkplumb","prompt":"x","session":"fork-child","sessionForkFrom":"base"}' > "$SANDBOX/forkplumb.json"
mkdir -p "$SANDBOX/data/sessions/base"
printf '{}' > "$SANDBOX/data/sessions/base/20260903T000000-plumb.jsonl"
expect_exit "fork task runs via mock" 0 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
scripts/run-task.sh run "$SANDBOX/forkplumb.json"
FL="$(ls -dt "$SANDBOX/data/runs"/r-* | head -1)"
grep -q '^MOSAIC_SESSION_FORK=/var/lib/mosaic/sessions/base/20260903T000000-plumb.jsonl$' "$FL/stderr.txt" 2>/dev/null \
&& grep -q '^MOSAIC_SESSION_DIR=/var/lib/mosaic/sessions/fork-child$' "$FL/stderr.txt" 2>/dev/null \
&& check "fork source + target delivered to adapter" 0 \
|| check "fork source + target delivered to adapter" 1
printf '{"taskVersion":1,"id":"t-f2","prompt":"x","sessionForkFrom":"base"}' > "$SANDBOX/noforktarget.json"
expect_exit "fork without session target exits 2" 2 -- $TASK validate "$SANDBOX/noforktarget.json"
printf '{"taskVersion":1,"id":"t-f3","prompt":"x","session":"base","sessionForkFrom":"base"}' > "$SANDBOX/selfork.json"
expect_exit "self-fork exits 2" 2 -- $TASK validate "$SANDBOX/selfork.json"
# capability policy (M9): least-privilege intersection
POL="$SANDBOX/data/workspaces"; mkdir -p "$POL"
pol_run() { # missionTools(ABSENT|json) taskTools(ABSENT|json) -> stderr MOSAIC_TOOLS value
if [ "$1" = "ABSENT" ]; then
printf '{"missionVersion":1,"id":"m-pol","objective":"o"}' > "$SANDBOX/pol-m.json"
else
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":[%s]}}' "$1" > "$SANDBOX/pol-m.json"
fi
if [ "$2" = "ABSENT" ]; then
printf '{"taskVersion":1,"id":"t-pol","prompt":"x","mission":"pol-m.json"}' > "$SANDBOX/pol-t.json"
else
printf '{"taskVersion":1,"id":"t-pol","prompt":"x","mission":"pol-m.json","capabilities":{"tools":[%s]}}' "$2" > "$SANDBOX/pol-t.json"
fi
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" MOSAIC_MOCK_RESPONSE=MOCKED \
scripts/run-task.sh run "$SANDBOX/pol-t.json" >/dev/null 2>&1
grep -o '^MOSAIC_TOOLS=.*' "$(ls -dt "$SANDBOX/data/runs"/r-* | head -1)/stderr.txt" 2>/dev/null
}
[ "$(pol_run '"read","bash"' ABSENT)" = 'MOSAIC_TOOLS=read,bash' ] \
&& check "policy: mission only -> mission tools" 0 || check "policy: mission only -> mission tools" 1
[ "$(pol_run ABSENT '"read","bash"')" = 'MOSAIC_TOOLS=read,bash' ] \
&& check "policy: task only -> task tools" 0 || check "policy: task only -> task tools" 1
[ "$(pol_run '"read"' '"bash","read"')" = 'MOSAIC_TOOLS=read' ] \
&& check "policy: both -> intersection (task narrowed)" 0 || check "policy: both -> intersection (task narrowed)" 1
[ "$(pol_run '"grep"' '"bash","read"')" = 'MOSAIC_TOOLS=' ] \
&& check "policy: empty intersection -> tool-free" 0 || check "policy: empty intersection -> tool-free" 1
printf '{"missionVersion":1,"id":"m-pol","objective":"o","capabilities":{"tools":["sudo"]}}' > "$SANDBOX/pol-m.json"
expect_exit "invalid mission capabilities rejected" 2 -- \
env MOSAIC_CONFIG="$SANDBOX/mock-adapters.json" $TASK validate "$SANDBOX/pol-t.json"
else
echo "skip adapter seam cases (docker daemon unavailable)"
fi
@@ -194,18 +288,12 @@ dump_latest_run() {
fi
}
latest_reason() {
local latest
latest="$(ls -dt "$SANDBOX/data/runs"/r-* 2>/dev/null | head -1)"
[ -n "$latest" ] && node -e 'try{const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));console.log(r.reason??"")}catch{console.log("")}' "$latest/result.json" 2>/dev/null
}
if docker info >/dev/null 2>&1; then
RUNS1=$(ls "$DATA_ROOT/runs" 2>/dev/null | wc -l)
if scripts/run-task.sh run "$SANDBOX/ok.json" >/dev/null 2>&1; then
PASS=$((PASS+1)); echo "ok live hello task succeeds with exact marker"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} live hello task succeeds with exact marker"
else
FAIL=$((FAIL+1)); echo "FAIL live hello task succeeds with exact marker" >&2
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} live hello task succeeds with exact marker" >&2
dump_latest_run
fi
@@ -222,9 +310,9 @@ process.exit(r.status === "succeeded" && r.response === "MOSAIC_HELLO_OK" && r.e
RC=$?
WRONG_REASON="$(latest_reason)"
if [ "$RC" -eq 1 ] && [ "$WRONG_REASON" = "expect-mismatch" ]; then
PASS=$((PASS+1)); echo "ok wrong expectExact fails with exit 1 (reason: expect-mismatch)"
PASS=$((PASS+1)); echo "${C_OK}OK${C_RESET} wrong expectExact fails with exit 1 (reason: expect-mismatch)"
else
FAIL=$((FAIL+1)); echo "FAIL wrong expectExact (exit $RC, reason: '${WRONG_REASON:-none}')" >&2
FAIL=$((FAIL+1)); echo "${C_FAIL}FAIL${C_RESET} wrong expectExact (exit $RC, reason: '${WRONG_REASON:-none}')" >&2
dump_latest_run
fi
@@ -234,6 +322,21 @@ process.exit(r.status === "succeeded" && r.response === "MOSAIC_HELLO_OK" && r.e
COUNT=$($TASK list | wc -l)
[ "$COUNT" -ge 2 ] && check "list shows both runs" 0 || check "list shows both runs" 1
# session fork (M11): teach in base, fork into child, child recalls;
# ancestor file count must be unchanged by the fork
printf '{"taskVersion":1,"id":"t-fork-base","prompt":"Remember this code word for later: mosaico. Reply with exactly: REMEMBERED","session":"suite-base","expectExact":"REMEMBERED","timeoutSeconds":180}' > "$SANDBOX/fb.json"
expect_exit "fork base: teach succeeds" 0 -- scripts/run-task.sh run "$SANDBOX/fb.json"
BASECOUNT=$(ls "$SANDBOX/data/sessions/suite-base" | wc -l)
printf '{"taskVersion":1,"id":"t-fork-child","prompt":"What code word did I ask you to remember? Reply with only the code word.","session":"suite-child","sessionForkFrom":"suite-base","timeoutSeconds":180}' > "$SANDBOX/fc.json"
expect_exit "fork child recalls ancestor context" 0 -- scripts/run-task.sh run "$SANDBOX/fc.json"
FR="$(ls -dt "$DATA_ROOT/runs"/r-* | head -1)"
node -e 'const r=JSON.parse(require("fs").readFileSync(process.argv[1],"utf8"));process.exit(String(r.response).toLowerCase().includes("mosaico")?0:1)' "$FR/result.json" \
&& check "forked child recalled ancestor code word" 0 || check "forked child recalled ancestor code word" 1
[ "$(ls "$SANDBOX/data/sessions/suite-base" | wc -l)" -eq "$BASECOUNT" ] \
&& check "ancestor session untouched by fork" 0 || check "ancestor session untouched by fork" 1
[ "$(ls "$SANDBOX/data/sessions/suite-child" | wc -l)" -ge 1 ] \
&& check "child session dir has its own branch file" 0 || check "child session dir has its own branch file" 1
else
echo "skip live task cases (docker unavailable)"
fi
+9 -2
View File
@@ -16,6 +16,13 @@ source scripts/common.sh
EXPECTED="${EXPECTED_MARKER:-MOSAIC_HELLO_OK}"
# Status colors: terminal-only, NO_COLOR-respecting; plain when piped.
if [ -t 1 ] && [ -z "${NO_COLOR:-}" ]; then
C_OK=$'\033[0;32m'; C_FAIL=$'\033[0;31m'; C_RESET=$'\033[0m'
else
C_OK=""; C_FAIL=""; C_RESET=""
fi
load_config
load_release
IMAGE="$MOSAIC_IMAGE_TAG"
@@ -51,11 +58,11 @@ TRIMMED="$(printf '%s' "$RESPONSE" | sed -e 's/^[[:space:]]*//' -e 's/[[:space:]
# 4-6. Exact comparison gate.
if [ "$TRIMMED" = "$EXPECTED" ]; then
echo "PASS: response matches expected marker"
echo "${C_OK}PASS${C_RESET}: response matches expected marker"
exit 0
fi
echo "FAIL: response does not match expected marker" >&2
echo "${C_FAIL}FAIL${C_RESET}: response does not match expected marker" >&2
printf 'expected: %s\n' "$EXPECTED" >&2
printf 'actual : %s\n' "$TRIMMED" >&2
exit 1
+9
View File
@@ -0,0 +1,9 @@
{
"taskVersion": 1,
"id": "t-worker-retry-refine",
"prompt": "Follow-up on your last change to scripts/mosaic-task.mjs (the retry operation). Conductor review found one defect: runTask builds spawnEnv from process.env, so a direct 'node scripts/mosaic-task.mjs retry <runId>' invocation lacks the launcher-exported variables and compose fails with: required variable MOSAIC_IMAGE_TAG is missing.\n\nFIX (in runTask, not retryRun): make spawnEnv self-sufficient —\n1. spawnEnv.PI_PROVIDER = resolved.execution.provider; spawnEnv.PI_MODEL = resolved.execution.model (from the already-loaded resolved config).\n2. spawnEnv.MOSAIC_IMAGE_TAG = 'mosaic-poc-agent:' + <pinned pi version> + '-r' + <RELEASE file content trimmed>, reading RELEASE and package.json the same way scripts/common.sh load_release does.\nKeep the change minimal; do not touch other files; keep code style.\n\nSELF-CHECK: node --check scripts/mosaic-task.mjs must pass.\nWhen finished, reply with exactly: REFINE_DONE",
"workspace": "stack-repo",
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
"session": "worker-1",
"timeoutSeconds": 600
}
+9
View File
@@ -0,0 +1,9 @@
{
"taskVersion": 1,
"id": "t-worker-retry",
"prompt": "You are working alone inside a git repository checkout (your current working directory). Implement a new feature in scripts/mosaic-task.mjs.\n\nFEATURE: add a 'retry <runId>' operation alongside the existing 'validate | run | show | list' operations.\n\nBEHAVIOR: 'retry <runId>' re-executes a previously recorded run. Steps: (1) validate the runId argument with the same pattern showRun uses; (2) load the config exactly like showRun does; (3) locate the run directory under the runs root; if it does not exist, fail with exit code 4 and the message 'run not found: <runId>'; (4) read that run directory's task.json snapshot; if unreadable, fail with exit code 4; (5) write the snapshot to a temp file and execute it through the EXISTING runTask(path) function so a brand-new run is created — do not duplicate any execution logic; (6) exit with whatever exit code the new run produced.\n\nCONSTRAINTS: reuse runTask; do not refactor unrelated code; do not touch any file other than scripts/mosaic-task.mjs; keep the existing code style; update the default usage error message to include retry.\n\nSELF-CHECK (run it yourself with the bash tool): node --check scripts/mosaic-task.mjs must pass.\n\nWhen finished, reply with exactly: RETRY_DONE",
"workspace": "stack-repo",
"capabilities": { "tools": ["read", "write", "edit", "bash"] },
"session": "worker-1",
"timeoutSeconds": 600
}