Collaborator-authored via conductor-loop calibration: task dispatched to the live ms-test seat (glm-5.3-flash) over agent-send.sh; diff reviewed line by line and every documented flag/exit code independently verified against tool source by the conductor; suites green at integration (config 24 / task 74 / release 14 / conductor 17 + verify). - new 'Tools (host-side)' section: agent-send.sh, agent-watch.sh, unslop-check.js - intro reading guide now points at tools/ (worker-flagged addition, accepted) - Maintenance suite counts corrected: test-task.sh 58 -> 74 Authored-by: ms-test collaborator (glm-5.3-flash) Integrated-by: conductor (dragon-lin:darkwing)
7.1 KiB
TOOLS.md — command and tool reference
On-demand reference for agent sessions (conductors, bootstrapping agents,
reviewers). AGENTS.md routes here; this file carries the depth: usage,
inputs/outputs, exit codes, and safety notes for every entry point.
Reading guide: system entry points are scripts/*.sh (bash) or invoked via
node scripts/mosaic-task.mjs (node). Host-side helpers under tools/
(tmux messaging, watchers, prose checker) are covered under Tools
(host-side) below. Every script fails closed — missing
or invalid configuration/policy refuses the operation with a nonzero exit
and changes nothing.
Lifecycle
| Command | Purpose | Notes |
|---|---|---|
scripts/bootstrap.sh |
Create ~/.config/mosaic-dev/config.json if absent |
Idempotent; existing config validated, never rewritten |
scripts/build.sh |
Build the release image | Tag derived from RELEASE + pinned pi version |
scripts/hello.sh |
One-shot startup request | Prints model response on stdout |
scripts/verify.sh |
Full gated test | Exit 0 only on exact MOSAIC_HELLO_OK; EXPECTED_MARKER overrides for negative drills |
Tasks (missions, runs, evidence)
| Command | Purpose | Notes |
|---|---|---|
scripts/run-task.sh run <task.json> |
Execute a task | Immutable run record under <dataRoot>/runs/ |
scripts/run-task.sh validate <task.json> |
Strict validation | Writes nothing |
node scripts/mosaic-task.mjs show <runId> |
Inspect a run | Full record + snapshots + artifacts |
node scripts/mosaic-task.mjs list |
List runs | task/workspace/session columns |
node scripts/mosaic-task.mjs retry <runId> |
Re-execute a run's snapshot | New run dir; retriedFrom lineage recorded |
node scripts/mosaic-task.mjs prune [--keep=N] [--yes] |
Retention | Dry-run default; receipt in runs/.pruned.log |
Task fields: prompt (required), mission (path), expectExact,
timeoutSeconds (5–600), workspace (:run or named), capabilities.tools
(allowlist: read write edit bash grep find ls), session,
sessionForkFrom (requires session). Mission fields: objective,
directives[], optional governing capabilities.tools. Policy: a task may
narrow a mission's tools, never widen; empty intersection = tool-free run.
Agent (interactive TUI)
scripts/agent.sh <name> [--mission <file>] [--workspace <ws>] [--session <s>] [--tools <list>]
Launches an interactive pi TUI inside the container with the four immutable
contracts + optional mission + agent identity as its system prompt,
persistent named session, optional workspace. Exit with /quit.
Release
| Command | Purpose | Notes |
|---|---|---|
scripts/release.sh package |
Build + tag the release image | Tag: mosaic-poc-agent:<pi>-r<release> |
scripts/release.sh activate |
Health gate → atomic pointer swap | --fault-injection proves the refusal path |
scripts/release.sh rollback |
Health-gated return to previous | Refuses if image missing |
scripts/release.sh status |
Release, tag, active pointer, log | Safe on empty state |
Conductor (worker patches)
scripts/conductor-apply.sh <runId> [--dry-run]
Auto-applies a worker's patch under roles/conductor-policy.json:
succeeded run → clean target tree → path allowlist → syntax gates →
apply → policy suites → attribution commit. Any failure reverts.
Push is never automatic.
Maintenance
| Command | Purpose | Notes |
|---|---|---|
scripts/reset.sh |
Delete the data root | Triple-safety-checked (path, symlink, ownership marker) |
scripts/test-config.sh |
Config selftests (no Docker) | 24 cases |
scripts/test-task.sh |
Task selftests + live cases | 74 cases |
scripts/test-release.sh |
Release selftests | 14 cases |
scripts/test-conductor.sh |
Auto-apply selftests (sandboxed) | 17 cases |
scripts/gitea-api.sh <METHOD> <path> [body] |
Gitea API helper | Token never on argv/stdout |
Tools (host-side)
Host-side helpers under tools/, outside the scripts/ command surface.
Per-tool READMEs: tools/tmux/README.md and tools/unslop-hook/README.md.
| Command | Purpose | Notes |
|---|---|---|
tools/tmux/agent-send.sh |
Inter-agent tmux message with addressing preamble | Reliable submit (bracketed paste, Enter flush, draft detection); ships send-message.sh over ssh for remote panes (remote needs only bash + tmux + base64) |
tools/agent-watch/agent-watch.sh |
Condition watcher per agent seat | One transient systemd --user timer + service per watch; fires agent-send.sh when the condition command exits 0 |
node tools/unslop-hook/unslop-check.js <file> |
Mechanical AI-tell prose check | Dependency-free node CLI + module driven by lists.json; extension.ts is the pi extension wrapper |
agent-send.sh prepends the preamble
[<src_host>:<src_session> -> <dst_host>:<dst_session>]; -C/--class adds
a class=<CLASS> token (terminal-log, actionable, human, reaction,
digest; consumers treat an absent class as actionable). Flags: -s dst
session (required) · -H ssh target for a remote pane · -L named tmux
socket · -n dst hostname for the preamble · -m/-f/stdin message body ·
-S source-label override · -r N Enter-flush attempts (default 2) · -v
verbose · -h help. Exit codes: 0 delivered/queued · 1 target not found ·
2 still draft · 3 usage error · 4 ambiguous socket (the session exists
on more than one tmux server; disambiguate with -L or MOSAIC_TMUX_SOCKET).
agent-watch.sh subcommands: start --name <id> --session <session> --when '<shell command; exit 0 = met>' --message <text> with --class,
--interval (default 30), --timeout (default 3600), --repeat,
--quiet-timeout, --socket · list · status [--json] · stop <name> ·
log <name> · meta-install [--interval 300] [--unit-name <unit>] ·
meta-remove [--unit-name <unit>]. Interval floor is 10s (a watcher is a
fallback cadence, never a tight poll); hidden _tick/_scan subcommands run
inside the systemd services. status exit codes: 0 clean · 3 any stale
watch or dead meta-watch · 6 systemd user bus unreachable. Delivery goes
through agent-send.sh; rc 2 means the text reached the pane as an
unsubmitted draft, which counts as delivered and is not retried (other
failures retry twice, then the watch gives up). Watches are one-shot by
default; --repeat re-arms. Notices carry a [watch:<name>] prefix.
unslop-check.js checks a file (or stdin) against the word, phrase,
punctuation-density, and pattern lists in lists.json, stripping fenced and
inline code first so a quoted mention never flags. Invocation:
node tools/unslop-hook/unslop-check.js <file>; UNSLOP_LISTS=<path>
overrides the lists location. Exit codes: 0 clean · 1 violations (findings
printed as JSON on stdout) · 2 gate broken (invalid lists or unreadable
input; error on stderr, never a clean verdict).
Exit-code convention
0 success · 1 operation failed · 2 invalid data/configuration ·
3 configuration missing for a read operation · 4 usage/file/environment
problem. Scripts print diagnostics on stderr; model responses (and only
model responses) on stdout.