Compare commits

...
Author SHA1 Message Date
jason.woltjeandClaude Opus 5.5 814426c6eb docs(sessions): Sage lands row 45
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:29:13 -05:00
jason.woltje f345ae851f queue: rev 219, row 45 done (#1527, candidate 5b067a9d landed as 9cdb6d82) 2026-10-09 09:28:24 -05:00
jason.woltjeandClaude Opus 5.5 5c680166df docs(build-log): row 45 landed; Rocko's round 1 and 2 entries and SESSIONS line
Rocko's three BUILD-LOG entries (row 45 started, round 1 ready, round 2
ready) and their round 2 SESSIONS line, plus my landing entry with the
gate counts and the reviewers' non-blocking follow-ups.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:27:54 -05:00
jason.woltjeandClaude Opus 5.5 2be51cf5b1 chore(sage): drop my trackers-boot copy, landed in packages/cli/tests (row 45)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:27:34 -05:00
jason.woltjeandClaude Opus 5.5 9cdb6d82e3 feat(cli): S4 follow-up, refusal backoff and tracker boot (row 45, #1527)
Rocko's round 2 candidate, packet agents/rocko/work/s4-follow-up/
(build.patch c8cec070, candidate manifest 5b067a9d, 8/8 OK).

- Definite DM refusals wait the full 30-minute cap, counted from the
  journal's last refusal, so five refusals span about two hours before
  gave-up (lead decision 73). Unknown outcomes keep doubling.
- Broker close sends at host.mjs:144/150/184/188 pass a callback, which
  closes the EPIPE window both reviewers found in round 1.
- README documents manual recovery for an open decision.
- trackers-boot test, X9, X14, and Darkwing's round 1 notes 1-4.
- The append type check stays out; Rocko's reason holds (both reviewers
  agree).

Reviews: Darkwing approve (comment 26884), Filbert approve (26886).
Landing gate on b13fef4c plus the patch: every node suite and every
scripts/test-*.sh green, test-task 98/0.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:27:34 -05:00
jason.woltjeandClaude Opus 5.5 f4714aa03b docs(sessions): Filbert row 45 round 2 review
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:26:32 -05:00
jason.woltjeandClaude Opus 5.5 9b067be173 docs(review): row 45 round 2 review record, approve (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:26:24 -05:00
jason.woltjeandClaude Opus 5.5 5eb9fa8ae7 queue: rev 218, row 45 round 2 approve by filbert on #1527 (comment 26886, candidate 5b067a9d)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:26:19 -05:00
jason.woltjeandClaude Opus 5.5 b13fef4c28 docs(review): row 45 round 2 review record, approve (darkwing)
Comment 26884 on #1527, candidate 5b067a9d, queue rev 217.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:20:38 -05:00
jason.woltjeandClaude Opus 5.5 354fa9f6bb queue: rev 217, row 45 round 2 approve by darkwing on #1527 (comment 26884, candidate 5b067a9d)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:20:34 -05:00
jason.woltje 19dcc5531d queue: rev 216, row 45 in review round 2 on #1527 (request posted by sage, comment 26882); includes revs 212-214 (rocko: in-progress, in-review round 2, candidate 5b067a9d) and rev 215 (request) 2026-10-09 09:11:22 -05:00
jason.woltjeandClaude Opus 5.5 af1377d7e8 docs(cli): row 45 S4 follow-up round 2 candidate packet (rocko)
build.patch c8cec070…6756, candidate-manifest 5b067a9d…0e0e, base d539d8d2.
Covers Darkwing 26872, Filbert 26880 and lead decision 73.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 09:10:55 -05:00
jason.woltjeandClaude Opus 5.5 d539d8d2cc docs(sessions): Filbert row 45 round 1 review
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:32:43 -05:00
jason.woltjeandClaude Opus 5.5 8bcdd9a6f8 docs: SESSIONS, Dewey row 46 review and Sage network cleanup, row 46 landing, row 45 rulings
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:32:43 -05:00
jason.woltjeandClaude Opus 5.5 46baf3f854 docs(review): row 45 round 1 review record, changes (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:32:38 -05:00
jason.woltjeandClaude Opus 5.5 6e4af4b7c2 queue: rev 211, row 45 round 1 changes (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:32:35 -05:00
jason.woltjeandClaude Opus 5.5 a690f09132 docs(plans): lead decision 73, row 45 round 1 rulings (definite refusals wait the full cap; close send callback)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:31:23 -05:00
jason.woltje 1aad1e09ee queue: rev 210, row 46 done (sage, landed 2d308abd); includes rev 208, row 45 changes from darkwing (comment 26872), and rev 209, row 46 approve from dewey (comment 26873) 2026-10-09 08:30:36 -05:00
jason.woltjeandClaude Opus 5.5 2d308abd16 test(conversation): fake-pi tool child signals ready after its TERM handler (row 46, #1528)
Fixes the K1, K3 and K10 race: a child could get SIGTERM before it
installed its handler. Author: Darkwing. Reviewer: Dewey (approve,
comment 26873). BUILD-LOG landing entry with gate counts and the
correction to decision 72's probe claim.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:30:01 -05:00
jason.woltjeandClaude Opus 5.5 1a481a1f4d docs(cli): row 45 round 1 review, changes on #1527 (darkwing)
R1: the guarded close send at host.mjs:144 and :150 can still fail with
EPIPE after a broker SIGKILL; reproduced, fix is a send callback.
R2: give-up lands 7.5 min after the first refusal, against decision 72's
reason; Sage's ruling.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:29:50 -05:00
jason.woltjeandClaude Opus 5.5 9e172fb494 queue: rev 208, row 45 round 1 changes by darkwing on #1527 (comment 26872)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:29:41 -05:00
jason.woltjeandClaude Opus 5.5 e52493f7ae docs: BUILD-LOG, Docker network removal result (17 removed, 32 to 15)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:26:29 -05:00
jason.woltjeandClaude Opus 5.5 7525431978 docs: BUILD-LOG, unused Docker networks listed before removal (Mos's ruling)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:26:07 -05:00
jason.woltje 112071c74b queue: rev 207, row 45 in review round 1 on #1527 (rocko; request posted by sage, comment 26869) 2026-10-09 08:16:10 -05:00
jason.woltjeandClaude Opus 5.5 156eb07add docs(cli): row 45 S4 follow-up candidate packet (rocko)
Packet for #1527, base 521597bb: build.patch, candidate manifest
b329fdcb, packet manifest, BUILD.md, gate and mutant receipts. The
candidate itself is not committed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:15:47 -05:00
jason.woltjeandClaude Opus 5.5 7478ee2c02 docs(conversation): row 46 cohort K1/K3/K10 diagnosis, receipts and candidate packet (darkwing)
The ignoreTerm tool child was TERMed before Node installed its handler
when the claim store sits on a fast disk. Candidate fixes the fixture;
the patch itself is not committed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:11:48 -05:00
jason.woltjeandClaude Opus 5.5 76506c7e96 queue: rev 203, row 46 in review round 1 on #1528 (darkwing); includes rev 201, row 45 in progress (rocko)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 08:11:43 -05:00
jason.woltjeandClaude Opus 5.5 521597bbe0 queue: rev 200, row 46 in progress (darkwing)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:56:57 -05:00
jason.woltjeandClaude Opus 5.5 41c5a8563f docs: SESSIONS, Sage lands rows 38 and 39 and opens rows 45 and 46
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:55:50 -05:00
jason.woltje d0c20a2f24 queue: rev 199, rows 45 and 46 added and briefed (sage, decision 72) 2026-10-09 07:55:41 -05:00
jason.woltjeandClaude Opus 5.5 ec2a2c6773 docs: lead decision 72, the S4 follow-up and cohort briefs, and the trackers boot test
Decision 72 rules on F2 (five definite refusals stop a DM; unknown
outcomes keep retrying) and J4 (type-check confirmed journal lines),
and opens two rows: the S4 follow-up for Rocko and the K1, K3 and K10
cohort failures for Darkwing. The trackers boot test from the row 39
gate is kept under agents/sage/work/s4-follow-up/ until the follow-up
moves it into packages/cli/tests/.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:54:52 -05:00
jason.woltjeandClaude Opus 5.5 e580ee2239 queue: rev 195, row 39 S4 done (sage, 2f5303c1)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:50:18 -05:00
jason.woltjeandClaude Opus 5.5 a78f64b1cf docs: BUILD-LOG, S4 round 1 and round 2 entries, the M28 correction and the row 39 landing
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:49:52 -05:00
jason.woltjeandClaude Opus 5.5 2f5303c1c7 feat(cli): the mosaic CLI, broker host and decision notifier (row 39, S4, rocko)
packages/cli adds mosaic inbox, decide, tasks, agents and trail over the
human-cli transport, and mosaic bus start, stop and status as the trusted
host (unit mosaic-bus@<business>, scripts/bus-service.sh). The host boots
packages/bus/src/process.mjs, passes config.trackers from the tracker.*
variables (lead decision 70), and runs a notifier child. The notifier DMs
each open blocking decision once and sends an 08:00 America/Chicago
digest, journaled in notify/<business>/sent.jsonl at 0600 with no Discord
ids. A torn journal tail is copied aside and truncated; a malformed line,
a directory looser than 0700 or a symlinked journal refuses (lead
decision 71). packages/discord gains dmRecipient, createDm and notify.mjs.

Candidate agents/rocko/work/slice1-s4, base b9b6cf00, build.patch
b52f7d68, manifest e858504e (29 files). Darkwing approved round 2 on
#1521 (comment 26855), Filbert approved round 2 (comment 26856). The
packet's mutant table lists M28 as killed; it survived, and BUILD-LOG
records the correction.

Integration gate in a worktree on 2557e29d with the patch applied:
bus 67, business 60, cli 49, control-board 124, discord 178, ledger 78,
mosaic 69, queue 148, seat 19, tasks 51 and webui 14, all with no
failures. Conversation is 149/3, the same K1, K3 and K10 cases that fail
on the base; S4 doesn't touch the package. Every scripts/test-*.sh is
green, with test-release 14/14 and test-task 98/98 on the existing gate2
compose network. A scratch test, not in this commit, booted the real
host with trackers against S3's fake Vikunja: the adapter went ready and
a task.close on a missing task answered task-not-found after a Vikunja
read.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:49:38 -05:00
jason.woltjeandClaude Opus 5.5 2557e29dc7 queue: rev 194, row 38 S3 done (sage, 7e73c2cd)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:41:46 -05:00
jason.woltjeandClaude Opus 5.5 ffd3a9d86c docs: BUILD-LOG, S2, S2b, S2c and S3 entries and the row 38 landing
The seats' uncommitted S2 to S3 entries land with row 38. Rocko's two S4
entries stay uncommitted until row 39 lands.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:41:17 -05:00
jason.woltjeandClaude Opus 5.5 7e73c2cd13 feat(tasks): the Vikunja v2 adapter, broker task verbs and sync (row 38, S3, darkwing)
packages/tasks adds the Vikunja v2 client, the eight task verbs, the
board-plus-cursor poll with its 60 s window and the digest. The broker
gains the task verbs and boots trackers from the boot config (lead
decisions 66 to 68). Due dates are truncated to the second and recorded
as truncated (B1). A write that lands but whose final read fails counts
as landed, in update and in create (B2).

Candidate agents/darkwing/work/slice1-s3, build-r2.patch 71ce87e6,
manifest e10e30e3 (28 files). Filbert approved round 2 on #1520
(comment 26853). Darkwing's post-reset rerun: test-release 14/14,
test-task 98/98 (comment 26857).

Integration gate in a worktree on c4baf779 with the patch applied:
bus 67, business 60, control-board 124, discord 173, ledger 78,
mosaic 69, queue 148, seat 19, tasks 51 and webui 14, all with no
failures. Conversation is 149/3. The three cohort kill cases (K1, K3,
K10) fail the same on the unpatched base, and the patch doesn't touch
the package. Every scripts/test-*.sh is green. test-release 14/14 and
test-task 98/98 ran on the existing gate2 compose network, because the
host's Docker address pools are exhausted. No network was created or
pruned.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:40:48 -05:00
jason.woltjeandClaude Opus 5.5 c4baf77916 docs(slice1): row 38 S3 round 2 gate rerun, test-release 14/14 and test-task 98/98
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-09 07:33:26 -05:00
jason.woltjeandClaude Opus 5.5 f5eb85ea4e docs: sessions, filbert row 39 S4 round 2 review
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 20:16:29 -05:00
jason.woltjeandClaude Opus 5.5 2de56a471c docs(slice1): filbert row 39 S4 round 2 review record (approve)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 20:16:18 -05:00
jason.woltjeandClaude Opus 5.5 1608e3da45 queue: rev 193, row 39 round 2 approve (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 20:16:12 -05:00
jason.woltjeandClaude Opus 5.5 9bdee45517 queue: rev 192, row 39 round 2 darkwing approves (comment 26855)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 20:01:45 -05:00
jason.woltjeandClaude Opus 5.5 3c18473d11 docs(slice1): darkwing row 39 S4 round 2 review record (approve)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 20:01:06 -05:00
jason.woltjeandClaude Opus 5.5 26bd829c1d queue: revs 188-191, row 39 round 2 in review (rocko moves, sage request)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:56:02 -05:00
jason.woltjeandClaude Opus 5.5 fdbb4e34c4 docs: sessions, filbert row 38 S3 round 2 review
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:41:41 -05:00
jason.woltjeandClaude Opus 5.5 d5c95244c8 docs(slice1): filbert row 38 S3 round 2 review record (approve)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:41:32 -05:00
jason.woltjeandClaude Opus 5.5 5d55cfc156 queue: rev 187, row 38 round 2 approve (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:41:27 -05:00
jason.woltjeandClaude Opus 5.5 bb267da534 queue: rev 186, row 38 round 2 in review (darkwing)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:28:10 -05:00
jason.woltjeandClaude Opus 5.5 b144a03d47 docs(slice1): darkwing row 38 S3 round 2 packet
B1 truncates due dates to the second; B2 records what landed when the
final read fails; tests for V9, V21 and S5. Candidate manifest
e10e30e3, 28 files over c9ef7ee4. Mutants 46/63. Live provider cases
in test-release and test-task failed on a zai usage limit (429).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:27:17 -05:00
jason.woltjeandClaude Opus 5.5 c9ef7ee49f queue: rev 184, row 38 back to in-progress for round 2 (sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:06:44 -05:00
jason.woltjeandClaude Opus 5.5 779569a69d docs: sessions, filbert row 38 S3 round 1 review
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:05:29 -05:00
jason.woltjeandClaude Opus 5.5 dad9b0b00a docs(slice1): filbert row 38 S3 round 1 review record (changes)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:05:20 -05:00
jason.woltjeandClaude Opus 5.5 944bd5770d queue: rev 183, row 38 round 1 changes (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 19:05:17 -05:00
jason.woltjeandClaude Opus 5.5 e328503bc0 docs: sessions, row 39 round 1 verdicts and decision 71 (sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:52:36 -05:00
jason.woltjeandClaude Opus 5.5 b9b6cf008a docs: lead decision 71, S4 round 2 scope and the notify journal's torn tail (sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:52:11 -05:00
jason.woltjeandClaude Opus 5.5 49f99f99aa queue: rev 182, row 39 back to in-progress for round 2 (sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:51:15 -05:00
jason.woltjeandClaude Opus 5.5 aba3aa9dad docs: sessions, filbert row 39 S4 round 1 review
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:49:41 -05:00
jason.woltjeandClaude Opus 5.5 9ec1ac309f docs(slice1): filbert row 39 S4 round 1 review record (changes)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:49:31 -05:00
jason.woltjeandClaude Opus 5.5 60882084cb queue: rev 181, row 39 round 1 filbert asks for changes (comment 26848)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:49:24 -05:00
jason.woltjeandClaude Opus 5.5 37cd4b4771 docs: sessions, darkwing row 38 S3 round 1 and row 39 review
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:36:52 -05:00
jason.woltjeandClaude Opus 5.5 481c6a7d34 queue: revs 179-180, row 38 S3 round 1 review request on #1520 (darkwing)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:35:32 -05:00
jason.woltjeandClaude Opus 5.5 4ac133dd8a docs(slice1): row 38 S3 build packet, probes and live rehearsal (darkwing)
Candidate for #1520 against 81339889: packages/tasks and the bus task
verbs. Live run rehearsed on a scratch v2.7.0 container; the estate run
waits for T236's base URL.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:33:50 -05:00
jason.woltjeandClaude Opus 5.5 81339889dc queue: rev 178, row 39 round 1 darkwing approves (comment 26843)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:16:54 -05:00
jason.woltjeandClaude Opus 5.5 848b26d1dd docs: sessions, row 39 S4 round 1 (rocko, sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:10:11 -05:00
jason.woltjeandClaude Opus 5.5 703b892966 queue: rev 177, row 39 follow-up note (sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:10:11 -05:00
jason.woltjeandClaude Opus 5.5 1423725a73 queue: revs 172-176, row 39 S4 round 1 on #1521 (rocko, sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 18:08:24 -05:00
jason.woltjeandClaude Opus 5.5 d68aa20f27 docs: decision 70, the host's boot config for S3 trackers (sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:31:45 -05:00
jason.woltjeandClaude Opus 5.5 020b2c17b3 queue: repin slice 1 brief for rows 35, 38, 40-42; only S4's section changed (sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:30:08 -05:00
jason.woltjeandClaude Opus 5.5 284cfcafa5 queue: row 39 briefed under lead decision 70 (sage)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:28:42 -05:00
jason.woltjeandClaude Opus 5.5 084c3a3cee docs: lead decision 70, S4 rulings for Rocko (sage)
human-cli.mjs is the decide transport. S4 owns the broker host and one
notifier for the blocking-decision DM and the 08:00 Central digest.
Darkwing is second reviewer for packages/discord and the host. Trail
prints in broker order. Row 39 no longer waits on S3. SESSIONS gets
Rocko's orientation line and a correction for Sage's estimated stamps.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:26:57 -05:00
jason.woltjeandClaude Opus 5.5 bee89d107d docs: lead decision 69, Rocko moves to a Claude thread under R26 (sage)
Jason's R26 via Mos: Claude first, Codex only on gpt-6.1-sol, never
Astra. Rocko's Codex thread ran gpt-6-astra; the seat now runs as T3
thread b84bb264 on Opus 5.5, because its only row, S4, builds mosaic
decide and the gated decision DM. Roster updated. Filbert's launcher
still pins Astra; recorded as a follow-up, unused while Filbert runs in T3.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:21:03 -05:00
jason.woltjeandClaude Opus 5.5 7780ef1e54 docs: runbook token listing reads the items envelope (sage)
Darkwing's correction, checked against the 2.7.0 OpenAPI: GET /tokens
returns PaginatedAPIToken, with items possibly null.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:17:18 -05:00
jason.woltjeandClaude Opus 5.5 67d29f7da5 docs: runbook section 5 lists a bot's token ids as svc (sage)
Darkwing's probe-s7f: svc-$BIZ revokes one bot token (204, then 401);
the owner gets 403. A plain GET /tokens shows only svc's own tokens, so
rotation names GET /tokens?owner_id=<bot id> for an old id. Decision 68
records the resolution.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:16:51 -05:00
jason.woltjeandClaude Opus 5.5 28fbdec97d docs: decision 68 correction, Jason creates the svc account (sage)
Mos: the svc-mosaic-stack password is a credential Jason keeps (R5), and
T236 creates no users. The runbook's Path A now has Jason create it with
the command from the instance's ops doc. Records the R20 confirmation.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:14:45 -05:00
jason.woltjeandClaude Opus 5.5 bc33faa38f docs: lead decision 68, a service account owns the Vikunja bots (sage)
Darkwing's row 38 labels probe on the pinned 2.7.0 image: a bot created
from the owner's account reads and attaches every label that account
created, in any project (upstream #3592). Bots owned by svc-<business>,
an account with no labels and no shares, see only labels on the shared
project's tasks. The runbook now creates the bots and mints and revokes
their tokens as svc-$BIZ; the owner keeps the project and the shares.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 17:13:51 -05:00
jason.woltje d228b03298 queue: row 38 S3 in progress (darkwing, LD67) 2026-10-08 17:02:37 -05:00
jason.woltje 341da4932c queue: revs 158-159, row 38 S3 after 36 and 37 only, briefed (sage, LD67) 2026-10-08 16:57:41 -05:00
jason.woltjeandClaude Opus 5.5 af4d3c0f6a docs(plans): lead decision 67, S3 starts before the Vikunja runbook half (sage)
The estate Vikunja (T236) won't serve by Jason's 10-09 run. S3 builds
against the recorded fake and a scratch container; its gate still needs
the live run. R20's PM and Lead match the slice 1 brief's staffing.
Merge plan step 2 no longer holds for T235 (brief c733ac92b).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 16:56:56 -05:00
jason.woltjeandClaude Opus 5.5 ba51c8d867 docs(plans): refactor-to-next merge plan and lead decision 66 (sage)
Jason's 2026-10-08 rulings via Mos (thread 1eba59e5). R10: next stays
the trunk; the plan pins refactor b8db39cf and next 2101c9b4, finds a
clean merge-tree, and names the CI gate, approval identity, frozen
next-lane publishing, root .mosaic/repo.json, 13 stale v1 PRs and the
hardcoded ledger branch as the work. Nothing merged.

R8: Path A on the new estate Vikunja meets slice 1 and replaces
decision 61's Path B. The runbook's Path A text names the estate
instance and the share-only isolation rule.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-08 16:53:57 -05:00
jason.woltje b8db39cf1e queue: rev 157, row 44 S2c done (sage) 2026-10-05 18:14:18 -05:00
jason.woltjeandClaude Opus 5.5 bb93a05ffd docs(plans): decision 65 correction, message.send classifies by sender role only
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 18:13:46 -05:00
jason.woltjeandClaude Opus 5.5 d27042faf2 feat(bus): a within-role message cites a decision without consuming it (row 44, S2c, rocko)
A within-role message.send stores its decision as a citation: the broker
checks it exists in the business, and doesn't class-match it, consume it
or put it in action.allowed (lead decision 65). Sends that aren't
within-role keep the S2b authority rules. The README wording on
within-role sends is fixed.

Candidate agents/rocko/work/slice1-s2c-r1, build.patch 2c4f8d9f,
manifest aa249ad9. Darkwing approved round 1 on #1526 (comment 26778).
Integration gate in a worktree on ac4a6499: every package green on
Node 26; bus 58/58 on Node 24; every scripts/test-*.sh green. Node 24
failures in conversation, ledger, queue, seat and webui match the base.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 18:13:46 -05:00
jason.woltjeandClaude Opus 5.5 ac4a6499f2 review(bus): darkwing S2c round 1, approve (row 44, #1526)
Verdict comment 26778, queue rev 156. Within-role citations never reach
the authority check or the grant event; other sends keep S2b. 58/58 on
Node 24 and 26, seven of seven mutants killed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 18:06:47 -05:00
jason.woltje 4a8d91edd0 queue: rev 156, row 44 S2c round 1 approve (darkwing) 2026-10-05 18:06:47 -05:00
jason.woltje 2af57513d6 queue: rev 155, row 7 note on the refused 2026-10-05 ledger run (sage) 2026-10-05 18:06:03 -05:00
jason.woltjeandClaude Opus 5.5 e250933de7 docs(records): 2026-10-05 weekly ledger run, refused on a T3 header conflict (row 7, #1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 18:05:19 -05:00
jason.woltje 24a83cf706 queue: revs 150-153, row 44 S2c round 1 on #1526 (rocko, sage) 2026-10-05 18:03:57 -05:00
jason.woltje 83cfd18c94 queue: rev 149, row 44 S2c briefed (sage) 2026-10-05 17:59:20 -05:00
jason.woltje 1f5d253b2e queue: rev 148, row 44 S2c added (sage, #1526) 2026-10-05 17:58:00 -05:00
jason.woltjeandClaude Opus 5.5 dd35a0abcb docs(plans): lead decision 65 and the S2c brief, within-role messages cite decisions
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:57:23 -05:00
jason.woltje 6b753ff7b1 queue: rev 147, row 43 S2b done (sage, 4afab552) 2026-10-05 17:53:18 -05:00
jason.woltjeandClaude Opus 5.5 4afab55254 feat(bus): decision-backed approvals are single-use (row 43, S2b, rocko)
authorize, message.send and role.revoke consume a cited decision once,
inside the write transaction (lead decision 64). A second use refuses
with decision-consumed; a class mismatch refuses with decision-mismatch.

Candidate agents/rocko/work/slice1-s2b-r2, build.patch 56d572fe,
manifest a89d64ca. Darkwing approved round 2 on #1525 (comment 26769).
Integration gate in a worktree on bba75b4a: every package green on
Node 26; bus 55/55 on Node 24; every scripts/test-*.sh green. Node 24
failures in conversation, ledger, queue, seat and webui are identical
on the unpatched base (container /tmp is overlayfs, no jsonschema).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:52:41 -05:00
jason.woltjeandClaude Opus 5.5 bba75b4a4f review(bus): darkwing S2b round 2, approve (row 43, #1525)
Verdict comment 26769, queue rev 146. R1, R2 and both lead decision 64
rulings are in; round 1 probes now refuse. 55/55 on Node 24 and 26,
nine of ten mutants killed, the survivor equivalent.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:43:09 -05:00
jason.woltje 4699419b8b queue: rev 146, row 43 S2b round 2 approve (darkwing) 2026-10-05 17:43:09 -05:00
jason.woltje e7bcf5a988 queue: revs 141-145, row 43 S2b round 2 on #1525 (rocko, sage) 2026-10-05 17:40:25 -05:00
jason.woltjeandClaude Opus 5.5 577380833c docs(plans): lead decision 64, decision-backed approvals are single-use
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:35:45 -05:00
jason.woltjeandClaude Opus 5.5 4f28c09069 review(bus): darkwing S2b round 1, changes (row 43, #1525)
Verdict comment 26765, queue rev 140. R1: message.send and role.revoke
verbs use gated decisions without consuming them. R2: a cross-role
decision authorizes without limit once policy makes the action gated.
Two rulings for Sage on cross-role reuse and class drift.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:35:15 -05:00
jason.woltje 490344d89c queue: rev 140, row 43 S2b round 1 changes (darkwing) 2026-10-05 17:35:09 -05:00
jason.woltje 975a0a7807 queue: revs 135-139, row 43 S2b review round on #1525 (rocko, sage) 2026-10-05 17:23:39 -05:00
jason.woltjeandClaude Opus 5.5 65d78d1227 docs(sessions): sage, S1 and S2 landed
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:16:41 -05:00
jason.woltje 1fe9fc056d queue: row 37 S2 done, row 43 note (sage) 2026-10-05 17:16:40 -05:00
jason.woltjeandClaude Opus 5.5 38828a2cb3 feat(bus): the bus and the broker core (row 37, S2, rocko)
Rocko's round 2 candidate, approved by Darkwing (#1519 comment 26757).
build.patch 40d7e838, manifest 61519059, 24 files under packages/bus,
schema v3b (179ffe35, lead decision 60). Integration gate in a git
worktree of 942dca9e (S1 in the tree) plus the patch: bus 43/43 and
business 60/60 on Node 24 and 26, every package test and every
scripts/test-*.sh green, test-task 98/98 with the live-provider cases.
Rulings from lead decisions 62 and 63: the human proof is cooperative in
slice 1, and a self-raised cross-role decision routes to the human.
Single-use gated approvals follow in row 43.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:15:53 -05:00
jason.woltje 942dca9ea8 queue: row 36 S1 done, row 37 note (sage) 2026-10-05 17:09:58 -05:00
jason.woltjeandClaude Opus 5.5 2d64c71eb2 feat(business): roles v2, business and project files, variable layers (row 36, S1, darkwing)
Darkwing's round 2 candidate, approved by Filbert (#1518 comment 26730).
build-r2.patch a27890d5, manifest 869168c7, 34 files, applied on HEAD and
checked 34/34. Integration gate on an export of HEAD plus the patch:
business 60/60 on Node 24 and 26, every package test and every
scripts/test-*.sh green, test-task 98/98 with the live-provider cases.
Conductor, queue, conversation and discord confirmed in git worktrees of
HEAD with and without the patch, identical results. Lead decision 63
accepts the vocabulary location, the example path and the business
branch.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:09:07 -05:00
jason.woltjeandClaude Opus 5.5 d1e633b220 queue: revs 126-130, row 37 round 1 approve on #1519 (darkwing)
Revs 126-129 are Sage's round opening (request comment 26756). Rev 130
records Darkwing's approval, comment 26757, record 1e3c06c7.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 17:02:05 -05:00
jason.woltje f594318d45 queue: revs 126-129, row 37 S2 review round on #1519 (rocko, sage) 2026-10-05 17:01:08 -05:00
jason.woltjeandClaude Opus 5.5 e16faafdc2 docs(sessions): sage, rocko and dewey lines
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 16:45:41 -05:00
jason.woltje d14b2c702e queue: revs 113-125, briefs re-pinned, rows 11 and 34 done, row 43 S2b (sage) 2026-10-05 16:45:30 -05:00
jason.woltjeandClaude Opus 5.5 623d1b6cc4 docs(plans): S2b brief, gated approvals are single-use
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 16:44:10 -05:00
jason.woltjeandClaude Opus 5.5 85e8f97dbf docs(plans): slice 1 brief accepted, S0 rely-on lines in S6; lead decisions 62-63
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-05 16:43:34 -05:00
jason.woltjeandClaude Opus 5.5 1e3c06c78b docs(review): row 37 S2 round 2, approve (darkwing)
Round 2 of #1519: R1 to R5 fixed and reproduced, 43/43 on Node 26 and
Node 24, 22 of 23 mutants killed against a clean baseline. Records a
reparented-CLI path past the human proof for Sage's ruling.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 23:31:49 -05:00
jason.woltjeandClaude Opus 5.5 79699c143b docs(review): row 37 S2 round 1, changes (darkwing)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 23:13:35 -05:00
jason.woltjeandClaude Opus 5.5 dc0c41c426 docs(sessions): filbert row 36 round 2 approve
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 23:10:31 -05:00
jason.woltjeandClaude Opus 5.5 17481148e2 docs(review): row 36 S1 round 2, approve (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 23:10:26 -05:00
jason.woltjeandClaude Opus 5.5 473155d463 queue: rev 112, row 36 S1 round 2 approve on #1518 (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 23:10:21 -05:00
jason.woltjeandClaude Opus 5.5 95c65b5d06 docs(review): row 34 S0 round 2, approve (darkwing)
Record for #1516 comment 26729: a and b fixed, c to f fixed, ten cc-7
cases reproduce 10/10.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 23:02:20 -05:00
jason.woltje 2b9af160f2 queue: row 34 S0 round 2 approved on #1516 (darkwing) 2026-10-04 23:02:14 -05:00
jason.woltje 768243ccff queue: revs 108-110, row 36 S1 round 2 on #1518 (darkwing) 2026-10-04 22:58:34 -05:00
jason.woltjeandClaude Opus 5.5 8da2ea4b79 docs(sessions): filbert row 34 round 2
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:56:57 -05:00
jason.woltjeandClaude Opus 5.5 db98d12e42 queue: revs 105-107, row 34 S0 round 2 in review on #1516 (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:56:44 -05:00
jason.woltjeandClaude Opus 5.5 d9568cec98 docs(slice1): row 36 S1 round 2 packet (darkwing)
Answers Filbert's round 1 (#1518 comment 26724): launch.by must hold
role.launch, launch is null when limits.authority drops it, a test for
cross-role narrowing, and the business file checked on its open
descriptor.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:55:37 -05:00
jason.woltjeandClaude Opus 5.5 4b7405b86d probes(slice1): row 34 S0 round 2, wrapped command hooks (filbert)
Adds ten cc-7 cases (Darkwing's six wrapper cases plus allow, block,
SIGTERM-ignoring gate and inner timeout above the hook timeout), amends
rely-on lines 2 to 4, removes the stale cc-1d-block-bare evidence, adds
compare.sh, reads SDK_ENTRY from the environment, fixes the cc-5c row and
the Pi exit-code sentence.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:54:10 -05:00
jason.woltjeandClaude Opus 5.5 1abebe5321 docs(review): row 34 S0 round 1, changes (darkwing)
Record for #1516 comment 26725: 34/34 cases reproduce, and a wrapped
Claude command hook fails closed on crash, missing path, noexec and hang.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:51:00 -05:00
jason.woltje 75f6358d0e queue: rev 104, row 34 S0 round 1 changes on #1516 (darkwing) 2026-10-04 22:50:54 -05:00
jason.woltjeandClaude Opus 5.5 7cbf78dab9 docs(review): row 36 S1 round 1, changes (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:40:19 -05:00
jason.woltjeandClaude Opus 5.5 779b01bd30 queue: rev 103, row 36 S1 round 1 changes on #1518 (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:40:12 -05:00
jason.woltjeandClaude Opus 5.5 e250750605 docs(prd): version 1.0, approved by Jason; lead decision 61
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:37:57 -05:00
jason.woltje e3b767155c queue: row 35 note, Path B ruled (sage) 2026-10-04 22:37:57 -05:00
jason.woltjeandClaude Opus 5.5 24c81a6142 docs(plans): lead decision 60, schema v3b accepted
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:34:37 -05:00
jason.woltje 6050d3da54 queue: row 37 note, S2 pins schema v3b (sage) 2026-10-04 22:34:37 -05:00
jason.woltjeandClaude Opus 5.5 48d76de7c6 docs(slice1): schema v3b, current snapshot view (darkwing)
Lead decision 59. task_current picks each task's current snapshot,
skipping a poll row whose read_at is at or before the latest self row's
at; tasks_open reads it. Prototype on Node 24 and 26, two mutants
caught, and Dewey's fixture replayed: tasks_open now differs from v3a
only on #42 (in-progress). The notes add an S3 poller rule: compare a
read against task_current, not the raw latest row.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:33:09 -05:00
jason.woltjeandClaude Opus 5.5 554b03f565 docs(plans): DEFERRED, Node 24 test dir form
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:31:20 -05:00
jason.woltjeandClaude Opus 5.5 c0e2b99afe docs(plans): DEFERRED, unguarded live call in test-task.sh
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:30:51 -05:00
jason.woltjeandClaude Opus 5.5 47742e0a97 queue: revs 99-100, row 36 S1 in review on #1518 (darkwing)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:30:18 -05:00
jason.woltjeandClaude Opus 5.5 b851f83c39 docs(slice1): row 36 S1 build packet, v3a trigger count correction (darkwing)
S1 candidate for Filbert's review on #1518: build.md, build.patch and
build-manifest.sha256 (34 files at base fef4b362). The candidate itself
is not committed.

proto-v3a-notes-correction: the schema has 36 triggers; the printed 35 is
counted after the tamper check's DROP TRIGGER. Found by Rocko.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:28:30 -05:00
jason.woltjeandClaude Opus 5.5 e153c3a3b2 docs(plans): lead decision 59, current snapshot view
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:28:17 -05:00
jason.woltjeandClaude Opus 5.5 b0a5dd8002 docs(sessions): row 34 S0 in review (filbert)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:27:03 -05:00
jason.woltje ec65dadf51 queue: revs 97-98, row 34 S0 in review on #1516 (filbert) 2026-10-04 22:26:44 -05:00
jason.woltjeandClaude Opus 5.5 771fc3d270 docs(slice1): row 34 S0 harness probe matrix (filbert)
Pi 0.85.1 and Claude Code 2.1.289 (Agent SDK 0.3.289), 34 cases against a
scripted local model in a loopback-only network namespace. MATRIX.md has
the command, outcome, model-visible message and fail-closed column per
case, and one rely-on line per case. Evidence under evidence/.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:25:41 -05:00
jason.woltje fef4b362fc queue: row 35 SR waiting on Jason 2026-10-04 22:17:11 -05:00
jason.woltjeandClaude Opus 5.5 ccbf81b1a8 docs(sessions): row 35 SR approved
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:16:27 -05:00
jason.woltjeandClaude Opus 5.5 8619d36bc4 docs(slice1): row 35 SR round 3 review record (darkwing)
Verdict approve, #1517 comment 26715.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:16:10 -05:00
jason.woltjeandClaude Opus 5.5 e7c6d9f8e0 queue: row 35 SR round 3 approve (darkwing, comment 26715)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:16:10 -05:00
jason.woltje ee628a4072 queue: row 35 SR round 3 requested 2026-10-04 22:14:55 -05:00
jason.woltjeandClaude Opus 5.5 98814a67d2 docs(guides): slice 1 identities round 3, revoke before rm
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:14:10 -05:00
jason.woltjeandClaude Opus 5.5 ee82aa5e4f docs(slice1): row 35 SR round 2 review record (darkwing)
Verdict changes, #1517 comment 26712: the Vikunja rotation removes the
owner header before the revoke that needs it.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:13:33 -05:00
jason.woltjeandClaude Opus 5.5 bc81699c58 queue: row 35 SR round 2 changes (darkwing, comment 26712)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:13:33 -05:00
jason.woltje 6352af080c queue: row 35 SR round 2 requested 2026-10-04 22:11:31 -05:00
jason.woltjeandClaude Opus 5.5 104cf4f3f0 docs(guides): slice 1 identities round 2; lead decision 58
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:10:32 -05:00
jason.woltjeandClaude Opus 5.5 4d51c1304d docs(slice1): row 35 SR round 1 review record (darkwing)
Verdict changes, #1517 comment 26709. Gitea 1.27.1 and Vikunja 2.7.0
source checks for Sage's three questions; two mint defects.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:09:21 -05:00
jason.woltjeandClaude Opus 5.5 fe537f68a1 queue: rev 86, row 35 SR round 1 changes (darkwing, comment 26709)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:09:21 -05:00
jason.woltje d0e1bee751 queue: rev 85, row 37 note, S2 builds from v3a 2026-10-04 22:03:36 -05:00
jason.woltjeandClaude Opus 5.5 c9ff79c802 docs(plans): lead decision 57, schema v3a accepted
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:03:02 -05:00
jason.woltjeandClaude Opus 5.5 29daa48216 docs(slice1): schema v3a, task event rules (lead decision 56, Q3)
Two triggers: every task.* event names its task in subject, and
task.created cites a human.input event and a requirement id. The task
ref pattern is tightened on snapshots, decisions and subjects. Runs on
Node 24.21.0 and 26.8.1 match apart from the version line.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 22:02:31 -05:00
jason.woltje 5e8f0f4749 queue: revs 80-83, row 35 SR in review on #1517 2026-10-04 21:59:50 -05:00
jason.woltjeandClaude Opus 5.5 c9c1699a05 docs(guides): slice 1 identities runbook (row SR); lead decision 56
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:58:49 -05:00
jason.woltjeandClaude Opus 5.5 8c4460f35e chore(queue): row 34 S0 probe matrix started by filbert
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:58:22 -05:00
jason.woltje 6ecbf00b94 queue: row 36 S1 started by darkwing 2026-10-04 21:57:07 -05:00
jason.woltjeandClaude Opus 5.5 b02aade221 docs(rocko): slice 1 S2 preparation packet (record)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:56:11 -05:00
jason.woltjeandClaude Opus 5.5 06cf0d33d6 docs(plans): lead decision 55, schema v3 accepted
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:55:50 -05:00
jason.woltjeandClaude Opus 5.5 7ed831781b docs(slice1): schema v3 prototype, addendum B section 5
task_snapshots gains via and read_at, an integer-bucket CHECK with a
tombstone exception, the read-ordering rule in task_external_changes and
the tasks_open view. task.missing names its task and a reason. Prototype
rerun on Node 24.21.0 and 26.8.1; a mutant without the read_at rule
reports the broker's own move as external. Lead decision 52 (B3).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:55:17 -05:00
jason.woltje 05a1cbf85a queue: revs 64-76, slice 1 rows 34-42, rows 34-37 briefed 2026-10-04 21:53:00 -05:00
jason.woltjeandClaude Opus 5.5 4837e1ea29 docs(plans): lead decision 54, slice 1 rows and owners
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:52:27 -05:00
jason.woltjeandClaude Opus 5.5 43c48d7a23 docs(plans): slice 1 brief, rows S0 to S7
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:50:59 -05:00
jason.woltjeandClaude Opus 5.5 dbfaad72c3 queue: rev 63, row 5 note, Gate E rides the slice 1 WebUI step
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:46:35 -05:00
jason.woltjeandClaude Opus 5.5 a59a976dda docs(prd): PRD draft 0.4 with PRDY round 3; lead decision 53
Tokens by hand from a runbook, Gitea bot users per role, slice 1 agents
as Jason's OS user, DMs through the existing connector, digest at 08:00
Central, Gate E with the slice 1 WebUI step. Approval as 1.0 is still
open.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 21:46:00 -05:00
jason.woltjeandClaude Opus 5.5 b6631e6b4f docs(plans): slice 1 addendum B; lead decision 52
Real Vikunja scope map, a read-only sync bot, a two-request poll that
sees column moves and deletions, and task_snapshots changes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 18:42:31 -05:00
jason.woltjeandClaude Opus 5.5 5175ca10f2 docs(plans): Vikunja probe record; lead decision 51
Researcher ran P1-P7 on a scratch Vikunja 2.7.0. Scope names in
addendum A are wrong, polling on updated misses column moves and
deletions, and If-Match isn't enforced. Decision 51 rules on each.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 18:31:47 -05:00
jason.woltjeandClaude Opus 5.5 1595553ad9 queue: revs 61-62, row 5 CHAT-03 I1 committed 243e153c, waiting on Jason for Gate E
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 15:48:48 -05:00
jason.woltjeandClaude Opus 5.5 243e153c8b feat(conversation): CHAT-03 I1, mediated control of a sealed headless Pi (#1507)
Controller, claim store, live-session guard, engine link and seal,
turn tracker, cohort force stop and recovery, client library,
transcript and mediated terminal, with the fake engine and tests.
Fixtures only; no live cutover.

Dewey built it. Darkwing (comment 26690) and Filbert (comment 26694)
approved round 2. Manifest I1-r2-manifest.sha256 (2b48e333, 27 files).
Suites on an export: conversation 152/152, control-board 124, webui 14,
seat 19, chat-00/01/01c checks, and all nine scripts/test-*.sh green.
Follow-ups for I3 are in DEFERRED. Gate E stays with Jason.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 15:47:53 -05:00
jason.woltjeandClaude Opus 5.5 ddd9cf3635 chore(queue): row 5 round 2 approved by filbert, comment 26694 on #1507
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 15:42:05 -05:00
jason.woltjeandClaude Opus 5.5 b846c44dba docs(review): Darkwing's CHAT-03 I1 round 2 review record
Approve, no blockers; follow-ups F1-F3 for I3. Comment 26690 on #1507.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 15:33:21 -05:00
jason.woltjeandClaude Opus 5.5 2cd6aad90f queue: rev 59, row 5 CHAT-03 I1 round 2 approved (Darkwing)
Verdict comment 26690 on #1507, candidate 2b48e333. B1 and B2 fixed;
follow-ups F1-F3 recorded in the review file, none blocking.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 15:33:02 -05:00
jason.woltjeandClaude Opus 5.5 5e7fa4c5ea docs: Dewey's CHAT-03 I1 build and round-1 rework log entries
BUILD-LOG phases for the I1 build and the round-1 rework, and Dewey's
SESSIONS line. The candidate itself stays uncommitted until both
reviewers approve round 2.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 15:27:08 -05:00
jason.woltjeandClaude Opus 5.5 665866f158 queue: revs 56-58, row 5 CHAT-03 I1 round 2 requested (Dewey)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 15:27:08 -05:00
jason.woltjeandClaude Opus 5.5 f2b8c92a84 docs(plans): slice 1 prototype v2 records; lead decision 50
Darkwing's schema-v2 adds task_snapshots, decisions.blocking and a
closed events.kind list. Sage reran proto-v2 on Node 26 with matching
output. Decision 50 accepts the five choices beyond addendum A.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 14:59:33 -05:00
jason.woltjeandClaude Opus 5.5 d7723e2237 docs(prd): PRD draft 0.3; slice 1 addendum A; lead decision 49
Darkwing's addendum as a record. Broker verbs enforce field ownership,
broker push from a bare repo, polling without webhooks, Vikunja probes
before the adapter interface is fixed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 14:56:14 -05:00
jason.woltjeandClaude Opus 5.5 c57998772d docs(prd): PRD draft 0.2 with requirement ids; lead decision 48
PRDY round 2 answers, and Researcher's Vikunja and Pocket ID report as
a record.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 14:48:15 -05:00
jason.woltjeandClaude Opus 5.5 df9e036214 docs(plans): meta-harness survey; lead decision 47
Filbert's survey of Pi, Claude Code and Codex hooks as a record. No hook
is a hard block alone; credential scope, verbs, tool ceiling and a
container are. Two DEFERRED items.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 14:26:50 -05:00
jason.woltjeandClaude Opus 5.5 bb7e37dda2 docs(plans): slice 1 data model note and prototype; lead decision 46
Darkwing's design note and SQLite prototype as records. Sage accepts
seven of eight open questions; the PM launching sessions goes to Jason.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 14:19:14 -05:00
jason.woltjeandClaude Opus 5.5 9c69f2fba3 docs(prd): Mosaic Stack PRD draft 0.1 from PRDY round 1; lead decision 45
Round 1 answers fill users, problems, hard limits, hosting and the v1
done measure. Decision 45 raises Sage's seat limit to 4 Opus 5.5 and 4
Sonnet 5.5 sessions.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 14:12:38 -05:00
jason.woltjeandClaude Opus 5.5 8a3e672f4a docs(records): row 5 round 1 review files, record-issue gap in DEFERRED
Both reviewers requested changes (26681, 26683). Darkwing's verdict
landed on #1508; pointer 26685 on #1507.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 14:01:37 -05:00
jason.woltjeandClaude Opus 5.5 14562ac227 docs(plans): lead decision 44, foundation direction ratified
Slice 1 becomes goal 1, working north star sentence on the goals page,
Vikunja local by default with existing or bundled at install, SQLite
for decisions and messages, Pocket ID after slice 1, row 32 parked.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 14:00:58 -05:00
jason.woltje 3270cc4ea9 chore(queue): row 32 parked on Jason's word (rev 55); carries filbert's row 5 r1 changes verdict (rev 54, comment 26683) 2026-10-04 14:00:57 -05:00
jason.woltjeandClaude Opus 5.5 f0b13a2feb chore(queue): row 5 round 1 changes requested by filbert, comment 26683 on #1507
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 13:59:59 -05:00
jason.woltjeandClaude Opus 5.5 2f4b947906 docs(plans): lead decision 43, Jason's first foundation rulings
Vikunja integrate, first roles agreed, Pocket ID recorded as an SSO
candidate, row 32 parking advised.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 13:54:27 -05:00
jason.woltjeandClaude Opus 5.5 0e45ccd586 docs(sessions): sage routes row 5 round 1 rework to dewey
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 13:52:41 -05:00
jason.woltjeandClaude Opus 5.5 5a5013c42e chore(queue): row 5 round 1 changes requested by darkwing, comment 26681 on #1508
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 13:51:41 -05:00
jason.woltjeandClaude Opus 5.5 8bcf079a58 docs(plans): foundation direction proposal after Jason's KPI message
Answers the five questions, maps the brain dump against the repo, and
proposes one end-to-end slice before measures. Not ratified.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 13:45:43 -05:00
jason.woltjeandClaude Opus 5.5 1c72495815 docs(records): row 33 landed, cold-start failure reopened (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 01:26:07 -05:00
jason.woltje 93eba94939 chore(queue): row 33 done, candidate committed (#1508) 2026-10-04 01:25:41 -05:00
jason.woltjeandClaude Opus 5.5 6fb50cc0a8 fix(queue): genesis owner-not-reviewer check, assign wording, queue-commit HEAD-moved message, calendar dates (row 33, #1508)
Built by filbert, approved by darkwing in round 1 (#1508 comment 26651).
Manifest fe7da3ef, 10 paths. Suites green on an index export.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 01:25:06 -05:00
jason.woltjeandClaude Opus 5.5 862f865a3c chore(queue): row 33 round 1 approved by darkwing, comment 26651 on #1508
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 01:10:04 -05:00
jason.woltjeandClaude Opus 5.5 930d27579a chore(queue): row 33 in-review, round 1 requested from darkwing, comment 26648 on #1508
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 00:59:34 -05:00
jason.woltje 3e39a26a5f chore(queue): row 5 note, CHAT-03 resumed (#1507) 2026-10-04 00:28:16 -05:00
jason.woltjeandClaude Opus 5.5 b48fc0f6cb docs(records): add --issue sets closes, session line (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 00:27:06 -05:00
jason.woltje e048e665d1 chore(queue): row 33 closes narrowed to none (#1508) 2026-10-04 00:26:26 -05:00
jason.woltje 48e9c147ab chore(queue): row 33 queue follow-ups briefed for filbert, row 5 reviewers darkwing and filbert (#1508) 2026-10-04 00:25:44 -05:00
jason.woltjeandClaude Opus 5.5 ec1debd5a9 docs(plans): brief, queue follow-ups from rows 12, 13 and 31 (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 00:25:04 -05:00
jason.woltje 41f373225c chore(queue): row 32, ledger guideposts, queued for Jason's acceptance (#1514) 2026-10-04 00:23:50 -05:00
jason.woltjeandClaude Opus 5.5 07dd203072 docs(plans): brief, ledger guideposts beyond messages per closed issue (#1514)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 00:22:56 -05:00
jason.woltjeandClaude Opus 5.5 921063d5d8 docs(records): Gate G pass and row 31 landed, follow-ups (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 00:17:43 -05:00
jason.woltje 7584fbd055 chore(queue): row 31 done, candidate committed as bb2fa970 (#1508) 2026-10-04 00:17:21 -05:00
jason.woltjeandClaude Opus 5.5 bb2fa9704c feat(queue): refuse a row's owner as its reviewer at add, set reviewers and assign (row 31, #1508)
Built by rocko (Gate G run), approved by filbert in round 1 (#1508 comment
26629). Manifest 35199daf. Suites green on an index export.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 00:16:41 -05:00
jason.woltjeandClaude Opus 5.5 1649f5e2a3 chore(queue): row 31 round 1 approved by filbert, comment 26629 on #1508
Filbert posted the verdict as comment 26629 with the seat's own token and
recorded it. The candidate is manifest 35199daf, uncommitted in the working
tree; no file under docs/plans/reviews/ was added.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 00:12:55 -05:00
jason.woltjeandClaude Opus 5.5 b800a2e60b docs(records): lead decision 42, Jason's unlock rulings (#1508, #1510)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-04 00:05:56 -05:00
jason.woltje 29f6db6290 chore(queue): Gate G pass, rows 9, 10, 16 done, row 31 request reposted (#1508, #1510) 2026-10-04 00:05:52 -05:00
jason.woltjeandClaude Opus 5.5 2eca954163 docs(plans): project relaunch with roles, idea and SetSpark evidence (goal 4 input)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-03 23:51:56 -05:00
jason.woltjeandClaude Opus 5.5 7a3c364b41 docs(records): lead decision 41, Gate G setup (row 31, #1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-28 08:28:00 -05:00
jason.woltjeandClaude Opus 5.5 c989dfe9d4 chore(queue): row 31, owner can't be a reviewer, briefed for rocko (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-28 08:27:44 -05:00
jason.woltjeandClaude Opus 5.5 92aeeeef2d docs(plans): brief, a queue row's owner can't be its reviewer (#1508)
Filbert's n2 from the Piece D round 2 review, written as a queue brief.
It is also the piece a fresh seat starts in Gate G.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-28 08:27:03 -05:00
jason.woltjeandClaude Opus 5.5 0d6658b57d docs(records): 2026-09-28 weekly ledger output (row 7, #1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-28 08:26:12 -05:00
jason.woltjeandClaude Opus 5.5 1310805e5d chore(queue): row 7 note, 2026-09-28 weekly ledger run
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-28 08:26:12 -05:00
jason.woltjeandClaude Opus 5.5 319ee1332c chore(queue): row 13 round 1 approved by filbert, comment 26589 on #1508
Filbert posted the verdict as comment 26589 with the seat's own token and
recorded it (rev 23). The candidate fd72d268 passes verify-commit, and no
file under docs/plans/reviews/ was added.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 11:38:32 -05:00
jason.woltjeandClaude Opus 5.5 ae2cfcb574 chore(queue): row 13 in-review, round 1 request posted as comment 26586 on #1508
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 11:36:23 -05:00
jason.woltjeandClaude Opus 5.5 a91df5ed78 docs(records): Piece E build note and follow-ups (row 13, #1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 11:35:21 -05:00
jason.woltjeandClaude Opus 5.5 0a6f58ddd7 chore(queue): row 7 brief re-pinned to the ledger README weekly routine after Piece E
Piece E (fd72d268) moved the weekly routine into packages/ledger/README.md.
Row 7 stays until Jason closes it (lead decision 40).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 11:35:07 -05:00
jason.woltjeandClaude Opus 5.5 fd72d26899 feat(ledger): Piece E, queue section in the weekly ledger (row 13, #1508)
The ledger prints a queue section above the weekly table. It checks four
things:
- open issues named by done rows;
- owner registrations for active rows;
- closed issues for done rows;
- the age of required rows.
The result is fail, incomplete or reduced pass. It uses its own Gitea
budget of the open list plus at most 10 lookups. A full open page counts
only while an issue in some row's closes has no known state (lead
decision 40). T3 seats are exempt per run with --unsupported-runtime.
The weekly routine is in packages/ledger/README.md.

Built by Darkwing (build.patch ab1f12ca, manifest 0b20bbca). Filbert
reviewed it: round 1 81f26f2e asked for changes (C1, ISO requiredSince
never aged); round 2 ce8ce150 approved. Also carries Filbert's plan
amendment for decision 40 (68a25ffe).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 11:33:44 -05:00
jason.woltjeandClaude Opus 5.5 2333d837e2 docs(records): lead decision 40, Piece E questions (row 13, #1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 11:04:28 -05:00
jason.woltjeandClaude Opus 5.5 f304eaa567 docs(records): row 12 live round and close, queue-commit message follow-up (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 10:14:35 -05:00
jason.woltjeandClaude Opus 5.5 ff61532aa8 chore(queue): row 12 done, Piece D live round approved (comment 26579, #1508)
Round 1 on f539466f: request comment 26577 by darkwing, approval
comment 26579 by filbert, recorded at rev 17. The rev 17 entry went into
f3f48cfd, whose message doesn't mention it. No file was added under
docs/plans/reviews/.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 10:14:25 -05:00
jason.woltjeandClaude Opus 5.5 f3f48cfdc4 chore(queue): row 13 in-progress, Piece E started (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 10:12:52 -05:00
jason.woltjeandClaude Opus 5.5 3c9f4244af chore(queue): row 12 in-review, round 1 request posted as comment 26577 on #1508
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 10:10:54 -05:00
jason.woltjeandClaude Opus 5.5 32ddb2f210 chore(queue): row 12 reviewer filbert for the live round on #1508
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 10:09:35 -05:00
jason.woltjeandClaude Opus 5.5 0550ab21cf docs(records): Piece D build note, TOOLS.md review verbs, DEFERRED updates (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 10:07:51 -05:00
jason.woltjeandClaude Opus 5.5 f539466fcb feat(queue): Piece D, reviews as issue comments, raw per-seat token helper (row 12, #1508)
queue move ID in-review posts the review request as a Gitea comment and
review record reads verdicts back, so reviews stop being files in
docs/plans/reviews/. On a comment round, in-review to waiting-on-jason
now needs every listed reviewer's approval for the current round, the
same as in-review to done (Filbert r1 C1). scripts/gitea-api.sh reads
the raw per-seat token files (lead decisions 37 to 39): config built and
checked before curl starts, export attribute cleared, fixed base URL.
test-queue.sh skips its live checks outside the canonical root.

Darkwing authored. Filbert approved D r2 (cf1d3fd0) after r1 (a2dc2302)
and corrected the plan (293747cd). Rocko reviewed the helper (e896192f,
2096b0a3), and Sage's lead check passed under decision 38. Manifest
b402fb38, 19 files.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-27 10:07:29 -05:00
jason.woltjeandClaude Opus 5.5 cdcedb2741 docs(records): lead decision 39, helper lead check passed, C1 in D (row 12)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 21:45:23 -05:00
jason.woltjeandClaude Opus 5.5 4c53f8c738 docs(records): lead decision 38, Gitea helper round 2 disposition (row 12)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 21:39:10 -05:00
jason.woltjeandClaude Opus 5.5 8efc0ff330 chore(queue): row 5 note, CHAT-03 source author Dewey (lead decision 36)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:25:24 -05:00
jason.woltjeandClaude Opus 5.5 d2813a4b6b docs(records): lead decision 37, raw per-seat token files in the Gitea helper (row 12)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:25:02 -05:00
jason.woltjeandClaude Opus 5.5 8ffbd73b76 docs(records): lead decision 36, row 10 build note, export recipe limit (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:24:24 -05:00
jason.woltjeandClaude Opus 5.5 eae341d3ca chore(queue): row 10 to Jason's gate, row 16 check, header points at the goals review (#1508)
Includes Darkwing's revs 3 to 8 (row 9 to waiting-on-jason, row 12 started).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:23:47 -05:00
jason.woltjeandClaude Opus 5.5 5efe28ab01 docs(agents): seats read the queue, not CURRENT.md (row 10, #1508)
AGENTS.md cadence, pointer and recovery rule name `scripts/mosaic queue
next <seat>` and the ratified goal order. The six seat CONTEXT files run
the queue instead of reading CURRENT.md for ownership and gates. The
queue README says why queue-commit.sh calls cli.mjs directly (Filbert
A2 r1 n3).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:23:20 -05:00
jason.woltjeandClaude Opus 5.5 9e23705724 chore(queue): row 9 to waiting-on-jason for Gate G, review round 1 on A2 6ca116b7 (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:19:07 -05:00
jason.woltjeandClaude Opus 5.5 c8236f7b78 docs(records): queue genesis build note and session line (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:17:18 -05:00
jason.woltjeandClaude Opus 5.5 42f3f2d94f chore(queue): assign row 10 to sage, row 13 gate after rows 9 to 12 (lead decision 35, #1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:17:06 -05:00
jason.woltjeandClaude Opus 5.5 3377b877ba chore(queue): genesis from reviewed map (#1508)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:16:10 -05:00
jason.woltjeandClaude Opus 5.5 6ca116b7ba feat(queue): queue as data A2, migration, render and dispatch (#1508)
Filbert approved round 1 (f167b85e). Manifest 782bcb62, 21 files, plus
the QUEUE.md markers and the TOOLS.md section. Lead decision 35.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 20:14:09 -05:00
jason.woltjeandClaude Opus 5.5 c9539baa0f docs(records): CHAT-03 build note for unexplained aborted, lead decision 34
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:45:26 -05:00
jason.woltjeandClaude Opus 5.5 8a465e891a docs(chat-03): brief pinned at 1ef15ac0 after r3 and scope check, lead decision 33 (#1507)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:45:04 -05:00
jason.woltjeandClaude Opus 5.5 93ee5054f4 docs(plans): north star and goal order ratified by Jason, AGENTS.md pointer
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:40:05 -05:00
jason.woltjeandClaude Opus 5.5 77680e0b22 docs(records): CHAT-03 r3 closed with two rulings, lead decision 32
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:39:00 -05:00
jason.woltjeandClaude Opus 5.5 8dba3ff797 docs(records): CHAT-03 seal loads no explicit extensions, lead decision 31
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:34:06 -05:00
jason.woltjeandClaude Opus 5.5 f0296e7763 docs(records): CHAT-03 rescoped against Gate E, lead decision 30
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:31:38 -05:00
jason.woltjeandClaude Opus 5.5 4bdd3fc3b0 docs(queue): close rows 1, 6, 23, 24, 25 and issues 1503, 1509, 1511, lead decision 29
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:28:13 -05:00
jason.woltjeandClaude Opus 5.5 a11e1153d7 docs(records): fix lead decision reference in DEFERRED SetSpark line
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:24:28 -05:00
jason.woltjeandClaude Opus 5.5 9da0d0895f docs(records): SetSpark approver fix deployed and surveyed, lead decision 28
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:24:20 -05:00
jason.woltjeandClaude Opus 5.5 4d999e2f61 docs(queue): row 5 CHAT-03 rescope hold, row 7 ledger owner and numbers
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:22:49 -05:00
jason.woltjeandClaude Opus 5.5 f2b9e622da docs(plans): goals review with ledger evidence, lead decision 27
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:21:33 -05:00
jason.woltjeandClaude Opus 5.5 4bbacf63ae docs(queue): rows 9 and 11, queue A1 committed, A2 in progress
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:10:39 -05:00
jason.woltjeandClaude Opus 5.5 08bd3a4cc5 fix(tests): clear NODE_TEST_CONTEXT for nested node --test in two suites (#1508)
test-foundation.sh and test-discord.sh ran a nested node --test that would
exit 0 on failure under a parent runner. Both clear the variable now, and
each has a check that fails the suite if it comes back. Darkwing wrote it,
Filbert approved it (e464be6c). Nine suites green on an index export.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:10:31 -05:00
jason.woltjeandClaude Opus 5.5 34a72af912 feat(queue): queue as data A1, journal, lock, CLI and verify (#1508)
packages/queue, scripts/queue-commit.sh, scripts/git-hooks and
scripts/test-queue.sh, plus docs/plans/BRIEF-TEMPLATE.md. There is no
queue.json yet, so verify skips until the genesis commit after A2.

Darkwing built it, and Filbert reviewed R0 (6933b885, changes requested)
and r1 (e464be6c, approved). The 20 files match manifest 85a8a453. The
nine suites passed on an index export, including the new queue suite.
test-queue.sh joins the suite list in AGENTS.md. Lead decisions 20, 23
and 26.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 19:07:48 -05:00
jason.woltjeandClaude Opus 5.5 91df7d6b54 docs(records): CHAT-03 deviation V-1 accepted with limits, lead decision 25
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 18:58:22 -05:00
jason.woltjeandClaude Opus 5.5 f87cd6e201 docs(records): SetSpark approver fix landed in shared-signals cc74d92, lead decision 24
Packet, Rocko's R1 and R2 reviews, DEFERRED outcome and SESSIONS.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 18:52:04 -05:00
jason.woltjeandClaude Opus 5.5 40a02d2bcb docs(records): queue A1 review rulings, lead decision 23
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 18:25:52 -05:00
jason.woltjeandClaude Opus 5.5 556772ceb2 docs(records): CHAT-03 chartered, SetSpark approver owner, lead decisions 20 to 22
Jason approved CHAT-03 and gave Sage the SetSpark approver owner choice.
Decision 20 records the queue A1/A2 split. Decision 21 makes the SetSpark
fix Mosaic's through a Sage subagent and keeps the sanctioned pending:
approver markers so the cutover migration still works. Decision 22 charters
CHAT-03 with Dewey writing the brief. QUEUE row 5, DEFERRED and SESSIONS
updated.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 18:14:02 -05:00
jason.woltjeandClaude Opus 5.5 bc73482045 docs(records): CHAT-02 Console live check passed, WebUI restarted
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 18:07:27 -05:00
jason.woltjeandClaude Opus 5.5 c9e771cf59 feat(webui): CHAT-02 Console, read-only conversation view (#1507)
History opens a seat's conversation from the Waiting card, table row
and inspector. It pages the whole branch through the CHAT-02 board
routes, renders untrusted text inert, polls with the follow cursor, and
marks every switch (branch, newer, reconcile, gone). The WebUI proxy
passes only the two conversation routes' queries upstream.

Dewey authored it. Filbert asked for changes on r1 (24b046af) and
approved r2 (d06de6a7) in review 160dd68d. A relaunch shows 'newer',
not 'reconcile', a deviation from brief 2.3 item 6 that Filbert
accepted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 18:05:11 -05:00
jason.woltjeandClaude Opus 5.5 3a209eeafe fix(ledger): Gate F follow-up, Filbert's notes 1 to 3 (#1506)
Darkwing's follow-up to the T3 thread source: manifest 382f5bb0 pins
t3.mjs, ledger.test.mjs and README.md. Filbert approved it (review
6fd693b6). Ledger 51/51; the eight suites pass on the index.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:57:51 -05:00
jason.woltjeandClaude Opus 5.5 a4d38a3d93 docs(records): row 25 live check passed, SetSpark approver gap has no owner
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:42:17 -05:00
jason.woltjeandClaude Opus 5.5 a68dc1740a docs(records): CHAT-02 build log, Gate F commit, snapshot isolation gap
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:40:23 -05:00
jason.woltjeandClaude Opus 5.5 136958c98b feat(ledger): Gate F, the ledger's T3 thread source (#1506)
packages/ledger/src/t3.mjs reads ~/.t3/userdata/state.sqlite read-only,
in one transaction. It maps each thread to a seat by title and checks
self-addressed headers. Unmatched threads go in a t3:unmapped row. A
missing or locked database exits 1 and names --no-t3. Gate F is on by
default (lead decision 12). The 6a uppercase-class fix rides here.

Separate item: the Pi session reader splits lines only on \n, so a raw
U+2028 or U+2029 in a string no longer splits a record. Node 26.8.1's
readline split there, and the live ledger refused on HEAD.

Darkwing built to brief R3 (f3c05c1b); manifest ba73a163. Filbert
approved the build (e47ec6da) and the U+2028 fix as its own item; brief
review be1aa414. On an index export: the eight suites
24/90/43/17/14/15/63/18, ledger 47/47. Four nonblocking notes go to a
small follow-up.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:39:56 -05:00
jason.woltjeandClaude Opus 5.5 e58783d278 docs(records): CHAT-02 backend commit and board restart
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:37:07 -05:00
jason.woltjeandClaude Opus 5.5 a5beb6d97d feat(conversation): CHAT-02 read-only Pi history reader and two board routes (#1507)
packages/conversation is a library with no server: safe-fs, the Pi session
parser, CHAT-01 pages, pinned snapshots, cursors and follow. The control
board adds GET /api/conversations and /api/conversation behind the Host
and Origin guard. Both are read-only, their queries are validated, and
each refusal code maps to a status.

Dewey authored it (packet 0cf177b1, revision 2). Filbert reviewed the code:
R1 revise (branch ids moving on append, the assumed-link bridge merging
branches, one unreadable seat directory turning the catalogue into a 500),
then R2 approve (3b14d66c). Darkwing reviewed the routes: R1 approve
(07b10ad1), R2 approve (b9d92003). The package lands with the routes,
because serve.mjs imports the reader at load.

On an index export: the eight suites 24/90/43/17/14/15/63/18,
conversation and control-board 153/153, webui 9/9.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:36:20 -05:00
jason.woltjeandClaude Opus 5.5 c5db8c8819 docs(records): row 25 approver fix, second Discord restart, SetSpark approver and docker network gaps
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:31:11 -05:00
jason.woltjeandClaude Opus 5.5 20ea5a0b64 fix(discord): row 25 approvers are user names, never Discord ids in tool text (#1509)
Jason's live check after the 20:58Z restart posted no Approve button. The
Discord Sage wrote DEC-009's required_approvers as names; the SetSpark
service stores approvers as discord:<id> and accepted the names, and the
connector correctly refused the approval request ("bad approver id").

- binding.mjs derives setspark.approvers from the binding's users (name to
  id); a binding-set approvers key and duplicate names are refused. With
  setspark set, a user id or name change refuses the reload (pi's approvers
  are fixed at start).
- setspark.mjs: record_create/record_update map required_approvers names to
  discord:<id> and refuse unknown names, ids, duplicates and non-lists
  before any request, without echoing the value. hideIds turns mentions,
  discord: values and standalone 17-20 digit runs into the user's name or
  "unknown user" in every verb's text and refusal, including the service
  message and code before they are cut. The connector's approval request
  keeps the bare ids.
- tests: boundary test over nested, keyed, numeric, mention and cut ids;
  a local contract fixture from create through validateRequest, with the
  old name-stored shape still refused.

Rocko: R1 revise, R2 revise, R3 approve (81379830..., report da75219f...).
Suites on an index export: 24/90/43/17/14/15/63/18; Discord node tests 173/173.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:30:11 -05:00
jason.woltjeandClaude Opus 5.5 1c5f6bc3a0 docs(queue): queue-as-data plan round 6, Rocko approved (#1508)
Plan 282fabbb (Filbert) and the six adversarial rounds (Rocko, r6 80cde839).
Lead item 15: Gate F first, then A1 and A2 as separate reviewed commits.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:05:16 -05:00
jason.woltjeandClaude Opus 5.5 ffc22c04c6 docs(ledger): Gate F T3 thread source brief R2, Filbert approved (#1506)
Brief e8300cb6 (Darkwing), review bb02d8d3 (Filbert, approve with three
nits), R1 record and R1-to-R2 diff. Lead rulings in item 14.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:04:56 -05:00
jason.woltjeandClaude Opus 5.5 34777c56bc docs(records): row 25 SetSpark fix, Discord restart receipt, Gate F rulings
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:58:59 -05:00
jason.woltjeandClaude Opus 5.5 6c06a6f358 fix(discord): row 25 SetSpark key reaches pi and the connector (#1509)
resolveToolRoots dropped tools.setspark, so the record verbs never reached
pi's --tools list and the connector never built its SetSpark client. The
resolved config now carries only baseUrl, keyFile, principal and timeoutMs,
the keys the extension's loadSetsparkConfig accepts; the extension restores
the response cap. The connector's client uses the validated binding object,
which keeps the cap. README: principal is required, timeoutMs is optional.

Test (failing first): binding -> resolveToolRoots -> JSON -> loadToolsConfig
gives the binding's config, and the verbs reach enabledToolNames. Eight
suites green on an index export. Rocko approved
(agents/rocko/work/row25-setspark-fix-review-2026-09-26.md, a28df89e).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:57:55 -05:00
jason.woltjeandClaude Opus 5.5 84d0965142 docs(deferred): 6b closes the leaked fake pi and concurrent hang entries; wedge-restart and startup-source follow-ups
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:51:05 -05:00
jason.woltjeandClaude Opus 5.5 3edb15eb96 fix(discord): engine tests stop in finally; a turn pi never started stops pi instead of guessing (#1509)
Every engine test stops its engine in finally, and commands() tolerates a
log that doesn't exist yet, so a failed assertion no longer leaks a fake pi
and hangs the six-package union. A turn that failed client-side stays at
the front of the queue and holds the next prompt. If pi has sent no
agent_start ABORT_GRACE_MS (30 s) after the failure, the engine marks
itself wedged, fails held prompts with engine-wedged, refuses new ones with
engine-down, and stops pi. The exit reaches onExit, the connector exits 1,
and the unit restarts it. Pi's events carry no prompt id, so R1's approach,
dropping the turn and sending on, let a late run answer the next prompt.
Rocko rejected R1 and approved R2.

Tests: engine 17/17 (R1 fails 4, HEAD fails 5). Union 408/408 and the eight
suites green on the committed index. Record:
agents/darkwing/work/discord-engine-busy/.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:50:37 -05:00
jason.woltjeandClaude Opus 5.5 f55b94e866 docs(records): board guard live, restart receipt
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:40:56 -05:00
jason.woltjeandClaude Opus 5.5 d1629d610d fix(board): refuse foreign Host and Origin on every control-board route (#1507)
After a DNS rebind, a web page could read /api/board and POST /api/reply,
which pastes into a live seat pane. foreignRequest() now runs first and
returns 403 for a non-loopback Host, a wrong port, userinfo or a path in
Host, or any Origin other than http://<Host>. A missing Origin still passes,
which covers the WebUI proxy. Dewey authored it; Rocko approved de9ff942
(review 5a12f08e) with one low wording finding, now fixed in the notes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:40:34 -05:00
jason.woltjeandClaude Opus 5.5 401cc850bb docs(records): correct walkthrough ruling times from the transcript, note 6b R1 verdict
Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:36:03 -05:00
jason.woltjeandClaude Opus 5.5 993673865a docs(records): Jason's walkthrough rulings, Sage moved to SetSpark, CHAT-02 go
Jason ruled on seven open items (20:27Z-20:45Z): seat Gitea tokens read in
place, one Discord restart after 6b with row 25 live, row 8 limited to the
dev seats, DYOR in dyor-stack-v4 with Sage moved to SetSpark, skills/aws-*
excluded locally, no second WebUI return defect, and go on CHAT-02 only.

Sage persona files name SetSpark as its business work. Darkwing's SOUL drops
harness names that were wrong for T3. DEFERRED adds the slash-prefix paste
hazard and the board Host/Origin gap, and moves the ledger T3 item to Done.
Dewey's approved CHAT-02 brief (636b0fac) and Filbert's review are recorded.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:35:15 -05:00
jason.woltjeandClaude Opus 5.5 ef0020ad85 fix(ledger): count the T3 header as agent, control-board sender as board (#1506)
messageKind knew only the tmux preamble, so a prompt opening with the T3
header [from: role (id) -> to: role (id)] counted as human in Table 2. The
first line now matches either form; anything short of the full header stays
human. Filbert approved R1 against the frozen hashes.

Proof: packages/ledger/tests/ledger.test.mjs, 22/22; against HEAD's
ledger.mjs it fails exactly the two new tests. No suite runs it. Eight
suites green on the staged tree.

The fix changes zero current counts: no Pi log under .pi/state contains a
T3 header, and the ledger does not read T3 transcripts. Gate F waits on a
T3 thread source.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:16:55 -05:00
jason.woltjeandClaude Opus 5.5 d41f81aafe feat(agents): Sage launch files, lead text across seat personas, lead decision record
- agents/sage/ launch files committed after Dewey's review (R1 revise, R2
  approve). The launcher test now covers sage with its zai/glm-5.3 high pin.
  The README states the lead role and its limits, and the seat reads but
  never writes the old DYOR records under ~/.mosaic.
- N6 (Dewey authored, Sage reviewed against pins): seat personas and the
  Rocko launcher name Sage as project lead and Darkwing as a collaborating
  engineering seat, per Jason's 2026-09-26 ruling.
- docs/plans/2026-09-26_lead-decisions.md: the push, no merge into next and
  its conditions, the board restart, queue-as-data rulings, Gate F waiting
  on a T3 source, and what stays with Jason.

Launcher tests 6/6 and 1/1, eight suites green. Not pushed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:11:10 -05:00
jason.woltjeandClaude Opus 5.5 42c08d5285 fix(webui): pending reply notice until the seat answers, relative Age (#1507)
Dewey's return-flow candidate on the #1512 R1 baseline. The inspector used to
show the previous answer while a seat worked on a reply, which looked like the
reply; it now shows a pending notice that clears on the new final answer.
Age shows a relative time beside the ISO time. Filbert approved R2, source
only; the patch reproduces the pinned hashes (app.js d1a51646,
return-flow.test.mjs a598c0d4). webui tests 9/9, all eight suites green.
Known limits are in agents/dewey/work/return-flow-age/NOTES.md. Not pushed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:00:59 -05:00
jason.woltjeandClaude Opus 5.5 0f5b7cb9be docs(records): Sage lead handover, rows 23-25 state, row 8 stub, BUILD-LOG rebuilt on HEAD
Records Jason's 2026-09-26 ruling: Sage leads the project, Darkwing is a
collaborating seat, development stays in T3, and the old ~/.mosaic fleet is
being retired. The lead role adds no push, merge or deploy authority.

- QUEUE rows 23-25 show their commits and pushes; row 8 links a parked stub
  brief listing the rulings Jason must make before fleet seats move.
- Shared records from Darkwing (rows 6, 16, 18, 22, #1511, #1512) and Dewey
  (row 5) that were waiting on one owner for the shared files.
- DEFERRED: T3 headers counted as human in the ledger (#1506), #1509 engine
  test leak and busy gap, #1512 re-run outcome.
- BUILD-LOG.md rebuilt as HEAD plus the uncommitted entries; the working copy
  had dropped the row 23-25 entries. Diff against HEAD is additions only.

All eight suites green. Not pushed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 14:59:24 -05:00
jason.woltjeandClaude Opus 5.5 af4203ca92 feat(board): session attention, Discord rows, task attribution and relaunch activity (rows 18, 22, #1511, #1512)
One cumulative control-board, webui and seat state. The four rows edit the
same files (scan.mjs, page.html, README.md, app.js), so they land together,
each on its own receipt:

- Row 18, Discord connector rows on the board (#1509): R3 approved by
  Darkwing and Dewey, Gitea comment 26257, manifest 254403b8. Jason
  accepted the visual test.
- Row 22, board attention status (#1503): Filbert approved R1, comment
  26248, manifest e40b58ec; restart receipt 26249.
- #1511, task attribution (row 6 code phase): R2 approved by Filbert and
  Dewey, manifest d4c96395. docs/TOOLS.md carries the approved --by usage
  line (tools-usage.patch 86bcba3c).
- #1512, relaunch activity (row 6 pilot): R1 approved by Darkwing and
  Dewey, candidate manifest 47769fad. All seven source files match it.

Row 16, internal development bootstrap (#1510): the seven files outside
shared records match Filbert's R1 pins, receipt 26204 (agents/researcher/*,
scripts/test-darkwing-launch.mjs, the bootstrap plan).

packages/webui/src/public/app.js is committed at its #1512 R1 pin ce7d79a4.
The working copy holds Dewey's unreviewed return-flow candidate on top of
that, and it stays uncommitted.

Also: the four row briefs and Darkwing's evidence records under
agents/darkwing/work, including the 2026-09-26 tree manifest and the #1512
re-run against 21e3e908. Serial acceptance command: 397/397, three runs.
The failures that only show when tests run concurrently are in #1509 engine
tests, and they reproduce on clean HEAD.

Suites on the exact staged tree: config 24, task 90, foundation 43,
conductor 17, release 14, auth 15, discord 63; package union 397/397
(serial); test-darkwing-launch 5/5.

Shared records (BUILD-LOG, QUEUE, CURRENT, DEFERRED, SESSIONS, AGENTS.md,
agents/README.md) follow in Sage's records commit.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 14:54:18 -05:00
jason.woltjeandClaude Opus 5.5 21e3e908b6 docs(comms): T3 threads as the temporary development channel beside tmux
Jason, 2026-09-26: while Mosaic Stack development runs in T3, development
agents message each other through T3 threads with the t3 MCP tools, until
Mosaic Stack has its own internal comms. ms-communications now picks the
transport by where the recipient runs, carries the T3 header and reply
rules, and counts a returned turnId as the delivery receipt.
docs/guides/T3-AGENT-COMMS.md is adapted to this repository: status,
thread lookup by agent-name title, replies with wait off, and record
locations; references to guides absent here are replaced.

Suites green before commit: config 24, task 90, foundation 43,
conductor 17, release 14, auth 15, discord 63. unslop-check clean on
both files. Reviewed by Sage (T3 thread 1ef1e4f8).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 14:33:24 -05:00
jason.woltjeandClaude Fable 5.1 43d7574d6a feat(discord): SetSpark record client for the Discord Sage, fixed verbs against setspark-api, connector-verified approvals (#1509)
Row 25, parts 2a and 2b, against the shared-signals contract a5425a2.

Model side: eight fixed verbs in the pi extension (record_list, record_get,
record_create, record_update, resolve_id, open_approval_request,
get_approval_request, create_document), each one HTTP call with arguments
checked before any request. Writes carry an idempotency key
<principal>:<message id>:<call index> and an audit context. The seat key is
read from a 0600 file on every call and never cached, printed or journaled.

Connector side: append-only approval ledger, Approve button and exact
"approve" reply resolved by the connector against the required approvers,
confirmation message posted as button evidence, bind and add_approval through
the service under connector keys, retry of unknown entries on start.

Evidence: node tests 162 pass, scripts/test-discord.sh 63/63. Review by
rev-code-02, round 1 approved (#1509 comment 26467, tree 7872d8c5).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-22 12:59:39 -05:00
jason.woltjeandClaude Fable 5.1 1949ed8d31 feat(discord): git verbs for the Discord Sage on the shared-signals root, seat identity through a package credential helper, vault record protocol (#1509)
Row 24. A writable root that is a git work tree may carry a git object in
the binding; the seat then has git_status, git_commit (explicit paths, seat
author, Requested-by trailer from the envelope requester, push at once per
D6), git_pull (ff-only) and git_push (one branch, never force), plus
reserve_id and per-write clone locks under protocol vault. Git children run
with no host config and one credential helper, bin/git-credential.mjs,
reading the 0600 seat token file named in the binding; the fleet helper
serves only the Gitea hosts. Suite 58/58, node 143. rev-code-02 APPROVED
round 1 (#1509 comment 26375, tree 82ab962f).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-18 07:52:35 -05:00
jason.woltjeandClaude Fable 5.1 1685deb423 feat(discord): writes on write-marked roots, web fetch and search, held prompts (#1509)
Row 23. write_file and edit_file for roots marked write: true under the
same fence as reads; web_fetch (https only, public addresses, pinned
connection, capped body) and web_search through SearXNG; extension
renamed to tools.mjs. Engine holds a prompt while pi is busy and sends
it as its own run, so a second message mid-turn no longer folds into
the first (live defect). fake-pi models the real follow-up folding.

Suite 52/52, node tests 129. rev-code-02 APPROVED round 3, comment
26362, tree dbd2ce9a. Records: QUEUE rows 23-24, CURRENT, BUILD-LOG
phase, SESSIONS, row 24 brief (git verbs, D5-D7 ruled).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-18 07:27:50 -05:00
jason.woltjeandClaude Opus 5 1ac812d3d5 feat(discord): read-only tools for the Discord Sage through a Mosaic pi extension confined to declared roots (#1509)
A binding may declare `tools` with named roots. pi starts with
--no-builtin-tools and the package's own extension, allowlisting
list_dir, read_file and search. src/tools.mjs holds the rules: names
not paths, per-segment lstat walk, one checked descriptor read that
refuses symlinks, swaps, FIFOs, hard links and oversize files, credential
shapes refusing the whole read, and a per-message call budget. The engine
settles on agent_end and records tool calls in the turn record.

Jason's rulings R1-R7 in the brief, section 7. rev-code-02 approved
round 2 (comment 26276) on tree 43f0329b after four round 1 fixes.
Suite 48/48, node tests 116. Not pushed.

Co-Authored-By: Claude Opus 5 <[email protected]>
2026-09-14 19:52:21 -05:00
jason.woltjeandClaude Fable 5.1 c4fc8e7d7f docs(discord): brief read-only tools for the Discord Sage as row 21, waiting on D1–D3 (#1509)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 19:33:45 -05:00
jason.woltjeandClaude Fable 5.1 1267e6a0a0 docs(discord): rows 19–20 done, Carmen enrolled live by reload; records and receipt (#1509)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 19:03:53 -05:00
jason.woltjeandClaude Fable 5.1 caaef941e6 feat(discord): binding reload without a restart, and a per-user channel allowlist (#1509)
`reload` validates the binding file and sends SIGHUP to the live owner;
the running connector re-reads it and swaps guildName, channels, users
and limits in place. name, seat, guildId, botUserId, tokenFile, engine
and context are fixed for the life of the process; a change there, an
invalid file or a channel outside the guild refuses the reload and keeps
the old binding. Every attempt is one line in reloads.jsonl. The service
unit maps `systemctl --user reload` to the same signal.

A user entry may carry `channels`, an allowlist of listed channel ids;
absent means every listed channel. Outside the list the message is
dropped as channel-not-for-user; threads count as their parent.

Suite 41/41, 101 node tests. QUEUE rows 19 and 20 opened.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 18:59:31 -05:00
jason.woltjeandClaude Fable 5.1 d9745a4510 docs(discord): row 17 operator check passed, service unit verified by Jason (#1509)
Jason ran traffic under the unit, SIGKILL recovery, the brake, and the
release, and reported all verified. Receipt in the private evidence dir.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 16:29:19 -05:00
jason.woltjeandClaude Fable 5.1 3affbab5e2 docs(discord): row 18 assigned to darkwing by Jason's ruling (#1509)
The control board row for the Discord connector is board-side work in
packages/control-board. Jason ruled darkwing builds it under the board
brief; the coordinator answers connector-side questions only.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 16:20:15 -05:00
jason.woltjeandClaude Fable 5.1 90cb31f56f docs(discord): brief the control board row as iteration 3, blocked on ownership (#1509)
QUEUE row 18. The board discovers rows from pi session directories and
decides liveness by tmux; the connector's session and run.lock live
elsewhere and it has no pane since iteration 2. The brief lists what the
board-side change needs. Owner for Jason to rule.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 14:40:29 -05:00
jason.woltjeandClaude Fable 5.1 436ba6ed6b feat(discord): systemd user service with a supervised run; brakes exit 3 and are never retried (#1509)
QUEUE row 17, MVP iteration 2. scripts/discord-service.sh renders and
installs mosaic-discord@<binding> from packages/discord/systemd/. The
unit's main process is `run --supervised`, which applies the new recover
policy first: a lock whose owner is gone is cleared and only the STOP
written for that is removed; an operator STOP or a held binding refuses
with exit 3, which RestartPreventExitStatus never retries. `recover` is
also a CLI verb. First cut used ExecStartPre and looped live, since systemd
honours the never-retry status only from the main process; replaced and
re-verified before any message traffic. Suite 40/40, 95 node tests.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 14:39:11 -05:00
jason.woltjeandClaude Fable 5.1 dc5902aafd docs(discord): read receipt live check passed, QUEUE row 15 done (#1509)
Jason's message in #sage-admin got the eyes reaction before the reply; the
turn record carries receipt.ok true and a private evidence receipt was
written. CURRENT.md gains the pilot narrative line.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 14:22:34 -05:00
jason.woltjeandClaude Fable 5.1 93d6b62457 feat(discord): eyes reaction as a read receipt on every admitted message (#1509)
QUEUE row 15, MVP iteration 1 after the Sage pilot. rest.react is best
effort (2xx true, anything else false, never throws); the connector reacts
at admission before the engine runs and records the outcome in the turn
record as receipt. Drops and refusals get no reaction. Suite 28/28, 90
node tests.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 14:20:30 -05:00
jason.woltjeandClaude Fable 5.1 788515dc69 docs(discord): close the Sage pilot, Gate H passed; correct the Discord profile wording (#1509)
Live pilot receipts summarized in BUILD-LOG (eight steps, private
evidence under the seat's work directory). Jason ruled the replies read
as Sage; QUEUE row 14 done. DISCORD-USER.md now says unlisted senders
are dropped silently rather than refused with a reply, which is what
the connector does.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 13:39:25 -05:00
jason.woltjeandClaude Fable 5.1 786e379c49 feat(discord): connector pilot for the Sage seat, reviewed candidate (#1509)
Zero-dependency Discord connector under packages/discord: binding
validation, REST and gateway clients, pi engine adapter, journal with
append-only inbox, outbox, admissions and notices, and a run.lock
ownership record {pid, start, boot} whose identity is checked three ways
and whose cleanup is gated by STOP. CLI check|run|stop|unlock via
scripts/discord.sh; offline suite scripts/test-discord.sh (28 checks,
87 node tests).

Reviewed by rev-code-02 on #1509 over nine rounds; approved exact tree
4e0feb6758c0a7e4a71483912a8e0d3e3ec95aef at comment 26170. Corrections
(1) to (12) recorded in BUILD-LOG. No listener started, no token read,
no Discord write; the live pilot follows this commit per the brief.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 01:15:37 -05:00
darkwing b023841c8a docs: publish reviewed CHAT-01C private readback contracts (#1507) 2026-09-13 00:48:26 -05:00
jason.woltje 28d4e98ad8 docs: publish reviewed CHAT-01 draft contracts and fixtures (#1507) 2026-09-12 23:48:03 -05:00
jason.woltje 370823b354 docs: pin CHAT-00 protocol research and synthetic checks (#1507) 2026-09-12 20:43:35 -05:00
jason.woltje 4c436f4ce4 plan: confirm all-seat WebUI session chat and gated delivery (#1507) 2026-09-12 20:28:22 -05:00
jason.woltjeandClaude Fable 5.1 0e77cfd10a Plan page for #1508: queue as data, pieces A-E, Gate G; QUEUE rows 9-13 link to it
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 19:45:14 -05:00
jason.woltjeandClaude Fable 5.1 cbf24b01b0 QUEUE rows 9-13 (#1508): process becomes data, one writer, ledger checks; AGENTS.md cadence reads QUEUE.md first
Jason: the process goes off the rails every time because the entry point is
prose. Rows 9-13 are marked required and cannot be parked.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 19:41:56 -05:00
jason.woltje 11659cf2a1 records: hand Console to Jason for Gate E after publication (#1507) 2026-09-12 19:39:14 -05:00
jason.woltje ea00ec66d9 webui: serve the Console first screen from the control board (#1507) 2026-09-12 19:37:35 -05:00
jason.woltjeandClaude Fable 5.1 7c8e530add QUEUE.md: one task table; CURRENT and DEFERRED point at it
Jason could not find what is next without reading the prose plans. QUEUE.md
is one row per piece with owner, issue, state, gate and brief location.
CURRENT.md becomes the narrative log; DEFERRED.md keeps only gaps.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 19:33:59 -05:00
jason.woltjeandClaude Fable 5.1 355b4308c3 CURRENT points at the piece 5 brief; DEFERRED: shared-checkout hazard (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 19:26:19 -05:00
jason.woltjeandClaude Fable 5.1 f80404f541 brief for piece 5, darkwing on point, with Gate F (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 19:25:53 -05:00
jason.woltje fab40f25f8 records: close ledger and seat registration after owner acceptance (#1506, #1504) 2026-09-12 19:18:47 -05:00
jason.woltjeandClaude Fable 5.1 b7f8d1cc35 Gate D written; brief for piece 4, the WebUI first screen on the Console design (#1506, #1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 19:05:44 -05:00
jason.woltjeandClaude Fable 5.1 46ded3fcba record Jason's comms rule and working preferences in the repository (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 18:59:33 -05:00
jason.woltje cd0aa5fbb3 ledger: count issue-tagged commits and repo seat messages (#1506) 2026-09-12 12:04:15 -05:00
jason.woltjeandClaude Fable 5.1 db0d784bd8 plans: add DEFERRED.md, the one list of gaps found and not yet handled (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:53:24 -05:00
jason.woltjeandClaude Fable 5.1 889d87500f CURRENT: next action is piece 3, the ledger, per Jason's go (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:42:56 -05:00
jason.woltjeandClaude Fable 5.1 0334026e13 control board: brief for piece 3, the ledger, with Gate D (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:42:28 -05:00
jason.woltjeandClaude Fable 5.1 c90ce3a836 control board: Gate C passed, #1505 closed, next action waits for the ledger brief (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:39:23 -05:00
jason.woltjeandClaude Fable 5.1 4f83068097 control board: log Gate C pass for reply-from-board (#1505)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:38:55 -05:00
jason.woltjeandClaude Fable 5.1 867619dca2 control board: reply from the board through agent-send.sh (#1505)
Piece 2 of the MVP (#1503). A one-line reply box and Send in the detail
of rows with a live registration; POST /api/reply runs
tools/tmux/agent-send.sh -s <session> -S <host>:control-board
[-L <socket>] -m <text> once for one seat and returns the exit code,
stdout and stderr. The page shows delivered or failed with the tool's
stderr; other rows say "reply needs a registered seat". No send-keys,
queue, retries, history or broadcast; packages/seat and agent-send.sh
untouched. Every message ends with a fixed trailer telling the seat to
answer in its own session (Jason's refinement after the first Gate C
exchange; the board has no pane). Board suite 98/98.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:30:30 -05:00
jason.woltjeandClaude Fable 5.1 a62ca1904f control board: brief for piece 2, reply-from-board, with Gate C (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:17:28 -05:00
jason.woltjeandClaude Fable 5.1 0bed9ba3db control board: show the model each seat is running (#1503)
Jason asked for the running model (sonnet, opus, gpt-6-astra) on the
board. readSession keeps model and provider from the log's latest
model_change entry or assistant turn, whichever is later, so a /model
switch shows on the next scan; scanAgent exposes both; the page shows the
model under the agent name with the provider in the hover and a Model
detail row. Blank when the log names none.

Live check on a scratch board against the real data root: all 42 rows
carried a model. Board 91/91. Sonnet review caught a double-escaped hover
title; fixed.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:13:32 -05:00
jason.woltjeandClaude Fable 5.1 d0fb5b0f20 control board: log Gate B pass, seat registration shown on the board (#1504)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:13:05 -05:00
jason.woltjeandClaude Fable 5.1 17153fe140 control board: stale registrations and a launch-test leak (#1504)
Two defects in 69f99323, reported by the professor session and verified.

The darkwing launch test's flock-contention spawn ran without the fixture
config, so launch.sh re-entered scripts/mosaic against the real data root
and wrote fixture records for darkwing, dewey and filbert there. That
spawn now names the fixture config, and both launch test files set
MOSAIC_CONFIG to a nonexistent path and clear MOSAIC_LAUNCH_REGISTERED
process-wide, so a spawn that forgets fails instead of polluting.

A registration is written before the launch script's own checks, so a
refused launch left a record with a dead pid that the board honoured. The
scanner now probes the recorded pid (pidAlive, signal 0); a gone pid makes
the record stale: still on the Registered line with alive false, derived
task, project and workspace win, index gains registrationStale, CLI
summary gains a stale count.

Fleet launchers marked not planned per Jason. Board 90/90, seat 15/15,
launch scripts 5/5. Sonnet review APPROVED.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 11:04:49 -05:00
jason.woltjeandClaude Fable 5.1 69f99323c7 Add mosaic launch <seat> with seat registration for the control board (#1504)
New package packages/seat and wrapper scripts/mosaic. `launch <seat>` writes
<dataRoot>/seats/<layout>/<seat>/registration.json and then execs the seat's
launch.sh unchanged; `seat task <seat> <text>` edits the task only. The board
reads registrations, matches by sessions directory, and lets a registered
task, project or workspace override the derived value with a source tag.
The four repository launch scripts register themselves unless already
registered or run with --check. Fleet launchers untouched; one-liner on the
plan page.

Review found the record path keyed by seat name alone (repo and fleet
"darkwing" would collide); fixed by keying on layout. Also: the Pi pin
refusal now names installed and required versions.

Tests: seat 15, control-board 89, launch scripts 5, registry 69, config 24.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 10:56:17 -05:00
jason.woltjeandClaude Fable 5.1 01d9a19612 control board: record Jason's selection of Dewey design 3 (Console) for the WebUI (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 10:39:54 -05:00
jason.woltjeandClaude Fable 5.1 4a7e16c3ec Show task, active project and workspace per control board row (#1503)
Gate A fix asked by the professor session on Jason's behalf. Each row
now carries three derived fields, shown as "unknown" when the log and
tmux do not hold them:

- task: the session's first user message (pi logs have no task envelope)
- workspace: the live pane path of the pane running pi, else session cwd
- activeProject: basename of the nearest git checkout above the workspace

tmuxInspect replaces the bare liveness call in the CLI and returns
{ alive, workspace }; tmuxIsAlive stays as a wrapper. The grouping column
and seen.json keys are unchanged. Fixture test per field, tmux parse
tests, page test; missing launcher signals are recorded in the plan page.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 09:52:45 -05:00
jason.woltjeandClaude Fable 5.1 6ec253de20 Pin the mid-tool-call rule in the control board scanner (#1503)
Acceptance rule (plan page, bf641e22): a seat mid-tool-call is working,
never waiting. The scanner already met it through pi's stopReason values;
deriveState now also checks the content for a toolCall block (working),
after the error stop reasons and before "stop" (waiting). Thinking blocks
do not keep a text turn from being waiting. Three JSONL fixture tests and
three state-table cases pin the rule. Live check on the real board:
orch-01 and rev-code-01 mid-tool-call are working, velma's finished
text-only turn is waiting. Sonnet review: APPROVED.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 09:29:19 -05:00
jason.woltjeandClaude Fable 5.1 bf641e220a control board: park D05/registry-3/new seats and add the mid-tool-call working rule (#1503)
Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 09:25:32 -05:00
jason.woltjeandClaude Fable 5.1 bf1391a424 Show "shown of total" in control board project headers while rows are hidden (#1503)
A project header now reads "fleet (13 of 38)" while Hide offline or Hide
seen hides at least one row, and "fleet (38)" when nothing is hidden. The
note under the table still says which filter hid how many. Numbers only,
so nothing new needs escaping. Static test pins the expression and the
removal of the raw-length header. Sonnet review: APPROVED, no findings.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 09:00:21 -05:00
jason.woltjeandClaude Fable 5.1 e90d15dd70 Add a per-project "Hide seen" checkbox to the control board page (#1503)
Each project table now has "Hide seen" beside "Hide offline", both on by
default, with a note saying how many rows each one hides. The choice
survives the 10-second refresh. Static test pins the markup, the filter,
the persistence guard, and the change handler. Sonnet review: APPROVED.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 08:55:22 -05:00
jason.woltjeandClaude Fable 5.1 b70749219f Add a collapsed "Seen" section to the control board page (#1503)
Jason asked for a way to recall rows he marked Seen. The page now lists
them under a collapsed "Seen (N)" section between "Waiting on you" and
"By project", each with Unsee. Open/closed state survives the 10-second
refresh because only the section body is re-rendered. Page-only change
plus one static test. Tests: control-board 64/64, registry 69/69.
Review APPROVED, no findings.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 08:45:51 -05:00
jason.woltjeandClaude Fable 5.1 88d21defde Check pi liveness per tmux pane and add "Seen" marks to the control board (#1503)
First step-3 refinement from Jason's daily use. Liveness now lists the
panes of the agent's tmux session and counts it alive only if a pane runs
pi, so killed pi sessions whose tmux session still exists show offline
instead of waiting. A "Seen" button on waiting and error rows stores the
row's lastActivity in <dataRoot>/board/seen.json (clicks only, never
rewritten by a scan, fail closed if corrupt) and drops the row from
"Waiting on you" until the agent writes anything newer; "Unsee" reverses
it. New POST /api/seen route: JSON only, 4 KB limit, 400 on bad input.

Tests: control-board 63/63 (30 new), registry 69/69. Review APPROVED;
receipt docs/plans/reviews/2026-09-12_control-board-step3-seen-marks.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 08:24:49 -05:00
jason.woltjeandClaude Fable 5.1 ebedd1281e Add control board web page and local server (#1503)
Step 2 of the control board MVP (MOSAIC-STACK-D-001): `serve` command starts
a loopback-only local server that serves one self-contained page and re-runs
the status scanner on each /api/board request. The page lists sessions
waiting on Jason first (errors on top), then one table per project with
plain-word states, ages, last messages, expandable detail rows, per-project
hide-offline, and a 10-second auto-refresh with pause.

Tests: control-board 33/33 (10 new: loopback rules, host refusal, all routes,
per-request rescan, 500 path, CLI refusals, live serve, page escaping guard);
registry 69/69 unchanged. Receipt:
docs/plans/reviews/2026-09-12_control-board-step2-review.md.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 07:58:38 -05:00
jason.woltjeandClaude Fable 5.1 b9f59a5903 Add control board status scanner and MVP plan (#1503)
Step 1 of the control board MVP (decision MOSAIC-STACK-D-001): a plan page,
Gitea #1503, and packages/control-board, which reads each agent's newest pi
session log plus tmux liveness and writes one status file per agent under
<dataRoot>/board/. 23/23 tests; independent review approved after three
fixes (length stopReason as error, unknown liveness state, secrets-boundary
test). CURRENT.md now points at step 2, the page.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 07:26:03 -05:00
jason.woltjeandClaude Fable 5.1 1993039c76 Record owner acceptance and closure of #1500 fixture increment
Jason reran the two-test fixture demonstration on the canonical checkout
at 5abbabb7 (2/2 pass). Records-only: acceptance receipt, BUILD-LOG,
SESSIONS and CURRENT.md next action (MVP re-plan). No source change.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 06:47:24 -05:00
jason.woltje 5abbabb74b Record fixture increment review, publication and owner test gate (#1500) 2026-09-10 20:23:37 -05:00
jason.woltje 3daee5ad89 Add fixture-only execution materialization and refresh (#1500) 2026-09-10 20:20:44 -05:00
jason.woltje 27e4873acc Record reviewed prerequisite correction and separate next approval gate (#1500) 2026-09-10 16:41:36 -05:00
jason.woltje 6335342873 Correct registry validation and contain metadata reads (#1500) 2026-09-10 16:40:16 -05:00
jason.woltje b3fa221060 Record owner acceptance of increment 1 (#1499) 2026-09-10 16:17:22 -05:00
jason.woltje 5446097483 Record increment 1 approval after contamination fix (#1499) 2026-09-10 15:23:58 -05:00
jason.woltje e5a9a05ba2 Drop uncommitted sage agent contamination and dangling export (#1499 I1-F1/F2) 2026-09-10 15:23:02 -05:00
jason.woltje e6e30b6f55 Record increment 1 review request (#1499) 2026-09-10 15:18:27 -05:00
jason.woltje 557aba0f9e Add packages/mosaic registry schemas, validation, read-only CLI; pin pi 0.85.1 (#1499) 2026-09-10 15:18:15 -05:00
jason.woltje 54c1231280 Draft M20 increment 1 charter for owner approval (#1499) 2026-09-10 15:10:20 -05:00
jason.woltje 21461c8674 Register completed registry gate 7 closure 2026-09-10 15:08:58 -05:00
jason.woltje 86e009dde5 Close registry gate 7 with owner-approved refresh mechanism 2026-09-10 15:08:40 -05:00
jason.woltje 224147eccc Record owner rulings resolving registry review gates 2026-09-10 14:42:31 -05:00
jason.woltje d5307d2b0a Register completed registry alignment records 2026-09-10 14:28:07 -05:00
jason.woltje b8008eda5d Record owner account-selection rulings and registry draft revision 2026-09-10 14:27:47 -05:00
jason.woltje 50d2a2eba9 Record independent verification and completion of skill repair (#1498) 2026-09-08 18:13:46 -05:00
jason.woltje f3dce32088 Repair native launcher skill rename and fail-closed regression coverage (#1498) 2026-09-08 18:12:20 -05:00
jason.woltje 12ff5da7df Record owner acceptance and close publication recovery trial (#1497) 2026-09-08 17:57:29 -05:00
jason.woltje 10448e41a2 Record verified publication and owner acceptance handoff (#1497) 2026-09-08 17:20:12 -05:00
jason.woltje 29c1defe29 Publish reviewed development agents and WUI draft with recovery evidence (#1497) 2026-09-08 17:15:03 -05:00
jason.woltje 3b7fd19d08 docs: verify published wave and set registry review next 2026-09-08 12:44:13 -05:00
jason.woltje 69f10a4062 fix(tmux): explicit transport-only dispatch and safe remote quoting (#1496) 2026-09-08 12:41:25 -05:00
jason.woltje 67eaf6fb47 docs(logs): pending session/build records through 2026-09-07
Append-only BUILD-LOG phases, SESSIONS registrations, and CURRENT
checkpoint state accumulated through the consolidation and inspector
review waves.
2026-09-07 14:07:16 -05:00
jason.woltje 3ea385223e feat(skills): six new ms-* skills
ms-archify (evidence-based architectural mapping), ms-sdlc,
ms-proactive-agent, ms-goal, ms-grill-me, ms-frontend-design.
2026-09-07 14:07:16 -05:00
jason.woltje 193479b52d docs: concept annexation, provider/reference docs, ACT-1 groundwork
Mosaic concepts pages now own the adapted content; source/license
metadata under docs/reference/concepts. Adds ACT-1 agent-context
planning capture, pinned concept test package + preparation utility,
foundation observation notes (durability, evidence, federation,
onboarding, workflow), and the #1495 consolidation assessment.
TOOLS.md updated for the host-dev launcher.
2026-09-07 14:07:05 -05:00
jason.woltje 7c580a5625 feat(agents): darkwing host-dev launcher + entry-point consolidation
scripts/agent.sh --host-dev delegates to scripts/agent-host-dev.sh;
agents/darkwing/launch.sh provides the native development TUI
(context, skills, coding tools, /goal). Root SOUL.md is the M14-era
default-collaborator persona captured by the ACT-1 context work.
Launcher regression checks pass (test-darkwing-launch.mjs).
2026-09-07 14:06:54 -05:00
jason.woltje 8ebddd6f93 feat(foundation): offline synthetic scope/permission inspector (FI-FILBERT-8 APPROVED r6)
Rocko-authored, Filbert-reviewed inspector (r6 manifest
a4a44930...) with full review/build/verdict evidence under
docs/plans/reviews. 43/0 selftests, oracle zero-disagreement,
foundation checker PASS. Owner A9 acceptance recorded separately.
2026-09-07 14:06:35 -05:00
jason.woltje 127a54fdff chore: consolidate new foundation and archive v1 (#1495) 2026-09-07 12:32:57 -05:00
Dewey 9a5fbdbda7 fix(goal): quiet waits and unify fleet NG ownership (#56, #57, #58) 2026-09-06 04:07:09 -05:00
jason.woltje 7345f330fc docs: map foundation to integrated rewrite baseline 2026-09-06 02:40:05 -05:00
jason.woltje d4696d09eb feat(extensions): establish canonical goal source (#54, #55) 2026-09-06 02:32:32 -05:00
jason.woltje 44f257cb06 docs: record accepted phase-2 foundation contract 2026-09-06 02:23:15 -05:00
jason.woltje 69d1bb3aa4 docs(plan): resolve harness IDs + lifecycle review gate (#50)
Owner adjudication:
- canonical harness IDs match executables: pi, claude, codex, opencode
- agent.json uses one scalar harness ID; registry/manifest resolution, no
  hard-coded schema enum
- target mosaic harness list/detect/install/rm/status lifecycle
- detection recognizes reviewed executables and records compatibility
  without reading/copying harness homes
- installs are exact-version/verified, Mosaic-managed, never global
- detected external harness is available but not container-ready until
  imported/installed, absent a separately reviewed host adapter

Gate 1 resolved. Remaining P0/review gates stay open; no implementation
authorized. Suites 24/15/90/14/17 + verify green; unslop clean.
2026-09-04 12:26:57 -05:00
jason.woltje 9ea7dea711 docs(review): record independent glm-5.3 auth-registry spec review (#50)
Verdict ACCEPT WITH CHANGES. Persist ten-gate recommendations and three
P0 blockers: per-seat launch provider/model resolution; rotating OAuth
persistence for long-running seats; role auth ceiling ∩ settings profile
plus data-map/reset alignment. Reviewer made no repo edits.

Implementation remains blocked pending owner/conductor adjudication.
Prior gate remains green: suites 24/15/90/14/17 + verify.
2026-09-04 12:14:48 -05:00
jason.woltje d9a94f51ba docs(plan): specify harness declaration + centralized auth/provider registry (#49)
Design only; implementation blocked pending owner review.

- agent.json: one harness identifier (pi first), resolved through a
  versioned adapter/harness manifest; reusable settingsProfile reference
- central data-root registry: providers, accounts (metadata + secret
  credential split), reusable settings profiles, audited runtime selection
- per-seat pi auth.json/models.json mechanically generated and atomically
  activated; no seat/provider registration ceremony
- mixed oauth/api-key accounts supported centrally; one active account per
  provider per pi materialization
- host-side centralized OAuth login/refresh; agents never authenticate
- local/remote Ollama modeled as endpoint providers, not accounts
- target mosaic auth/provider/agent settings CLI; secrets never on argv
- migration, fail-closed acceptance suites, and ten explicit review gates

CURRENT.md points only to spec review. Suites 24/15/90/14/17 + verify
green; unslop clean.
2026-09-04 11:57:40 -05:00
jason.woltje 975084abe2 fix(auth): mosaic-managed auth lives under the data root, never ~/.pi (#48)
Owner direction: the stack must never impact default harness usage.
Correction to M19 as shipped (nothing had been created in ~/.pi — the
move breaks nothing).

- Mosaic-managed accounts: <dataRoot>/auth/<account>.json, perms 0600
  enforced (loose perms flagged in listings, refused by --auth — mirrors
  gitea-api.sh credential hygiene).
- ~/.pi is read-only to the stack, permanently; the only interaction
  remains the existing read-only container mount of the default
  credential. Recorded as a ROADMAP standing decision.
- auth.sh is now config-driven (data root from config.json, fail closed,
  consistent with every other tool); status reports both sources labeled.
- agent.sh --auth resolution moved after load_config (needs the data
  root); missing/symlinked/non-0600 accounts refuse.
- test-auth.sh: 15 no-Docker cases (accounts-create-nothing, loose-perms
  refusal, invalid-config refusal added). Test-authoring correction
  recorded in BUILD-LOG (fixture-state mismatch caught before running).

Suites 24/15/90/14/17 + verify green.
2026-09-03 22:53:33 -05:00
jason.woltje 073bbfdb6a feat(auth): M19 harness auth tooling — auth.sh checkpoint + per-launch account injection (#47)
Investigation (pi 0.84.4 docs + host auth.json metadata, values never
read): provider stacking is native (one auth.json keyed by provider;
resolution --api-key > auth.json > env > models.json; OAuth auto-refresh).
Multi-account per provider is NOT native -> named-file design:
auth.<account>.json + per-launch injection.

- scripts/auth.sh: status (provider names, credential types, perms,
  env-side names informational — never credential material) and accounts
  (named files, active marker). Exit codes per convention: 3 missing for
  a read, 2 unparseable, 4 file/environment (symlinks refuse).
- scripts/agent.sh --auth <account>: resolves auth.<account>.json and
  exports PI_AUTH_FILE (the existing compose read-only mount source — no
  new plumbing); missing/invalid account refuses pre-container.
- scripts/test-auth.sh: 13 no-Docker cases; core assertion is the safety
  property itself — fixture key/token/env VALUES never reach output.
- Docs: TOOLS.md Auth section, AGENTS.md command surface + suites.

Headless task runs keep the default credential (worker auth selection is
a separate policy decision). Real-host smoke: anthropic/openai-codex
oauth + zai api_key reported, perms 600, no named accounts yet.

Suites 24/90/14/17/13 + verify green. Agreed sequence M16-M19 complete;
M20 owner-gated.
2026-09-03 19:58:50 -05:00
jason.woltje d1d7b5598d feat(agent): fail-closed seat resolution under MOSAIC_AGENTS_DIR override (#46)
Owner decision after live verification of M18: an explicit agents-dir
override that cannot resolve the named seat now refuses the launch
(exit 4, names the seat and dir) instead of launching seatless and
unbounded. Unsetting the override keeps the M13 plain governed TUI.
MOSAIC_ROLES_DIR needs no symmetric change - the M18 gate already
refuses unresolvable role contracts.

Task suite 88 -> 90 (refusal + refusal-names-the-seat). TOOLS.md Agent
section documents the refusal.

Suites 24/90/14/17 + verify green.
2026-09-03 19:45:10 -05:00
jason.woltje ca8135d70c feat(roles): M18 seat-role progressive capability restriction (#45)
Role contracts (roles/<role>.json): roleVersion, name bound to filename,
tools ceiling (subset of pi built-ins), network declared (none|api-only|
open; enforced when network policy lands). Strict schema, fail closed -
a non-role document refuses resolution.

mosaic-task.mjs resolve-role: config-free contract validation, emits
MOSAIC_ROLE_TOOLS / MOSAIC_ROLE_NETWORK.

agent.sh: a declared role binds to its contract. Missing/invalid contract
refuses the launch (exit 2, names the role - the under-equipped-seat
failure mode, mirroring M17 skills). Effective tools = ceiling ∩ requested
(CLI --tools or agent.json caps); no request -> ceiling stands; narrowing
and tool-free outcomes loud on stderr. Adapters unchanged; headless M9
chain (mission ∩ task) untouched.

Ships roles/researcher.json (existing seat declares the role; without the
contract the fail-closed gate would refuse its launch).

Task suite 74 -> 88: contract resolution, wrong-kind/name/network/
duplicate/unsupported/missing refusals, ceiling narrowing E2E (mock
adapter), tool-free E2E, missing-contract refusal. Test-authoring
correction recorded in BUILD-LOG (a check that registered on one path
only, caught by count arithmetic).

Suites 24/88/14/17 + verify green.
2026-09-03 17:25:17 -05:00
jason.woltje bf56583a49 docs(skills): owner loop-doctrine + collaborator delivery discipline, remediated (#44)
ms-communications: integrated as-authored - owner preamble restructure +
collaborator delivery-discipline hunks from the #43 calibration (own
session output is not a send path; the tool performs the preamble flip;
receiving rule 3 requires actually running agent-send.sh).

ms-conductor: collaborator redraft integrated (canon-aligned tracking
surfaces, one-action cadence, fail-closed core) with one conductor
remediation - step 3 now distinguishes refusal (fail closed, never
bypass) from runner outage (direct dispatch to a qualified live seat via
ms-communications permitted, recorded loudly as degraded: no sandbox, no
run record; suites still gate integration). Preserves the owner's
outage-dispatch intent inside invariant 6.

docs/TOOLS.md: release.sh ensure row added (M16 subcommand existed in
code but not in the doc - flagged by the collaborator, verified in
release.sh usage).

Authorship: owner (ms-conductor doctrine, preamble restructure) +
ms-test collaborator (delivery hunks, redraft); remediation + integration
by conductor (dragon-lin:darkwing). Suites 24/74/14/17 + verify green;
unslop clean.
2026-09-03 17:17:29 -05:00
jason.woltje c8f433131c docs: calibration phase 22 record; CURRENT.md staleness corrected (M16/M17 late-logged, next M18); session registered (#43) 2026-09-03 16:55:31 -05:00
jason.woltje d1e75f855c docs(tools): document tools/ tree in TOOLS.md + fix stale suite counts (#43)
Collaborator-authored via conductor-loop calibration: task dispatched to the
live ms-test seat (glm-5.3-flash) over agent-send.sh; diff reviewed line by
line and every documented flag/exit code independently verified against tool
source by the conductor; suites green at integration (config 24 / task 74 /
release 14 / conductor 17 + verify).

- new 'Tools (host-side)' section: agent-send.sh, agent-watch.sh, unslop-check.js
- intro reading guide now points at tools/ (worker-flagged addition, accepted)
- Maintenance suite counts corrected: test-task.sh 58 -> 74

Authored-by: ms-test collaborator (glm-5.3-flash)
Integrated-by: conductor (dragon-lin:darkwing)
2026-09-03 16:55:25 -05:00
jason.woltje f710bf1a68 docs: KICKSTART file 2026-09-03 16:34:20 -05:00
jason.woltje 2085d75190 docs: ms-communications skill (owner-authored) + session registry entry + KICKSTART recovery file 2026-09-03 16:34:03 -05:00
jason.woltje 2f5a8d2cec feat(skills): skill lifecycle + ms-* skill set completion (#40, #41, #42)
- scripts/skill.sh: install (bundled or path) / activate / deactivate /
  uninstall (refuses while enabled) / list
- skills-enabled + skills-available dirs under the data root; a skill not
  in skills-enabled is not enabled or available for use
- pi adapter: MOSAIC_SKILLS -> --skill per dir; --no-skills when none
- agent.sh: seat definitions declare skills[]; resolution against
  skills-enabled refuses the launch loudly when missing
- ms-* skills completed (owner-authored canon, hands-off): ms-tools
  adapted to the runtime, ms-file-read/write/agent/conductor bodies
  written in the owner's style; ms-agent-watch + ms-unslop untouched
- tasks/USER.md onboarding fixtures; suite hardening (nested def path,
  user seed, mock-adapter dispatch evidence)

Suites: config 24, task 74, release 14, conductor 17, verify PASS.
RELEASE 0.0.12 packaged; health-gated activation on merge.

Closes #40, closes #41, closes #42
2026-09-03 16:23:57 -05:00
jason.woltje e14ad9ab52 Merge: skill lifecycle, skills completion, M16 self-determination hardening, M20 decision
- skill.sh lifecycle (install/activate/deactivate/uninstall/list)
- 8 ms-* skills completed (owner canon preserved)
- seat skills dispatch + enabled-dir resolution
- M16: ensure at launch, drift warnings, recursion guard
- M20 decision: packages/* monorepo at usurpation

Closes #40, closes #41, closes #42
2026-09-03 16:03:05 -05:00
jason.woltje 9fd16b9739 feat(release): recursion guard for the health gate; run-task drift warning; M20 packages/* decision recorded (#39)
- release.sh health gate runs with MOSAIC_ENSURE_SKIP=1: the gated task run
  cannot re-enter release self-determination
- run-task.sh warns on release drift instead of silently using a stale image
- ROADMAP: M20 decision recorded (packages/* monorepo at usurpation,
  continuity-first); restructure sequenced as M20 phase 1

Closes #39
2026-09-03 15:58:46 -05:00
jason.woltje 9051ad179b docs(roadmap): restructure sequencing (M20 phase 1, not first) + skills-as-discipline doctrine 2026-09-03 14:50:59 -05:00
jason.woltje 2f649ed930 docs: README release ensure row 2026-09-03 14:43:25 -05:00
jason.woltje db330c12c7 feat(release): self-determination - ensure at launch, drift warnings, M20 packages/* decision (#38)
- release.sh ensure: aligned no-op; drift -> package-if-needed + health-gated
  activate (M3 gate-then-flip, automated)
- ensure_release_aligned in common.sh: invoked by hello/verify/agent;
  MOSAIC_ENSURE_SKIP guards recursion; run-task warns on drift without
  auto-aligning (workers/suites never trigger builds or model gates)
- ROADMAP: M20 decision recorded - v2 adopts packages/* monorepo at usurpation
- BUILD-LOG Phase 20 + tool-race process note

Live-verified: post-reset pointer loss auto-restored via health-gated
ensure; drift warning fires on desired-version bump; idempotent no-op on
aligned state.

Closes #38
2026-09-03 14:42:09 -05:00
jason.woltje 0a17d29bce docs(roadmap): flush 2026-09-03 14:28:58 -05:00
jason.woltje 3458f6a7ad docs(roadmap): M17 skill lifecycle (skills-enabled/available, role-scoped subsets, --skill negates --no-skills); M20 stack succession path + monorepo question 2026-09-03 14:28:45 -05:00
jason.woltje b033952cd5 docs(harvest): pattern ledger from stack/next + fleet runtime (12 patterns, skips, owner-corrected M17-M19 designs) 2026-09-03 13:55:30 -05:00
jason.woltje c102980ad4 docs(plan): CURRENT.md - queue aligned to ROADMAP (M16 next, CI deferred) 2026-09-03 12:59:28 -05:00
jason.woltje 58b96cb715 docs(plan): ROADMAP.md - M16 release self-determination, M17 ms-tools skill, M18 seat-role restriction, M19 auth tooling; CI deferred per owner 2026-09-03 12:57:26 -05:00
jason.woltje dd3ff944a1 chore(release): 0.0.11 2026-09-03 12:30:52 -05:00
jason.woltje 8eb81ebec1 feat(onboard): user onboarding - no default USER.md, guided creation (#37)
- bootstrap no longer creates user/USER.md (owner direction)
- scripts/onboard.sh: name REQUIRED (interactive loop or --name),
  optional fields prompted (profession, marital, age, gender, education,
  location, timezone, skillset, interests, hobbies, pets); flag-driven
  non-interactive mode for automation
- templates/USER.md: canon skeleton, placeholder rendering, unfilled
  optional = (not provided)
- agent.sh: auto-runs onboarding when profile missing (TTY gate);
  headless run-task warns and continues without user context
- user profile dispatched to all launches (M14 layer)

Closes #37 (onboarding requirements from owner layout review)
2026-09-03 12:30:42 -05:00
jason.woltje 530597cc84 docs(plan): CURRENT.md flush 2026-09-03 11:59:49 -05:00
jason.woltje 9a0d44f96a docs(plan): CURRENT.md - deduplicated completed log (marked correction), M15 review queued 2026-09-03 11:59:37 -05:00
jason.woltje 121b331c6c docs(log): back-fill phases 16-19 (M10, M12, M14, M15) - recorded retroactively with ground-truth sources 2026-09-03 11:59:14 -05:00
jason.woltje 34e06e7de7 Merge M15: agent seats - per-agent SOUL and role contracts
Closes #36
2026-09-03 11:56:38 -05:00
jason.woltje 9bd4f1c405 feat(agents): agent seats - per-agent SOUL, role, definitions dir (#36)
- agents/<name>/ holds agent.json (strictly validated: version, name,
  role?, capabilities?, workspace?, session?) + SOUL.md (persona prose)
- agent.sh: definition loading (quote-safe node defaults file), runtime
  SOUL copy to dataRoot/agents/<name>/, MOSAIC_AGENT_SOUL_FILE ->
  loader fills the SOUL slot from the seat's persona (contract SOUL =
  default persona; governance never overridden)
- seat.json written once at instantiation (seatVersion, name, role, at)
- identity section gains agent role; compose passthrough for role+SOUL
- live user context (M14) + seat SOUL compose the full persona:
  governance -> persona -> identity -> user -> mission
- RELEASE -> 0.0.10; packaged and health-gated activated
- example seat committed: agents/researcher

Closes #36
2026-09-03 11:56:38 -05:00
jason.woltje a7b612435b docs: BUILD-LOG Phase 15, CURRENT.md - M13 shipped 2026-09-03 11:25:41 -05:00
jason.woltje 87f10772ce Merge M13: interactive TUI agent + TOOLS.md
Closes #35
2026-09-03 11:24:56 -05:00
jason.woltje 7db4c5c2ed feat(agent): interactive TUI launcher + identity + TOOLS.md (#35)
- scripts/agent.sh <name>: launches interactive pi TUI in the container
  with contracts + optional mission + agent identity + named session +
  optional workspace/tools; the Mosaic alternative to vanilla pi
- pi adapter: MOSAIC_INTERACTIVE branch (clean TUI, no -p, no initial
  prompt); headless exec rebuilt via positional args (no word-splitting
  on the request); MOSAIC_AGENT_NAME optional in headless
- loader: AGENT IDENTITY section when the launcher names the agent
- compose: fixed command removed (request defaults live in run-agent.sh);
  MOSAIC_INTERACTIVE/MOSAIC_AGENT_NAME passthrough
- docs/TOOLS.md: full on-demand tool reference; AGENTS.md routes to it
- RELEASE -> 0.0.8 (container change); build verified

Closes #35
2026-09-03 11:24:56 -05:00
jason.woltje 0273a84549 docs: AGENTS.md - session recovery shim, invariants canon, session registry
- AGENTS.md at root: pi loads it automatically at every session start
  (conductor-level sessions; workers deliberately exclude it via
  --no-context-files). Deliberately short: invariants, session protocol,
  role model, command surface, data map, pointers - depth stays in docs/.
- docs/SESSIONS.md: append-only session registry, mandatory per session.
- Recovery rule encoded: compaction/restart loses nothing - AGENTS.md +
  CURRENT.md + git log + suites reconstruct state; never guess.
2026-09-03 11:02:19 -05:00
jason.woltje 3b674b7a66 Merge: roles/ directory convention - root is bootstrap-only 2026-09-03 10:56:07 -05:00
jason.woltje 527bc581ca refactor(layout): role contracts move to roles/ - root is bootstrap-only
Owner direction: the repository root holds first-class, bootstrap-required
configuration only. conductor-policy.json is a ROLE contract (the
conductor's authority), one of scores of future role contracts
(agent-policy, coder-policy, ...) - such files get a dedicated home.

- roles/conductor-policy.json (git mv)
- conductor-apply.sh + test-conductor.sh read the new path
- CONDUCTOR.md records the roles/ convention

Closes UX follow-up from owner layout review; no issue (convention change).
2026-09-03 10:56:07 -05:00
jason.woltje 1249714a9a docs(plan): CURRENT.md - M12 shipped 2026-09-03 07:04:37 -05:00
jason.woltje e175616885 feat(conductor): auto-apply policy gate for worker patches (#34)
- conductor-policy.json (tracked, strictly validated): enabled switch,
  path allowlist globs, gating suites - the autonomy decision lives in a
  declarative file the owner controls
- scripts/conductor-apply.sh <runId> [--dry-run]: succeeded-run check ->
  clean target tree -> diff from worker workspace -> allowlist -> syntax
  gates (node/bash/json) -> apply -> policy suites -> attribution commit;
  ANY failure reverts the tree; push is never automatic
- scripts/test-conductor.sh: 17 sandbox cases covering every gate incl.
  suite-failure auto-revert and disabled policy
- policy defaults: scripts/docs/tasks/missions/adapters + README; all
  three suites gate

Closes #34
2026-09-03 07:03:29 -05:00
jason.woltje 6955717612 Merge M11: session forking from a common ancestor
Closes #33
2026-09-03 06:43:02 -05:00
jason.woltje 88d9cf750f feat(sessions): sessionForkFrom - branch conversations from a common ancestor (#33)
- task schema: optional sessionForkFrom (source session name); requires
  session target; self-fork rejected
- runner: resolves source newest .jsonl (fail 4 if none/outside dataRoot);
  passes MOSAIC_SESSION_FORK + MOSAIC_SESSION_DIR; result records lineage
- pi adapter: --fork <source> --session-dir <target> when forking;
  ephemeral default unchanged; plain session resume unchanged
- compose passthrough; RELEASE -> 0.0.7 (adapter changed)
- suite +9 cases (58 total): plumbing via mock stderr, validation
  negatives, live fork - child recalls ancestor code word, ancestor
  session file untouched

Closes #33
2026-09-03 06:43:02 -05:00
jason.woltje c038706eed docs(plan): CURRENT.md - M10 shipped, retention next in review 2026-09-03 06:33:54 -05:00
jason.woltje 8622c9d826 Merge M10: run-record retention
Closes #32
2026-09-03 06:33:27 -05:00
jason.woltje 88eef507b0 feat(retention): run-record pruning - keep newest N, dry-run default (#32)
- mosaic-task.mjs prune [--keep=N] [--yes]: default keep 50; without
  --yes lists candidates without deleting
- only r-* directories under the runs root; symlinks skipped;
  sessions/workspaces/state/config untouched (asserted by suite sentinels)
- append-only receipt runs/.pruned.log records every pruned id
- test-task.sh: +8 retention cases (dry-run no-delete, keep-N, newest
  kept, receipt, isolation, invalid keep, empty no-op)

Also: suite hardening - prune section scopes its config per-command
(no export/unset leaking into later sections); duplicated check()
removed; latest_reason hoisted to helpers; status colors now green OK /
red FAIL (terminal-only, NO_COLOR-aware) per owner UX feedback.

Closes #32
2026-09-03 06:33:27 -05:00
jason.woltje 439bea6915 ui(test): green OK/PASS, red FAIL - terminal-only, NO_COLOR-aware
Owner feedback: grep match-highlighting made the word 'policy' red while
status words were plain - counter-indicative. Suites + verify now emit
ANSI colors (green success, red failure) when stdout is a terminal;
piped/machine-parsed output stays plain, honoring NO_COLOR. Word 'ok'
promoted to 'OK' for scannability.

Verified byte-level via forced-pty run; piped output unchanged; suites
41/24/14 + verify green.
2026-09-03 06:23:47 -05:00
jason.woltje fad8a4718c Merge M9: mission-level capability policy
Closes #30
2026-09-03 06:16:01 -05:00
jason.woltje 2ff49adff4 feat(policy): mission-level capability policy - least-privilege intersection (#30)
- mission schema: optional capabilities.tools (same validation as task)
- merge semantics in runTask: neither -> none; mission only -> mission;
  task only -> task; both -> intersection (task narrows, never widens);
  empty intersection -> tool-free run with an explicit stderr note
- result.json records EFFECTIVE tools; task/mission snapshots remain the
  immutable declaration of intent
- adapters unchanged; host-side only (no image change, 0.0.6 still active)
- task suite +5 cases (41 total): all four merge cases asserted from run
  evidence + invalid mission capabilities rejected

Policy decision recorded: missions govern; tasks cannot escalate.

Closes #30
2026-09-03 06:16:01 -05:00
jason.woltje 44c476ebbf fix(ops): show displays retriedFrom lineage (#29)
result.json recorded lineage correctly; the human-facing show command
omitted the field. Found by owner test: show | grep retriedFrom was
empty on a run whose result.json contained it.

Closes #29
2026-09-03 06:09:51 -05:00
jason.woltje cde480eb60 docs(plan): CURRENT.md — retry lineage shipped, M9 queued for decision 2026-09-03 05:31:24 -05:00
jason.woltje 5808248707 Merge retry lineage + relative mission resolution
Closes #28
2026-09-03 05:30:57 -05:00
jason.woltje afd5827db8 fix(retry): lineage tracking + relative mission path resolution (#28)
- retryRun rewrites a snapshot's relative mission path to the run's own
  recorded mission.json (absolute) before execution — retries stay
  faithful to what originally ran
- runTask accepts options.retriedFrom; retry records lineage in
  result.json (additive optional field, no schema break)
- task suite +4 cases: retry succeeds, lineage recorded, mission section
  present after retry (36 total), missing-run retry exits 4

Closes #28
2026-09-03 05:30:57 -05:00
jason.woltje d9cc990376 Merge M8: conductor loop - self-orchestration
Closes #25, closes #26, closes #27
2026-09-02 22:42:52 -05:00
jason.woltje 83c4e9851e feat(orchestration): retry <runId> — authored by headless pi worker (#26, #27)
Collaboration record (conductor loop, docs/plans/CONDUCTOR.md):
- round 1 (worker session worker-1, 2m28s): retry implemented per spec
- conductor live test exposed spec gap: direct invocation lacked
  launcher env exports
- round 2 (same worker session, 59s): spawnEnv made self-sufficient,
  but used PI_* where compose interpolates MOSAIC_*
- conductor hotfix: 3-line rename to MOSAIC_PROVIDER/MOSAIC_MODEL/
  MOSAIC_DATA_ROOT

Final: node scripts/mosaic-task.mjs retry <runId> re-executes a run's
task snapshot as a new run; live retry replied REMEMBERED; all suites
green (24/32/14 + verify).

Known limitation: retrying a run whose task used a RELATIVE mission path
resolves it against the temp dir; lineage tracking deferred.

Closes #25, closes #26, closes #27
2026-09-02 22:42:52 -05:00
jason.woltje 22508170a2 docs(plan): CURRENT.md — single next-action pointer for cadence-driven work 2026-09-02 22:25:56 -05:00
jason.woltje 90a67d050e Merge M7: operator ergonomics + release 0.0.6
Closes #24
2026-09-02 22:14:18 -05:00
jason.woltje 24bdef75fa feat(ops): run inspection, release 0.0.6, docs (#24)
- mosaic-task.mjs show <runId>: full record + snapshots + artifacts;
  uppercase-tolerant id validation; missing/traversal ids exit 4
- list: task/workspace/session columns
- RELEASE -> 0.0.6; README workspaces/capabilities/sessions sections;
  BUILD-LOG Phases 9-11; autonomous-run tracker results filled

Closes #24
2026-09-02 22:14:18 -05:00
jason.woltje 4e2a413640 Merge M6: named sessions - persistence and resume
Closes #22, closes #23
2026-09-02 22:08:34 -05:00
jason.woltje be55549700 feat(sessions): named persistent sessions with resume (L1) (#22, #23)
- task schema: optional session (named id) -> persistent session dir at
  dataRoot/sessions/<name>, isolated per name
- pi adapter: --session-dir when declared (ephemeral --no-session stays
  the default otherwise); -c resumes the most recent session when present
- compose passthrough; result.json records session
- fixtures: tasks/session-demo-1.json (teach) + session-demo-2.json (recall)
- E2E: teach -> REMEMBERED + host-side session JSONL; resume -> recalled
  'mosaico' exactly; single continued session file

Closes #22, closes #23
2026-09-02 22:08:34 -05:00
jason.woltje ddb1554e5b Merge M5: task workspaces + capability envelope
Closes #20, closes #21
2026-09-02 22:06:43 -05:00
jason.woltje b017e66e17 test(capabilities): workspace/tooling selftests + live demo fixture (#21)
- mock plumbing cases: workspace path + tools delivered (asserted from
  run-record stderr), host workspace created, absent fields = empty vars
- validation negatives: unknown tool, workspace traversal
- tasks/workspace-demo.json: pi uses bash inside the persistent demo
  workspace; host-visible proof.txt verified live

Closes #21
2026-09-02 22:06:43 -05:00
jason.woltje 172368612c feat(capabilities): task workspaces + tools allowlist plumbing (#20)
- task schema: optional workspace (absent | :run ephemeral | named
  persistent under dataRoot/workspaces) and capabilities.tools (pi
  documented tool allowlist); strict validation, traversal-proof names
- runner: creates host workspace, passes MOSAIC_WORKSPACE (container
  path) + MOSAIC_TOOLS; result.json records both
- pi adapter: cds into workspace; --tools when allowlist present else
  --no-tools
- mock adapter: logs delivered MOSAIC_* vars to stderr as deterministic
  plumbing evidence (dash prints 'export K=v', so use env not export)

Closes #20
2026-09-02 22:03:46 -05:00
jason.woltje 1387231e57 docs(plan): autonomous work run tracker (M5-M7 scope, test plan, review checklist) 2026-09-02 21:57:30 -05:00
jason.woltje 0292392e64 Merge M4: runtime adapter seam
Closes #16, closes #17, closes #18, closes #19
2026-09-02 21:30:35 -05:00
jason.woltje 594b8d711c docs(adapters): adapter seam docs + recorded M4 E2E, release 0.0.5 (#19)
- README: Runtime adapters section (contract summary, selection, mission
  injection point); BUILD-LOG Phase 8 entries
- E2E: 24+24+14 selftests green; verify PASS; 0.0.5 packaged and
  health-gated activated; mission-bearing fixture task succeeded through
  the real pi adapter; config checksum unchanged

Closes #19
2026-09-02 21:30:35 -05:00
jason.woltje 3c1ffd2c2d test(adapters): seam selftests — deterministic mock cases + mission injection (#18)
- test-config: absent adapter defaults to pi; mock validates; unknown
  adapter exits 2; env exports MOSAIC_ADAPTER (24 cases total)
- test-task: mock adapter gate pass/mismatch (no provider needed),
  mismatch reason asserted, unknown adapter fails closed, mission section
  injected into generated prompt asserted by content (24 cases total)
- harness fixes: helpers defined before use; per-case config files (no
  cross-case leakage); newest-run selection for the live case; deduped
  accidentally duplicated live block

Closes #18
2026-09-02 21:28:57 -05:00
jason.woltje 4ebb123ba3 feat(adapters): sanctioned mission directives injection (#17)
- load-contracts.sh: MOSAIC_MISSION_FILE (readable) appends a MISSION
  (runtime) section — objective + directives — after the immutable
  contracts; unreadable path is a hard error, absent env changes nothing
- mosaic-task.mjs: exports MOSAIC_MISSION_FILE as the run snapshot's
  container path (/var/lib/mosaic/runs/<id>/mission.json), with an
  outside-dataRoot guard; also exports the configured adapter

Verified: contract-only prompt has no mission section; mission-bearing
run shows objective + directives in the generated prompt, snapshot
recorded, real provider returns exactly MOSAIC_HELLO_OK.

Closes #17
2026-09-02 21:19:22 -05:00
jason.woltje bb5cecb348 feat(adapters): adapter contract, dispatch, pi + mock adapters (#16)
- adapters/README.md: the harness boundary contract (env in, response on
  stdout, diagnostics stderr, exit 0 success)
- adapters/pi: extracted current invocation unchanged
- adapters/mock: deterministic MOSAIC_MOCK_RESPONSE echo (test-only)
- run-agent.sh: name-validated dispatch to adapters/<name>/adapter.sh
- config: optional execution.adapter (pi|mock), default pi, configVersion
  stays 1 — existing configs remain valid; selection authority is the
  config file (load_config exports it)
- compose: MOSAIC_ADAPTER / MOSAIC_MOCK_RESPONSE passthrough; Containerfile
  installs adapters read-only; RELEASE -> 0.0.5

Verified: hello unchanged; mock verbatim via config; unknown adapter and
path-traversal names refused in-container; invalid adapter exits 2.

Closes #16
2026-09-02 21:18:07 -05:00
jason.woltje 88f55d9135 test(task): live failures self-report evidence; wrong-exit no longer masks (#15)
- live hello failure dumps latest run result.json + stderr tail before
  sandbox cleanup destroys them
- wrong-expectExact case asserts reason == expect-mismatch (was: any
  exit 1, which masked compose-level failures)
- repair dangling if/else from the docker-guard refactor

Closes #15
2026-09-02 20:56:58 -05:00
jason.woltje e7e1bd26eb fix(launcher): resolve release identity for direct task runs (#14)
M3 made MOSAIC_IMAGE_TAG required in compose, but run-task.sh never
called load_release — direct task runs failed in compose before any
model call. release.sh paths masked it by exporting the tag to children.

Found by owner-run test-task.sh; failure receipts were in the run
records' stderr.txt.

Closes #14
2026-09-02 20:44:31 -05:00
jason.woltje 5d86b8fa93 Merge M3: release model and safe updates
Closes #10, closes #11, closes #12, closes #13
2026-09-02 20:24:33 -05:00
jason.woltje 35ea464661 docs(release): release model usage + recorded M3 drills (#13)
- README: Release model section (package/activate/rollback/status,
  pointer + append-only log, gate-then-flip guarantee)
- BUILD-LOG Phase 7: drills recorded (update, refusal, rollback),
  two harness/product corrections documented

Drill evidence: 0.0.3 -> 0.0.4 update with unchanged config checksum and
green verify; fault-injected refusal left pointer untouched; health-gated
rollback restored 0.0.3; full append-only event history.

Closes #13
2026-09-02 20:24:33 -05:00
jason.woltje e87ecdb3e5 test(release): release-layer selftests (#12)
14 cases: RELEASE validation (valid/invalid/missing), tag consistency,
status on empty state, fault-injected refusal with no pointer + single
valid refusal log line, healthy activation, pointer fields, repeat
activation append-only log, rollback-without-previous refusal.

Harness fix learned the hard way: restore RELEASE from backup inline
after the missing-file case (mv-back restored the mutated file); single
exit trap self-heals the repo state.

Closes #12
2026-09-02 20:23:18 -05:00
jason.woltje a947db7bfd feat(release): package/activate/rollback/status with health gate (#11)
- activate: image-presence pre-check + M2 task-runner health gate
  (tasks/hello-marker.json exact marker) before atomic pointer replace
  (tmp+rename); every attempt appended to activation-log.jsonl
- --fault-injection flips the health expectation to prove the refusal path
- rollback: health-gated re-activation of the previous activated imageTag
  from the log; refuses when the image is gone or no previous exists
- status: release, tag, pointer, recent log; safe on empty state
- state lives under <dataRoot>/state/ (config-independent, reset-scoped)

Verified: activate OK; fault-injected refuse with pointer unchanged;
rollback-without-previous refuse.

Closes #11
2026-09-02 20:20:32 -05:00
jason.woltje a35ea62ab1 feat(release): RELEASE identity + image tag single-sourcing (#10)
- RELEASE file: single source of release version (0.0.X until declared stable)
- common.sh load_release(): validates version, derives
  MOSAIC_IMAGE_TAG=mosaic-poc-agent:<pi>-r<release> from the pinned pi dep
- compose.yaml: image tag is required env; build/hello/verify call load_release
- verify.sh derives the image name instead of hardcoding it
- package.json version aligned to the same 0.0.X line

Closes #10
2026-09-02 20:18:17 -05:00
jason.woltje f6c93dcb5c Merge M2: mission and task abstraction
Closes #6, closes #7, closes #8, closes #9
2026-09-02 19:58:53 -05:00
jason.woltje 7f0408d417 feat(task): fixture mission/task, docs, recorded M2 E2E (#9)
- missions/hello.json + tasks/hello-marker.json committed fixtures
- README: Missions & tasks section (usage, run record layout, M2 scope note)
- BUILD-LOG: Phase 6 before/after entries

E2E: fixture run succeeded with exactly MOSAIC_HELLO_OK; wrong expectExact
recorded status failed (expect-mismatch) and exited 1; runs listed; config
checksum unchanged.

Closes #9
2026-09-02 19:58:53 -05:00
jason.woltje cfd2a19bd7 test(task): mission/task selftests — schema negatives + live runs (#8)
18 cases: validation negatives (unknown keys, versions, ids, prompt,
expectExact NUL, timeout range, missing/invalid mission, writes-nothing)
plus live cases: exact-marker success, wrong expectExact fails, distinct
run dirs, result.json contents, list output.

Closes #8
2026-09-02 19:57:34 -05:00
jason.woltje d2a9e26395 feat(task): mission/task schemas, validation, and run runner (#6) (#7)
- scripts/mosaic-task.mjs: validate | run | list
- Strict v1 schemas: unknown keys rejected; ids/prompt/expectExact/
  timeoutSeconds bounds enforced; optional mission file resolved against
  the task file and validated too
- run: executes through the config-driven container path with stdin
  detached (issue #5 class), SIGKILL timeout (default 120s), trimmed
  response capture
- Immutable run records under <dataRoot>/runs/r-<utcstamp>-<rand>/:
  task.json + mission.json snapshots (write-once), stderr.txt, result.json
- expectExact gate: mismatch -> status failed, exit 1; result.json is
  always written
- scripts/run-task.sh: load_config + bootstrap_runtime_dir before exec
- M2 scope: mission directives are snapshotted for provenance, not yet
  injected into the runtime prompt (later policy layer)

Closes #6, closes #7
2026-09-02 19:56:29 -05:00
jason.woltje 0734b1f3a5 fix(launcher): detach stdin on agent container run (#5)
pi print mode reads piped stdin until EOF; an attached terminal stdin
blocked the one-shot run forever. Automated contexts (closed stdin)
never exposed it. Request text comes from the compose command.

Proven: tail -f /dev/null | scripts/hello.sh now returns MOSAIC_HELLO_OK
in ~4s (previously timed out at 30s); verify.sh remains green.

Closes #5
2026-09-02 19:50:07 -05:00
jason.woltje ce6420f3de Merge M1: configuration-driven Hello World
Closes #1, closes #2, closes #3, closes #4
2026-09-02 18:32:50 -05:00
jason.woltje 81f58b15c8 fix(launcher): ensure configured data root before verify mount; docs for M1 (#4)
- verify.sh now calls bootstrap_runtime_dir after load_config; previously a
  reset-then-verify flow let Docker auto-create a root-owned mount source
- common.sh: fail with clear guidance when data root exists but is not writable
- README: configuration section, bootstrap usage, selftest entry point
- BUILD-LOG: Phase 5 entries with corrections

E2E (clean slate): 20/20 selftests; bootstrap idempotent; config-driven
hello/verify MOSAIC_HELLO_OK; negative marker exit 1; reset + rerun green;
config checksum unchanged across the entire flow.

Closes #4
2026-09-02 18:32:50 -05:00
jason.woltje 0c2113f710 test(config): sandboxed config-layer selftests (#3)
20 cases: bootstrap create/idempotency, missing config, malformed JSON,
unknown keys/version/backend/environment, relative and non-canonical
dataRoot, filesystem root, home dir, ancestor-of-config, control chars,
symlinked config file, env export resolution, validation-writes-nothing.

Closes #3
2026-09-02 18:30:46 -05:00
jason.woltje 900a506c1f feat(config): wire launcher scripts and compose to config.json (#2)
- common.sh: load_config() exports MOSAIC_DATA_ROOT/PROVIDER/MODEL; fails closed
- compose.yaml: dataRoot mount and provider/model are required env (:? errors)
- build/hello/verify load config before any mutation; no silent bootstrap
- reset.sh: target resolved from configured dataRoot; all safety checks kept

Verified: compose fails without launcher env; verify/reset fail on missing
config; config-driven hello+verify pass; symlink refusal with sandboxed
config (canary survived); config checksum unchanged across reset+rerun.

Closes #2
2026-09-02 18:30:08 -05:00
jason.woltje c3d29e796a feat(config): config module with idempotent bootstrap and strict v1 validation (#1)
- scripts/mosaic-config.mjs: bootstrap | validate | env operations
- Exclusive creation (O_EXCL 'wx'); existing config validated, never rewritten
- Strict schema: unknown keys rejected, configVersion===1, backend docker only
- dataRoot guards: absolute, canonical, not root/home/ancestor-of-config
- MOSAIC_CONFIG override for sandboxed tests; exit codes 0/2/3
- scripts/bootstrap.sh: explicit bootstrap entry point

Closes #1
2026-09-02 18:28:14 -05:00
jason.woltje c2365ae519 chore: baseline container POC and atomic foundation plan
- Containerized Pi hello-world proof (image mosaic-poc-agent:0.84.4, non-root)
- Four immutable contract fixtures loaded into a generated system prompt
- build/hello/verify/reset scripts with exact-match gating and reset safety
- Documented Pi discovery (v0.84.4, -p mode, --system-prompt, container auth)
- Append-only BUILD-LOG with corrections; deferred layers in LAYERS.md
- Architecture plan: docs/plans/2026-09-02_atomic-mosaic-foundation.md
2026-09-02 18:24:36 -05:00
6039 changed files with 950524 additions and 858 deletions
+23 -15
View File
@@ -1,16 +1,24 @@
# Mosaic Stack standalone deployment (compose `stack` profile)
# Copy to .env and adjust. Port overrides exist because the defaults
# collide with common host services (and with the dev compose itself).
PG_HOST_PORT=5433
VALKEY_HOST_PORT=6380
GATEWAY_HOST_PORT=14242
# Registry image override (defaults to a local build of docker/gateway.Dockerfile):
# GATEWAY_IMAGE=git.mosaicstack.dev/mosaicstack/stack/gateway:sha-acf640d
# Non-secret runtime settings for the mosaic-poc-agent container.
# Copy to .env if you want to override the defaults in compose.yaml.
#
# NEVER put credentials in this file. Authentication is supplied at
# runtime only, via one of the two documented paths:
# 1. read-only mounted pi auth file (default: ~/.pi/agent/auth.json,
# override the host path with PI_AUTH_FILE)
# 2. provider API key environment variable (ZAI_API_KEY or
# ANTHROPIC_API_KEY), passed through by compose.yaml when set
# Optional explicit dogfood overlay (docker-compose.dogfood.yml).
# All three paths are required when that overlay is used. Use a dedicated
# next-based worktree, its canonical clone's .git directory, and the external
# home of the unprivileged code-dogfood-01 functional seat.
# MOSAIC_DOGFOOD_WORKTREE=/home/example/src/mosaic-stack-worktrees/dogfood-1487
# MOSAIC_DOGFOOD_COMMON_GIT_DIR=/home/example/src/mosaic-stack/.git
# MOSAIC_DOGFOOD_SEAT_HOME=/home/example/.mosaic/fleet/agents/code-dogfood-01
# Model provider (built-in pi provider name)
PI_PROVIDER=zai
# Model ID within the provider
PI_MODEL=glm-5.3-flash
# Optional: alternative host path of the pi credential file mounted
# read-only at /home/node/.pi/agent/auth.json in the container
#PI_AUTH_FILE=/home/jwoltje/.pi/agent/auth.json
# Optional: documented env-var auth alternative (secret! set in your
# shell or a gitignored .env, never commit)
#ZAI_API_KEY=
#ANTHROPIC_API_KEY=
+6 -27
View File
@@ -1,29 +1,8 @@
logs/
node_modules
dist
.turbo
.next
coverage
# build/deps
node_modules/
# runtime credentials — never commit, never copy into the image
.env
.env.local
*.tsbuildinfo
.pnpm-store
__pycache__/
docs/.obsidian
secrets/
# Step-CA dev password — real file is gitignored; commit only the .example
infra/step-ca/dev-password
# Scratch dirs created by the framework git-wrapper shell test harnesses
.mosaic-test-work/
# Transient config files vite/vitest/esbuild write next to a *.config.ts while
# loading it, then unlink. They are untracked but were not ignored, so turbo's
# package traversal hashed them and intermittently failed CI with "Package
# traversal error: ... .timestamp-*.mjs: No such file or directory" when the
# file vanished mid-scan. Ignoring them removes the race.
*.timestamp-*.mjs
# Playwright run artifacts (#1445, P6 E2E gate)
apps/web/test-results/
apps/web/playwright-report/
# generated runtime state lives in /home/jwoltje/.mosaic-dev (outside this project)
+6
View File
@@ -0,0 +1,6 @@
extensions/
extensions.installed.sha256
.extensions-*
state/
evidence/
native-test-*.log
+34
View File
@@ -0,0 +1,34 @@
# Native goal development copy
From this repository, start a fresh native Pi session:
```sh
bash scripts/goal-dev.sh
```
Canonical source lives under `extensions/`. The launcher first runs `scripts/sync-dev-extensions.sh`, which installs verified ordinary-file copies under `.pi/extensions/`, then loads only the generated goal extension. Global extensions remain unloaded. The launcher keeps your usual native Pi provider authentication; it copies no credentials. Goal state and new conversation files live under `.pi/state/`, which is ignored by Git. Each process gets a fresh incarnation; `/reload` and `/new` in the same process retain its goal. Restarting Pi does not adopt an earlier process's active goal.
Plain `pi` also discovers `.pi/extensions/goal/index.ts` after project trust, but may load global extensions too. Use the launcher to avoid duplicate `/goal` registrations. This is a local development test, not a sandbox or the managed Mosaic runtime. Docker and `~/.mosaic` are unchanged.
## Try it
1. Set `/goal <a long goal with acceptance criteria>`. This starts work immediately.
2. Look below the editor for `Goal: Active`. The old above-editor goal widget is gone.
3. Run bare `/goal`, then press `Alt+G`. Both show the entire stored goal and its status. Tab remains autocomplete.
4. Use `/goal stop` and `/goal resume`. Expect Paused and Active, or Waiting if an untimed wait remains recorded.
5. A blocked `goal_report` displays Blocked. A satisfied report displays Complete and retains the full goal for recall without continuing work.
6. `/goal clear` removes the retained goal. Try `NO_COLOR=1 bash scripts/goal-dev.sh` to check text-only labels.
Use terminal scrollback for recall longer than the screen. At narrow widths Pi may truncate its footer status row; bare `/goal` and Alt+G remain available.
## Checks
```sh
node --test extensions/goal/test/*.test.ts
bash scripts/test-extension-package.sh
python3 scripts/test-goal-native.py
```
Contract tests use ordinary read-only fixture copies in `test/fixtures/skills-local/`, not live brain files. The executive-update fixture SHA-256 matches the parser's pinned contract, `bbea48a46b1f8da7bc759f86856fb52830b7dde456b826317163c6dc6ccab319`.
`SOURCE-SNAPSHOT.json` records the original external-source baseline, not the edited candidate. No symlinks are used. Never edit `.pi/extensions/`; the sync script refuses to overwrite installation drift. Make changes under `extensions/`, run the checks, and relaunch. To disable the test, stop its Pi process and remove `.pi/extensions/`. Keep `.pi/state/` only if you need local test state.
+10
View File
@@ -0,0 +1,10 @@
{
"snapshotVersion": 1,
"copiedAt": "2026-09-06T04:58:22Z",
"source": "~/.mosaic/fleet/extensions",
"goalTreeSha256": "8853f2b72dde3e87c4573648b9a931c1c75da87ccde995c3224e6d2e707a75f0",
"mosaicCoreLibTreeSha256": "d1194dce31209e5773c6cc5ce571cbca3c39b29d943a79dea06665e05d29f319",
"symlinks": false,
"autoDiscoveredExtensions": ["goal"],
"purpose": "Issue #54 native Pi NG development copy; never loaded by Docker"
}
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Compatibility entrypoint for the accepted native test command.
set -euo pipefail
cd "$(dirname "${BASH_SOURCE[0]}")/.."
exec scripts/goal-dev.sh "$@"
+184 -153
View File
@@ -1,186 +1,217 @@
# Agent Guidelines — Mosaic Stack
# AGENTS.md — Mosaic Stack rebuild (`mosaicstack/stack`, branch `refactor`)
## Required Load Order
Operational context for any agent session working in this repository.
Read top to bottom; it is deliberately short — depth lives in the files it
points to, not here.
1. `~/.config/mosaic/SOUL.md`
2. `~/.config/mosaic/STANDARDS.md`
3. `~/.config/mosaic/AGENTS.md`
4. `~/.config/mosaic/guides/E2E-DELIVERY.md`
5. `AGENTS.md` (this file)
6. Runtime-specific guide: `~/.config/mosaic/runtime/<runtime>/RUNTIME.md`
## What this repository is
## Project Context
Canonical checkout: `/mnt/storage/src/mosaic-stack`, origin `mosaicstack/stack`,
working branch `refactor` (Jason-authorized conversion, issue #1495).
The new foundation is at the root. `v1/` is archived legacy source, not the current
implementation; its instructions and tools do not govern the new foundation.
`~/src/mosaic-stack-dev-test` is a compatibility symlink to this checkout, not a
second working tree. Both original Git histories are retained. Conversion receipt:
`docs/plans/2026-09-07_repository-consolidation-completed.md`.
Mosaic Stack is a self-hosted, multi-user AI agent platform. It is a TypeScript monorepo with a NestJS gateway, Next.js dashboard, Pi SDK agent runtime, and Discord/Telegram plugin architecture.
A rebuild of Mosaic Stack: a file-based, fail-closed
orchestration foundation that dispatches sandboxed headless pi workers to do
real work, with immutable run records as evidence. Thirteen-plus tagged
milestones (`git tag -l`) from `poc-container-hello-v0` to today; suites
green at every step. Not production software — a proven foundation.
### Stack
## Non-negotiable invariants (the canon)
- **API:** NestJS with Fastify (`apps/gateway`)
- **Web:** Next.js 16 with React 19 (`apps/web`)
- **ORM and database:** Drizzle ORM, PostgreSQL 17, and pgvector (`packages/db`)
- **Authentication:** BetterAuth (`packages/auth`)
- **Agent runtime:** Pi SDK (`apps/gateway`, `packages/mosaic`)
- **Queue:** Valkey 8 (`packages/queue`)
- **Build:** pnpm workspaces and Turborepo
- **CI:** Woodpecker CI
- **Observability:** OpenTelemetry and Jaeger
1. **Root is bootstrap-only.** First-class system configuration lives at the
repository root; everything else gets a dedicated directory (`roles/`,
`contracts/`, `missions/`, `tasks/`, `docs/`). Do not add new files to root.
2. **Configuration**: `~/.config/mosaic-dev/config.json` is the sole system
config — created only by `scripts/bootstrap.sh`, never overwritten,
fail-closed on any problem. Repo-scoped role authority lives in
`roles/*.json` (versioned, reviewed commits only).
3. **Secrets** never enter the repository or container images; auth is
runtime-only (read-only mount or environment variable).
4. **Contracts** (`contracts/`) are immutable and image-baked. Missions and
tasks are declarative JSON with strict schemas.
5. **Run records** under `<dataRoot>/runs/` are write-once evidence — never
rewritten, only pruned via `prune` with a receipt.
6. **Fail closed**: missing or invalid config/policy refuses the operation.
Never improvise around a refusal; diagnose it.
7. **Policy**: missions govern tasks (least-privilege intersection — a task
narrows, never widens). Role authority is declared in `roles/` and changes
only via reviewed commits.
8. **Git**: commit only after applicable suites are green. Work on the
owner-authorized `refactor` branch; never force-push. Push remains an explicit
act. Do not merge into `next` or `main` without separate authorization.
`scripts/conductor-apply.sh` commits locally; it does not authorize a push.
9. **Append-only logs**: BUILD-LOG.md (phases), `activation-log.jsonl`,
`.pruned.log`, docs/SESSIONS.md. Corrections are new entries, never edits.
### Package Map
## Autonomous operation within an agreed plan
| Package | Purpose | Key Dependencies |
| ------------------ | ----------------------------- | -------------------------------- |
| `apps/gateway` | NestJS API + WebSocket hub | Fastify, Socket.IO, Pi SDK, OTEL |
| `apps/web` | Next.js dashboard | React 19, Tailwind |
| `packages/types` | Shared TypeScript contracts | class-validator |
| `packages/db` | Drizzle schema and migrations | drizzle-orm, postgres |
| `packages/auth` | BetterAuth configuration | better-auth, @mosaicstack/db |
| `packages/brain` | Structured data layer | @mosaicstack/db |
| `packages/queue` | Valkey task queue and MCP | ioredis |
| `packages/coord` | Mission coordination | @mosaicstack/queue |
| `packages/mosaic` | Unified `mosaic` CLI and TUI | Ink, Pi SDK, commander |
| `plugins/discord` | Discord channel plugin | discord.js |
| `plugins/telegram` | Telegram channel plugin | Telegraf |
Autonomy starts after alignment, not before it. For a new substantial assignment,
recover the applicable mission, goal, task, `CURRENT.md` state, and prior owner
decisions, then work with the user to establish a plan of action: the intended
outcome, acceptance evidence, boundaries, and any gated actions. Recommend a
concrete plan instead of presenting an open-ended menu. A direct request or
existing approved plan that already settles those points is sufficient alignment;
do not ask for ceremonial reconfirmation.
## Architecture and Code Conventions
Once the plan is established, carry it to verified completion without prompting
for routine decisions or permission to take the next in-scope step. Authorization
persists for the life of that assignment unless the user changes or revokes it.
Treat mid-session user input as steering: incorporate it, update the plan or
tracking record when needed, and continue.
1. Gateway is the single API surface; all clients connect through it.
2. Pi SDK is ESM-only; gateway and CLI code must remain ESM.
3. Use `"type": "module"`, NodeNext module resolution, and `.js` extensions in imports.
4. Keep typed Socket.IO events in `@mosaicstack/types` to enforce client/server contracts.
5. Import OTEL tracing before NestJS bootstrap (`import './tracing.js'`).
6. Use explicit `@Inject()` decorators in NestJS because tsx/esbuild does not emit decorator metadata.
7. Keep DTOs in `*.dto.ts` files at module boundaries.
8. BetterAuth owns authentication tables; their schema is defined in `@mosaicstack/db`.
9. Create a task-specific scratchpad for non-trivial work.
### Decide and continue
## Development Workflow
- Resolve naming, implementation approach, layout, and similar non-breaking
choices from, in order: repository invariants and role policy, the approved
plan and acceptance criteria, established repository conventions, then the
smallest reversible option. Record a consequential choice and its tradeoff.
- Perform the in-scope investigation, edits, tests, documentation, and tracking
needed for end-to-end acceptance. Do not ask whether to add obviously required
tests or documentation.
- Diagnose failures and retry or remediate within the agreed scope. Fix a defect
when it blocks acceptance or is local to files already being changed; otherwise
record a bounded follow-up without expanding the assignment.
- Resolve minor ambiguity in favor of the mission, goal, north star, and prior
owner decisions. State the assumption in the completion report.
- Never stop merely to ask whether to proceed, which routine option to use, or
whether to execute the next step already contained in the plan.
Requirements: Node.js 20+, pnpm 10.6.2, and Docker Compose when optional local services are needed.
### Re-align or stop only at a real boundary
```bash
pnpm install --frozen-lockfile
pnpm preflight
Finish all independent work first, then ask one focused question only when:
# Optional local queue service only; do not start the full Compose stack.
docker compose up -d valkey
```
1. Two plausible readings materially change the outcome and the choice is costly
to reverse.
2. The next action would exceed the agreed scope or authority, introduce an
unapproved breaking public/API/schema/data/policy change, or alter a security
boundary.
3. Credentials or access are missing and no in-scope path remains.
4. The action is destructive, irreversible, production-affecting, incurs spend,
or communicates externally on the user's behalf without explicit authority.
5. Objectives or owner decisions genuinely conflict and repository evidence
cannot resolve them.
6. A fail-closed policy refusal or another agent's overlapping ownership prevents
safe progress. Diagnose and report it; never route around it.
The pre-push hook requires:
Repository gates still apply. In particular, a successful implementation or a
broad request to “finish” does not by itself authorize push, merge, deployment,
release, production changes, policy/role expansion, or access to secrets. Perform
such an action only when the established plan explicitly includes it. If blocked,
report the exact boundary, what is complete, the recommended resolution, and the
specific action that will resume; do not use “waiting for confirmation” as a
substitute for a real blocker.
```bash
pnpm preflight && pnpm typecheck && pnpm lint && pnpm format:check
```
## Session protocol (mandatory)
Software delivery also requires the applicable tests. Common repository commands are:
- **Register** your session in `docs/SESSIONS.md` — one append-only line
(date, actor, scope, outcome). Never rewrite or remove entries.
- **Cadence**: run `scripts/mosaic queue next <your seat>` first. It names
the row to resume, review or start, or says there is nothing. The goal order
in `docs/plans/2026-09-27_goals-review.md` sets priority, not CURRENT.md.
Open only the brief that row links to. Execute it through every authorized
stage (implement → test → verify against acceptance criteria; commit, push,
or close only when the established plan authorizes each) → move the row with
`scripts/mosaic queue move` (never by editing QUEUE.md) → register in
SESSIONS.md.
- "next" means one action. A batch mandate ("run the queue") repeats the
loop until green or truly blocked under the boundary rules above.
- Substantial work gets a Gitea issue and a BUILD-LOG phase entry
(before/after, with corrections recorded honestly).
```bash
pnpm typecheck # TypeScript checks across the workspace
pnpm lint # ESLint across the workspace
pnpm test # Checkout tests and package Vitest suites
pnpm format:check # Prettier check
pnpm build # Build all packages and applications
```
## Internal development bootstrap
## Branch Model and Merge Process — `main` and `next` (CANONICAL)
Jason's current direction is repository-native development in
`/mnt/storage/src/mosaic-stack`. Sage leads the project (Jason's ruling,
2026-09-26) and coordinates coding, review and research through Darkwing, Dewey,
Filbert, Rocko, Researcher and any further seats Jason launches under `agents/`.
Darkwing is a collaborating agent seat, not the coordinator. Development sessions
run in T3 for now. Work moves to the new stack; the old `~/.mosaic` fleet is being
retired, and a fleet seat acting outside Jason's instructions is the failure this
transition exists to prevent.
Do not assign new development work to fleet seats during this bootstrap phase.
Do not modify `~/.mosaic` launchers, provisioning or other state, or stop/migrate
live fleet processes as part of this work. Preserve existing work and histories.
Use the repository bootstrap/configuration and launch entry points; missing
configuration still fails closed. This changes development coordination, not
managed worker role policy or deployment authority. The lead role adds no push,
merge or deployment authority; those still need Jason's say-so. See
`agents/README.md` for the internal roster.
**Every contribution targets `next` first. No exceptions.** Features, fixes, tests,
docs, and policy changes all take the same route; urgency changes queue priority,
never the route. Agents never commit to or merge into `main`.
For control-board attention, start a completed reply with `Input needed: ` and
one specific nonempty request only when Jason must provide a decision or input.
Put that line at column zero, before other text. Do not use it for routine
completion or a wait on another agent. Ordinary completed replies are idle.
Use code fences or blockquotes when showing this convention as an example.
The signal is advisory status, never permission for a protected action. Seen
acknowledges a request; it does not resolve it. See `packages/control-board/README.md`.
| Branch | Role | Who merges into it |
| ------ | ---------------------------------------------------------------- | --------------------------------------------------------------------------- |
| `next` | Integration trunk — the only PR target for contributions | The designated merge-gate agent, after all gates pass. Never the PR author. |
| `main` | Stable/release line — receives promotion merges from `next` only | Jason only (or an agent he explicitly delegates for a named promotion). |
## Role model
### Contribution sequencing (in order, no skipping)
- **Conductor**: a system-scoped role — not an agent, not a daemon. Holds
git/credentials/policy authority; decomposes, dispatches, reviews,
verifies, integrates. Protocol: `docs/plans/CONDUCTOR.md`. Exists only
when invoked; push is never automatic.
- **Workers**: headless pi via `scripts/run-task.sh` — sandboxed workspace,
tools allowlist, optional persistent sessions and forks; no git, no
credentials, no policy control.
- Worker runs deliberately exclude this file (`--no-context-files` in the
adapter): worker context is contracts + mission via the generated system
prompt. This file is for conductor-level sessions.
1. **Issue first.** Work is tracked in a Gitea issue before a branch exists. The
issue number appears in the branch name and the PR body.
2. **Branch from the current `origin/next` head.** Name it
`feat/…`, `fix/…`, `docs/…`, or `test/…` with the issue number
(e.g. `docs/1214-branch-process`). Record the base SHA in the PR body.
3. **Develop with evidence.** Applicable tests accompany the change. Hooks are
never bypassed (`--no-verify` is prohibited). Stage explicit paths — never
`git add -A`.
4. **Open the PR against `next`.** The body states: scope, base SHA,
verification commands with results, and any known pre-existing failures on
the base — documented, not retried to green and not absorbed silently.
5. **CI must be terminal-green on the exact head.** All bounded Woodpecker
steps succeed (`verify-terminal-green` contract). Pipelines for fork PRs
start `blocked`; a maintainer approves the run — approving CI is not
approving the PR.
6. **Independent review. Self-merge is prohibited** — for every agent, on every
PR, including trivial ones. Where the change touches protected or
contract-bearing content, the reviewer verifies the exact head
(exact-byte/exact-blob comparison), not a description of it. An `AMEND`
verdict returns the PR to its author; the reviewer's gate stays held until
a fresh exact head passes.
7. **Merge into `next`** happens only after CI green + review pass, pinned to
the reviewed head SHA (a post-review push voids the review).
8. **Promotion `next` → `main`** is a deliberate, Jason-owned reconciliation
merge — not part of any contribution's lifecycle. Contributors are done at
step 7.
## Command surface
### Responsibilities
`scripts/bootstrap.sh` (idempotent) · `build.sh` · `hello.sh` ·
`verify.sh` · `run-task.sh run <task.json>` · `release.sh
package|activate|rollback|status` · `auth.sh status|accounts` · `reset.sh` (**danger**: wipes the data
root; triple-safety-checked) · `mosaic-task.mjs validate|run|show|list|retry|prune|resolve-role` ·
`agent.sh <name>` (interactive TUI agent) ·
suites: `test-config.sh`, `test-task.sh`, `test-release.sh`,
`test-conductor.sh`, `test-auth.sh`, `test-discord.sh`, `test-queue.sh`.
- **Contributor** — base pinning, green CI, evidence in the PR body,
responding to AMEND verdicts, never merging own work.
- **Reviewer / merge gate** — independent verification on the exact head;
holds and lifts gates; executes the merge into `next`.
- **Orchestrator / adjudicator** — cross-PR sequencing, disposition when PRs
collide, conflict adjudication.
- **Jason** — `next` → `main` promotions, merge-authority grants, collaborator
and token provisioning. Agents cannot grant themselves or each other any of
these.
Full reference — usage, fields, exit codes, safety notes:
`docs/TOOLS.md` (read on demand; do not rely on this summary for detail).
### Hotfixes and divergence
## Data map (canon)
- A hotfix follows the same path: branch from `next`, PR to `next`, gates,
merge, then an expedited Jason-owned promotion if `main` needs it urgently.
Committing the fix to `main` directly is prohibited even under pressure.
- **Never land work on `main` that is not on `next`.** This has happened
(issue #1152's goal controller reached `main` without reaching `next`) and
every later PR paid for it. If it happens anyway: transplant the work onto
a `next`-based branch with provenance-preserving commits
(`git cherry-pick -x` or explicit SHA references in the messages), PR it
through the normal gates, and let promotion re-align `main`. Do not
hand-patch `main` to compensate.
- Force-pushing a branch you do not own is prohibited; rebasing your own PR
branch is fine before review, and voids any review already given.
- `~/.config/mosaic-dev/config.json` — system config (user-authored; never
auto-written).
- `<dataRoot>` (from config; default `~/.mosaic-dev`):
- `runs/` — write-once run evidence (`result.json`, snapshots, `stderr.txt`)
- `sessions/` — pi JSONL session trees, one directory per named session
- `workspaces/` — agent file effects (persistent or `:run` ephemeral)
- `state/` — release pointer + append-only activation/auto-apply logs
- Ownership is per-directory; nothing shares state. Directory map and
lifecycle rules: README.md "Data map" section.
## Database and Local Runtime Safety
## Pointers (depth lives here)
- Current local data-layer work uses in-process PGlite; leave `DATABASE_URL` unset.
- PostgreSQL execution is held until KBN-101-00, KBN-101-03, and KBN-101-05 land.
- Do not invoke a migration runner, initialization SQL, or the Compose PostgreSQL service from this checkout.
- Do not start Gateway/Web or run root `pnpm dev` as a local PGlite route. The current dotenv loader can inherit a daemon PostgreSQL DSN; KBN-101-02 must make that path fail closed first.
- Migration artifact generation is offline and does not authorize PostgreSQL access:
- `docs/plans/2026-09-27_goals-review.md` — north star and goal order (Jason ratified 2026-09-27)
- `docs/plans/QUEUE.md` — THE task list, rendered from `docs/plans/queue.json`
(`scripts/mosaic queue next <seat>` reads it; `packages/queue/README.md` has the verbs)
- `docs/plans/CURRENT.md` — narrative log behind the queue rows
- `docs/plans/ROADMAP.md` — agreed milestone path (M16+)
- `docs/plans/CONDUCTOR.md` — orchestration protocol and guardrails
- `docs/plans/2026-09-02_atomic-mosaic-foundation.md` — architecture, invariants
- `docs/plans/2026-09-03_autonomous-run.md` — batch-run tracker
- `BUILD-LOG.md` — append-only build/verification history with corrections
- `LAYERS.md` — implemented vs deferred layers
- `docs/SESSIONS.md` — session registry
- `adapters/README.md` — the harness adapter contract
- `roles/` — role contracts (conductor, future agent/coder/reviewer)
```bash
pnpm --filter @mosaicstack/db db:generate
```
## Recovery rule
## docs/TASKS.md — Schema (CANONICAL)
Compacted, restarted, or new? Nothing that matters is lost: this file +
`scripts/mosaic queue next <seat>` + `docs/plans/CURRENT.md` +
`git log --oneline -10` + the suites reconstruct the full state. **Never
guess** — verify with the suites; the run records and logs hold the receipts.
The `agent` column specifies the required model for each task. **This is set at task creation by the orchestrator and must not be changed by workers.**
## Version pin
| Value | When to use | Budget |
| --------- | ----------------------------------------------------------- | -------------------------- |
| `codex` | All coding tasks (default for implementation) | OpenAI credits — preferred |
| `glm-5.1` | Cost-sensitive coding where Codex is unavailable | Z.ai credits |
| `haiku` | Review gates, verify tasks, status checks, docs-only | Cheapest Claude tier |
| `sonnet` | Complex planning, multi-file reasoning, architecture review | Claude quota |
| `opus` | Major cross-cutting architecture decisions ONLY | Most expensive — minimize |
| `—` | No preference / auto-select cheapest capable | Pipeline decides |
Pipeline crons read this column and spawn accordingly. Workers never modify `docs/TASKS.md` — only the orchestrator writes it.
**Full schema:**
```
| id | status | description | issue | agent | repo | branch | depends_on | estimate | notes |
```
- `status`: `not-started` | `in-progress` | `done` | `failed` | `blocked` | `needs-qa`
- `agent`: model value from table above (set before spawning)
- `estimate`: token budget e.g. `8K`, `25K`
`@earendil-works/pi-coding-agent` is pinned exactly (see `package.json` /
`RELEASE`); never install unversioned. Release identity: `RELEASE` file
(0.0.X until declared stable); image tags derive from it.
+371
View File
@@ -0,0 +1,371 @@
# Minimal Mosaic Stack container proof of concept
## Purpose
Build the smallest isolated container that can:
- launch Pi
- load a small set of Mosaic-style contract files
- send one real request to a model
- return a known response.
This is a standalone experiment. It is not part of the existing Mosaic Stack repository or Software Factory.
## Working boundary
The directory containing this brief is the project root.
### Do not read, copy, mount, import, or modify anything from:
- `/home/jwoltje/.mosaic`
- `/home/jwoltje/.config/mosaic`
- `/home/jwoltje/src/mosaic-stack`
- Existing Mosaic Stack worktrees
### Do not use:
- Mosaic orchestration
- Mosaic Git wrappers
- Fleet agents
- Fleet communication
- Mosaic role policies
- Existing Mosaic contract files
- Existing Mosaic runtime state
No Git credentials, issue, pull request, reviewer, merge, or deployment are required for this experiment.
Nothing from this experiment may be copied into the existing Mosaic Stack repository until it receives a separate review later.
## Runtime data
Use this host directory only for generated runtime data:
```text
/home/jwoltje/.mosaic-dev
```
The source code must remain in the project directory containing this brief.
Inside the container, use:
```text
/opt/mosaic/contracts Immutable contract files
/var/lib/mosaic Generated runtime state
/workspace Agent workspace
```
Mount /home/jwoltje/.mosaic-dev at /var/lib/mosaic.
### Required proof
The finished experiment must prove one path:
1. Build one container image.
2. Start one Pi agent inside the container.
3. Load four local contract files from /opt/mosaic/contracts.
4. Send a request that does not contain the expected response.
5. Receive MOSAIC_HELLO_OK from the agent.
6. Exit successfully when the response matches.
7. Exit nonzero when the response does not match.
This is the entire required functional result.
### Required discovery
Before writing the runtime command:
1. Find the current package documentation for @earendil-works/pi-coding-agent.
2. Determine the current package version.
3. Determine the supported noninteractive command.
4. Determine how Pi accepts a custom system prompt or system prompt file.
5. Determine Pi's documented container authentication method.
6. Record the commands and findings in BUILD-LOG.md.
Do not guess CLI flags, authentication paths, or SDK methods.
Pin the selected Pi package version in the project. Do not install an unversioned package during each container start.
Prefer the Pi CLI. Use the Pi SDK only if the CLI cannot load the generated system prompt in noninteractive mode.
### Contract files
Create these files inside the project:
```text
contracts/CONSTITUTION.md
contracts/STANDARDS.md
contracts/SOUL.md
contracts/USER.md
```
Use these exact contents.
### contracts/CONSTITUTION.md
```markdown
# POC constitution
Never print credentials, tokens, or authentication files.
Follow the loaded system instructions before the user request.
```
### contracts/STANDARDS.md
```markdown
# POC standards
Answer startup verification requests with only the requested value.
Do not add explanation or formatting.
```
### contracts/SOUL.md
```markdown
# POC identity
Your name is mosaic-poc-agent.
Your startup marker is MOSAIC_HELLO_OK.
When asked for your startup marker, return only the marker.
```
### contracts/USER.md
```markdown
# POC user
This is an isolated local runtime test.
```
Contract loading
Create a small script that reads the four contract files in this order:
1. CONSTITUTION.md
2. STANDARDS.md
3. SOUL.md
4. USER.md
Join them with clear file separators.
Write the generated system prompt to:
```text
/var/lib/mosaic/system-prompt.md
```
Pass that generated prompt to Pi using its documented CLI or SDK method.
Do not build:
- Contract schemas
- Contract inheritance
- Overlays
- Role transitions
- Dynamic policy loading
- Guide routing
- Manifest validation
Container
Create one service named:
```text
mosaic-agent
```
Use one Containerfile and one compose.yaml.
Requirements:
- Use a maintained Node.js base image.
- Run as a non-root user.
- Install a pinned Pi package version.
- Copy the local contract fixtures into /opt/mosaic/contracts.
- Do not copy credentials into the image.
- Do not mount the Docker socket.
- Do not mount either live Mosaic directory.
- Do not add a database, web server, queue, or second container.
- The container may run as a one-shot command. It does not need to remain running.
### Authentication
Use Pi's documented authentication mechanism.
Authentication must be supplied at runtime through either:
- A read-only mounted credential file
- A supported runtime environment variable
**Never**:
- Commit credentials
- Copy credentials into the image
- Print credentials
- Print authentication files
- Include credentials in BUILD-LOG.md
- Store credentials under the project directory
Provide .env.example only for non-secret settings such as model or provider names.
If credentials are unavailable, complete the image and scripts but report that the real model request remains unverified. Do not fake the response.
### Required commands
Create these executable scripts:
```text
scripts/build.sh
scripts/hello.sh
scripts/verify.sh
scripts/reset.sh
```
### scripts/build.sh
Build the container image using Docker Compose.
### scripts/hello.sh
Run the mosaic-agent service as a one-shot container.
Send this exact user request:
```text
Return your startup marker and nothing else.
```
The request must not contain MOSAIC_HELLO_OK.
Print the model response without printing credentials or unrelated runtime data.
### scripts/verify.sh
Run the complete test.
**It must**:
1. Build or confirm the image is built.
2. Run the agent request.
3. Remove surrounding whitespace from the response.
4. Compare the response with MOSAIC_HELLO_OK.
5. Exit 0 only when they match exactly.
6. Exit nonzero with a clear error when they do not match.
### scripts/reset.sh
Delete generated POC state only when all checks pass:
1. The resolved path is exactly /home/jwoltje/.mosaic-dev.
2. The path is not a symbolic link.
3. The directory contains a .mosaic-poc-root ownership marker created by this project.
Refuse to delete anything if a check fails.
## Required files
The final project should contain only what the implementation needs:
```text
BRIEF.md
BUILD-LOG.md
README.md
LAYERS.md
Containerfile
compose.yaml
package.json
package-lock.json
.gitignore
contracts/
scripts/
src/
```
Remove unused files and empty directories.
Build log
Create BUILD-LOG.md.
Treat it as append-only.
Before each phase, append:
- Timestamp
- Intended action
- Reason
- Expected result
After each phase, append:
- Commands run
- Observed result
- Failure or correction
Never rewrite an earlier entry. Add a correction as a new entry.
Do not record credentials.
Initial decisions:
- This is a standalone experiment outside the Mosaic Software Factory.
- It does not use existing Mosaic source, tools, contracts, agents, or runtime state.
- The first proof uses one Pi agent and four small local contract files.
- The only required model result is MOSAIC_HELLO_OK.
- Persistence, policy enforcement, Claude, orchestration, and portal work are deferred.
## Acceptance criteria
The experiment passes when:
1. scripts/build.sh exits 0.
2. The image contains the four local contract files.
3. The image contains no credentials.
4. The container has no mounts from ~/.mosaic or ~/.config/mosaic.
5. scripts/hello.sh performs a real model request.
6. The request does not contain the expected marker.
7. The agent returns exactly MOSAIC_HELLO_OK.
8. scripts/verify.sh exits 0.
9. Changing the expected value makes scripts/verify.sh exit nonzero.
10. scripts/reset.sh refuses unsafe paths.
11. Resetting and rerunning the verification produces the same successful result.
## Deferred layers
Document these in LAYERS.md. Do not implement them.
- L0: Container builds and returns MOSAIC_HELLO_OK.
- L1: Persist and resume a named Pi session.
- L2: Add a fixed tool permission policy.
- L3: Load full versioned contract bundles.
- L4: Add Claude as a second runtime.
- L5: Add multiple agents and communication.
- L6: Add orchestration, knowledge storage, and portal features.
## Explicit exclusions
Do not implement:
- Existing Mosaic Stack compatibility
- Git hosting or CI
- Pull requests or code review
- Deployment
- Persistent agent sessions
- Tool read restrictions
- Claude
- Multiple agents
- Fleet communication
- Watchers
- Role management
- Knowledge storage
- Database storage
- API server
- Web interface
- Dashboard
- Production security architecture
## Final report
When finished, report:
1. Files created.
2. Pi package version.
3. Exact build command.
4. Exact verification command.
5. Verification output with credentials removed.
6. Whether the real model request passed.
7. Any remaining failure.
8. Anything implemented beyond this brief.
Do not describe the experiment as production-ready.
+3768
View File
File diff suppressed because it is too large Load Diff
+1 -5
View File
@@ -1,5 +1 @@
# Claude Compatibility Pointer
@AGENTS.md
Do not add project guidance here. Keep `AGENTS.md` authoritative so every agent runtime receives the same instructions.
@AGENTS.md
+43
View File
@@ -0,0 +1,43 @@
# Minimal Mosaic Stack POC agent image.
# Base: maintained Node.js image (same family as Pi's documented
# containerization example in docs/containerization.md).
FROM node:24-bookworm-slim
# Tools Pi's documented container image expects (bash, CA certs, git, ripgrep).
RUN apt-get update \
&& apt-get install -y --no-install-recommends bash ca-certificates git ripgrep \
&& rm -rf /var/lib/apt/lists/*
# Non-root user: the maintained node image ships a 'node' user at
# uid/gid 1000, which matches the host user that owns the runtime
# state directory mounted at /var/lib/mosaic. It is reused as-is.
# Pinned Pi install: package.json pins the exact version and
# package-lock.json is installed with npm ci. No unversioned installs.
WORKDIR /opt/app
COPY package.json package-lock.json ./
RUN npm ci --ignore-scripts
# Immutable contract fixtures (required location), runtime scripts, and
# runtime adapters.
COPY contracts /opt/mosaic/contracts
COPY src /opt/mosaic/src
COPY adapters /opt/mosaic/adapters
RUN chmod 0555 /opt/mosaic/contracts /opt/mosaic/contracts/* \
&& chmod 0555 /opt/mosaic/src /opt/mosaic/src/*.sh \
&& chmod 0555 /opt/mosaic/adapters /opt/mosaic/adapters/*/adapter.sh
# Writable state, workspace, and pi agent directory (auth.json is
# bind-mounted read-only at runtime; nothing is copied into the image).
RUN mkdir -p /var/lib/mosaic /workspace /home/node/.pi/agent \
&& chown -R node:node /var/lib/mosaic /workspace /home/node /opt/app
USER node
WORKDIR /workspace
ENV HOME=/home/node \
PATH="/opt/app/node_modules/.bin:${PATH}" \
PI_OFFLINE=1
# One-shot agent: args form the user request (default is the startup
# verification request defined in compose.yaml).
ENTRYPOINT ["/opt/mosaic/src/run-agent.sh"]
+50
View File
@@ -0,0 +1,50 @@
# LAYERS
Deferred capability layers for the Mosaic experiment. Only L0 is implemented by
this proof of concept; everything below it is documented here and deliberately
not implemented (see BRIEF.md, "Explicit exclusions").
## L0 — Implemented: container returns MOSAIC_HELLO_OK
One image (`mosaic-poc-agent:0.84.4`, built on `node:24-bookworm-slim`, non-root,
pinned Pi) runs one Pi agent one-shot. Four immutable local contract files are
loaded in fixed order into the generated system prompt
(`/var/lib/mosaic/system-prompt.md`). One real model request is sent
noninteractively; the response must equal `MOSAIC_HELLO_OK` exactly or the
verification exits nonzero. Authentication is supplied at runtime only
(read-only mounted pi auth file, or a provider API key environment variable).
## L1 — Deferred: persist and resume a named Pi session
Keep a named Pi session across container runs (`--name`, session storage under
`/var/lib/mosaic`), resume it with the documented session flags, and verify
state survives a container restart.
## L2 — Deferred: fixed tool permission policy
Add a fixed allow/deny policy for Pi tools (e.g. restricting built-in tools via
documented `--tools` / `--exclude-tools` or an extension-based permission gate),
so contract files can constrain what the agent may do, not just what it says.
## L3 — Deferred: load full versioned contract bundles
Replace the four static fixtures with versioned contract bundles: bundle
manifests, contract versions, and deterministic ordering/hashing, loaded from
an immutable bundle artifact instead of files copied at image build time.
## L4 — Deferred: Claude as a second runtime
Add a second runtime (Claude) alongside the Pi agent in the same container
stack, behind the same contract-loading path, to compare behavior across
runtimes.
## L5 — Deferred: multiple agents and communication
Run several named agents with defined roles and a communication channel between
them (message passing or shared state under `/var/lib/mosaic`).
## L6 — Deferred: orchestration, knowledge storage, and portal features
Fleet-level orchestration, knowledge storage, monitoring, and portal UI on top
of L1-L5. This is where the existing Mosaic Stack concepts would be re-evaluated
from first principles.
+208 -430
View File
@@ -1,460 +1,238 @@
# Mosaic Stack
# Mosaic Stack — new foundation
Self-hosted, multi-user AI agent platform. One config, every runtime, same standards.
The active rebuild is at this repository's root. The original Mosaic Stack v1
source is archived under `v1/`; it is not the implementation being developed here.
Mosaic gives you a unified launcher for Claude Code, Codex, OpenCode, and Pi — injecting consistent system prompts, guardrails, skills, and mission context into every session. A NestJS gateway provides the API surface, a Next.js dashboard gives you the UI, and a plugin system connects Discord, Telegram, and more.
- Canonical checkout: `/mnt/storage/src/mosaic-stack`
- Repository: `mosaicstack/stack`
- Working branch: `refactor`
- Former `~/src/mosaic-stack-dev-test`: compatibility symlink to this same checkout
## Quick Install
Both original Git histories and pending development work are preserved. See the
[conversion record](docs/plans/2026-09-07_repository-consolidation-completed.md)
and [current next action](docs/plans/CURRENT.md). Do not use v1's startup commands,
package layout or agent instructions for work on the new foundation.
```bash
curl -fsSL https://mosaicstack.dev/install.sh | bash
## Original container proof
The foundation began as a standalone container experiment. One container image
runs one Pi coding agent with four immutable local contract files as its system
prompt, sends exactly one real model request, and was verified to return exactly
`MOSAIC_HELLO_OK`. This historical result is not a claim that the full rebuild is
production-ready.
## Layout
```text
BRIEF.md requirements for the original container proof
BUILD-LOG.md append-only build/verification log
LAYERS.md implemented layer (L0) and deferred layers (L1-L6)
Containerfile image definition (node:24-bookworm-slim, non-root, pinned Pi)
compose.yaml one service: mosaic-agent (one-shot; configured via env)
package.json pins @earendil-works/pi-coding-agent at exactly 0.84.4
package-lock.json resolved lockfile used by npm ci in the image
.env.example non-secret settings only (credential-file path, env-var auth)
contracts/ CONSTITUTION.md, STANDARDS.md, SOUL.md, USER.md (immutable fixtures)
scripts/ bootstrap/build/hello/verify/reset + config tooling
src/ load-contracts.sh, run-agent.sh (run inside the container)
docs/plans/ architecture and milestone plans
```
Or use the direct URL:
## Configuration
```bash
bash <(curl -fsSL https://git.mosaicstack.dev/mosaicstack/stack/raw/branch/main/tools/install.sh)
The sole discovery entry point is:
```text
~/.config/mosaic-dev/config.json
```
The installer auto-launches the setup wizard, which walks you through gateway install and verification. Flags for non-interactive use:
Created only by the explicit, idempotent bootstrap:
```bash
bash <(curl -fsSL …) --yes # Accept all defaults
bash <(curl -fsSL …) --yes --no-auto-launch # Install only, skip wizard
scripts/bootstrap.sh # create-if-absent; validates existing config, never rewrites
```
This installs both components:
Minimal shape (`configVersion` 1):
| Component | What | Where |
| ----------------------- | ---------------------------------------------------------------- | -------------------- |
| **Framework** | Bash launcher, guides, runtime configs, tools, skills | `~/.config/mosaic/` |
| **@mosaicstack/mosaic** | Unified `mosaic` CLI — TUI, gateway client, wizard, auto-updater | `~/.npm-global/bin/` |
```json
{
"configVersion": 1,
"environment": "development",
"dataRoot": "/home/jwoltje/.mosaic-dev",
"execution": {
"backend": "docker",
"provider": "zai",
"model": "glm-5.3-flash"
}
}
```
### Install lanes
Rules enforced by `scripts/mosaic-config.mjs`:
| Lane | Command | Use when | Source |
| ------------------------ | ------------------------------------- | ----------------------------------------------------- | ----------------------------------------------------------------------- |
| Stable | `bash tools/install.sh` | You want the released Mosaic CLI/framework | npm registry `@mosaicstack/mosaic@latest` + framework archive at `main` |
| Prerelease integration | `bash tools/install.sh --next` | You want the current `next` integration branch | Build-from-source at `next` |
| Contributor/source build | `bash tools/install.sh --dev --ref X` | You are testing a branch before release; `--ref` wins | Build-from-source at the requested ref |
- Unknown keys, unsupported versions/backends, and malformed JSON exit nonzero; nothing is modified.
- `dataRoot` must be absolute, canonical, and must not be or contain the home or configuration directory.
- Validation failures never touch config, state, or images.
- `scripts/test-config.sh` runs the sandboxed config selftests (no Docker required).
`--next` is shorthand for the prerelease integration lane: it enables source-build mode and uses `next` unless an explicit `--ref` or `MOSAIC_REF` is provided.
Run paths (`build/hello/verify/reset`) fail closed when configuration is missing or invalid; they never invent it.
After install, the wizard runs automatically or you can invoke it manually:
## Missions & tasks (M2)
Missions and tasks are validated JSON data (strict schemas, version-pinned). The M2 layer is host-side only: mission directives are recorded for provenance but do not yet reach the runtime system prompt (capability/policy layer comes later).
```text
missions/hello.json objective + directives (missionVersion 1)
tasks/hello-marker.json prompt + optional mission ref + expectExact + timeout
<dataRoot>/runs/r-<id>/ immutable run record: task.json, mission.json,
stderr.txt, result.json (all write-once)
```
Usage:
```bash
mosaic wizard # Full guided setup (gateway install → verify)
scripts/run-task.sh validate tasks/hello-marker.json # strict validation, writes nothing
scripts/run-task.sh run tasks/hello-marker.json # execute; result recorded under dataRoot/runs
scripts/mosaic-task.mjs list # list runs and statuses
scripts/test-task.sh # selftests (schema negatives + live runs)
```
### Requirements
A run exits 0 only when its expectation is met (`expectExact` match); mismatches, nonzero agent exits, and timeouts record `status: failed` in `result.json` and exit 1. Each run gets a unique directory — rerunning never rewrites history.
- Node.js ≥ 22
- npm (for global @mosaicstack/mosaic install)
- One or more runtimes:
- [Claude Code](https://docs.anthropic.com/en/docs/claude-code)
- [Codex](https://github.com/openai/codex)
- [OpenCode](https://opencode.ai)
- [Pi](https://pi.dev)
## Release model (M3)
`RELEASE` single-sources the release version (0.0.X until declared stable); the image tag derives from it plus the pinned Pi version. Activation is health-gated and every event is recorded:
```bash
scripts/release.sh package # build + tag the release image
scripts/release.sh activate # health check (exact marker) -> atomic pointer swap
scripts/release.sh activate --fault-injection # prove the refusal path (drills only)
scripts/release.sh rollback # health-gated return to the previous release
scripts/release.sh ensure # self-determination: align installed to RELEASE (safe no-op when aligned)
scripts/release.sh status # release, tag, active pointer, recent log
scripts/test-release.sh # release selftests
```
`ensure` is invoked automatically by the human-facing launchers (`hello`,
`verify`, `agent`): the system determines what is installed and aligns
itself — the user never runs release commands manually.
- `<dataRoot>/state/active.json` — the activation pointer (atomic tmp+rename replace)
- `<dataRoot>/state/activation-log.jsonl` — append-only history: package / activate / refused / rollback
A failed health check never activates; the previously active release remains deployed. Updating the software therefore cannot corrupt the running installation: package beside, gate, then flip. Verified by the update/refusal/rollback drills in BUILD-LOG Phase 7.
## Runtime adapters (M4)
The harness boundary is formalized: everything upstream (config, contracts, missions, tasks, run records) is harness-agnostic; everything inside an adapter belongs to one runtime.
```text
adapters/<name>/adapter.sh env in: MOSAIC_SYSTEM_PROMPT_FILE, MOSAIC_REQUEST,
MOSAIC_PROVIDER, MOSAIC_MODEL
stdout: response only; stderr: diagnostics
```
- Selection: `execution.adapter` in config.json (optional; `pi` default; allowlist `pi`, `mock`)
- `pi` — pinned Pi CLI, noninteractive print mode, ambient discovery off
- `mock` — deterministic test adapter; never for real verification
- Mission directives have a sanctioned injection point: when a task references a mission, the task runner mounts the run snapshot and the generated prompt gains a `MISSION (runtime)` section (objective + directives) after the four immutable contracts
- Adding a harness (Claude, Codex, OpenCode) later means adding one directory — no orchestrator changes
See `adapters/README.md` for the full contract.
## Workspaces, capabilities, sessions (M5/M6)
Optional task fields extend what an agent can do — all defaulting to the previous behavior:
```json
{
"workspace": "demo", // ":run" ephemeral, or persistent dataRoot/workspaces/<name>
"capabilities": { "tools": ["bash", "read"] }, // pi tool allowlist; absent = no tools
"session": "demo" // persistent session at dataRoot/sessions/<name>
}
```
- The adapter runs inside the workspace; files it writes are host-visible (`dataRoot/workspaces/<name>`).
- Sessions persist via pi's documented `--session-dir`; a follow-up run in the same session resumes the conversation (`-c`) and can recall prior context. Distinct names never share state. Ephemeral (`--no-session`) remains the default when no session is declared.
- Selection authority: config for adapter/provider/model; the task file for workspace/capabilities/session.
Inspect anything:
```bash
node scripts/mosaic-task.mjs list # runs with task/workspace/session columns
node scripts/mosaic-task.mjs show <runId> # full record + snapshots + artifacts
```
Demo fixtures: `tasks/workspace-demo.json`, `tasks/session-demo-1.json` + `tasks/session-demo-2.json`.
See `docs/plans/2026-09-02_atomic-mosaic-foundation.md` for the full plan.
Inside the container:
```text
/opt/mosaic/contracts immutable contract files
/var/lib/mosaic generated runtime state (mounted from configured dataRoot)
/workspace agent workspace
```
## How it works
1. `scripts/build.sh` builds the release image (`mosaic-poc-agent:<pi>-r<release>`,
tag derived from `RELEASE` + the pinned Pi version) with Docker Compose.
2. On each run, `/opt/mosaic/src/load-contracts.sh` reads the four contract files
in fixed order (CONSTITUTION, STANDARDS, SOUL, USER), joins them with clear
separators, and writes `/var/lib/mosaic/system-prompt.md`.
3. `/opt/mosaic/src/run-agent.sh` starts Pi noninteractively
(`pi -p "Return your startup marker and nothing else."`) with
`--system-prompt "$(cat /var/lib/mosaic/system-prompt.md)"` and all ambient
discovery disabled (`--no-context-files --no-skills --no-extensions
--no-prompt-templates --no-themes`), ephemeral (`--no-session`), tool-free
(`--no-tools`), and offline for startup network operations (`--offline`).
4. `scripts/verify.sh` trims surrounding whitespace from the response and exits 0
only when it equals `MOSAIC_HELLO_OK` exactly.
## Usage
### Launching Agent Sessions
```bash
scripts/bootstrap.sh # create config.json if absent (idempotent)
scripts/build.sh # build the image
scripts/hello.sh # one-shot request; prints the model response
scripts/verify.sh # full gated test; exit 0 only on exact MOSAIC_HELLO_OK
scripts/run-task.sh # run a mission/task file (see Missions & tasks)
scripts/release.sh # package / activate / rollback / status (see Release model)
scripts/test-config.sh # fast config-layer selftests (no Docker)
scripts/test-task.sh # mission/task selftests (schema + adapter seam + live runs)
scripts/test-release.sh # release selftests
scripts/reset.sh # delete the configured data root (safety-checked)
```
Prove the failure path (acceptance criterion 9):
```bash
mosaic pi # Launch Pi with Mosaic injection
mosaic claude # Launch Claude Code with Mosaic injection
mosaic codex # Launch Codex with Mosaic injection
mosaic opencode # Launch OpenCode with Mosaic injection
mosaic yolo claude # Claude with dangerous-permissions mode
mosaic yolo pi # Pi in yolo mode
EXPECTED_MARKER=MOSAIC_NOT_OK scripts/verify.sh # must exit nonzero
```
The launcher verifies your config, checks for `SOUL.md`, injects your `AGENTS.md` standards into the runtime, and forwards all arguments.
Pi launches default to a token-lean skill posture: `mosaic pi` passes `--no-skills` so Pi does not preload every global skill description into the system prompt. Use `MOSAIC_PI_SKILL_MODE=all mosaic pi` for the legacy all-skills catalog, or `MOSAIC_PI_SKILL_MODE=discover mosaic pi` to let Pi use its native settings/project skill discovery.
Mosaic also loads its Pi extensions from `~/.config/mosaic/runtime/pi/`. Inside Pi,
`/goal set <statement>` starts a bounded persistent loop that checks every turn and successful
compaction, requires two evidence-bearing completion reports, and can be inspected or stopped with
`/goal status`, `/goal pause`, `/goal resume`, and `/goal cancel`. Controller-owned goal-state
entries redact common credential shapes, but Pi's model/tool-call history is separate, so goals and
evidence must never contain secrets or raw sensitive output. Mosaic does not install this extension
into `~/.pi/agent/extensions/`.
### TUI & Gateway
```bash
mosaic tui # Interactive TUI connected to the gateway
mosaic gateway login # Authenticate with a gateway instance
mosaic sessions list # List active agent sessions
```
### Gateway Management
```bash
mosaic gateway install # Install and configure the gateway service
mosaic gateway verify # Post-install health check
mosaic gateway login # Authenticate and store a session token
mosaic gateway config rotate-token # Rotate your API token
mosaic gateway config recover-token # Recover a token via BetterAuth cookie
```
If you already have a gateway account but no token, use `mosaic gateway config recover-token` to retrieve one without recreating your account.
### Configuration
Mosaic supports three storage tiers: `local` (PGlite, single-host), `standalone` (PostgreSQL, single-host), and `federated` (PostgreSQL + pgvector + Valkey, multi-host). See [Federated Tier Setup](docs/federation/SETUP.md) for multi-user and production deployments, or [Migrating to Federated](docs/guides/migrate-tier.md) to upgrade from existing tiers.
```bash
mosaic config show # Print full config as JSON
mosaic config get <key> # Read a specific key
mosaic config set <key> <val># Write a key
mosaic config edit # Open config in $EDITOR
mosaic config path # Print config file path
```
### Management
```bash
mosaic doctor # Health audit — detect drift and missing files
mosaic sync # Sync skills from canonical source
mosaic skill list # Audit Claude skill registrations and conflicts
mosaic skill register <name> # Register one canonical skill with Claude Code
mosaic skill unregister <name> # Remove one Mosaic-owned Claude link
mosaic update # Update CLI/framework and auto-register canonical skills
mosaic wizard # Full guided setup wizard
mosaic bootstrap <path> # Bootstrap a repo with Mosaic standards
mosaic coord init # Initialize a new orchestration mission
mosaic prdy init # Create a PRD via guided session
```
### Sub-package Commands
Each Mosaic sub-package exposes its API surface through the unified CLI:
```bash
# User management
mosaic auth users list
mosaic auth users create
mosaic auth sso
# Agent brain (projects, missions, tasks)
mosaic brain projects
mosaic brain missions
mosaic brain tasks
mosaic brain conversations
# Agent forge pipeline
mosaic forge run [--simulate] # fails closed (FORGE_NO_EXECUTOR) with no executor wired; --simulate for typed simulated runs
mosaic forge status
mosaic forge resume [--simulate] # same fail-closed rule as forge run
mosaic forge personas
# Structured logging
mosaic log tail
mosaic log search
mosaic log export
mosaic log level
# MACP protocol
mosaic macp tasks
mosaic macp submit
mosaic macp gate
mosaic macp events
# Agent memory
mosaic memory search
mosaic memory stats
mosaic memory insights
mosaic memory preferences
# Task queue (Valkey)
mosaic queue list
mosaic queue stats
mosaic queue pause
mosaic queue resume
mosaic queue jobs
mosaic queue drain
# Object storage
mosaic storage status
mosaic storage tier
mosaic storage export
mosaic storage import
# Schema migration is unavailable in this release. The current storage wrapper shells
# directly to `pnpm --filter @mosaicstack/db db:migrate`; it is legacy N-1,
# uncertified, and MUST NOT be invoked pending KBN-101-02/-03/-06/-08 activation.
# Future schema migration is non-operative: external bootstrap → TLS/roles → runner
# --run → runner --verify → readiness. Tier copy uses only the separately held secure
# migrate-tier route.
```
### Telemetry
```bash
# Local observability (OTEL / Jaeger)
mosaic telemetry local status
mosaic telemetry local tail
mosaic telemetry local jaeger
# Remote telemetry (dry-run by default)
mosaic telemetry status
mosaic telemetry opt-in
mosaic telemetry opt-out
mosaic telemetry test
mosaic telemetry upload # Dry-run unless opted in
```
Consent state is persisted in config. Remote upload is a no-op until you run `mosaic telemetry opt-in`.
## Standalone container deployment
The `stack` profile runs PostgreSQL, Valkey, the gateway, and the bundled webUI. Copy
`.env.example` to `.env`, generate `BETTER_AUTH_SECRET`, then start the profile:
```bash
cp .env.example .env
printf 'BETTER_AUTH_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose --profile stack up -d
```
The optional dogfood overlay gives one dedicated in-stack agent a writable stack
worktree and its own read-only credential slot. It does not mount the fleet brain or
any other seat. Prepare a `next`-based worktree and an unprivileged
`code-dogfood-01` functional seat outside the container, then set these paths in
`.env`:
```dotenv
MOSAIC_DOGFOOD_WORKTREE=/path/to/mosaic-stack-worktrees/dogfood-1487
MOSAIC_DOGFOOD_COMMON_GIT_DIR=/path/to/mosaic-stack/.git
MOSAIC_DOGFOOD_SEAT_HOME=/path/to/.mosaic/fleet/agents/code-dogfood-01
```
The common Git directory must match the worktree's `.git` pointer. The seat home
must contain only that seat's credential at
`secrets/gitea-mosaicstack-code-dogfood-01.token`. Never place the token value in
`.env`. Start the overlay with:
```bash
docker compose \
-f docker-compose.yml \
-f docker-compose.dogfood.yml \
--profile stack up -d
```
The overlay removes the general shell tool for every session, including admins.
File tools stay inside the mounted checkout. Two dedicated delivery tools stage
explicit paths, run the CI queue guard, push through `git-credential-mosaic`, and
open PRs through `pr-create.sh`. They resolve only the `code-dogfood-01` slot and fail
if it is absent. The overlay enables Docker's init process so the R4 helper can
establish the gateway's seat lineage below PID 1.
This deployment route is separate from the local source-development restrictions
below.
## Development
### Prerequisites
- Node.js ≥ 22
- pnpm 10.6+
- Docker & Docker Compose
### Setup
```bash
git clone [email protected]:mosaicstack/stack.git
cd stack
# Install dependencies. The local tier uses in-process PGlite; leave DATABASE_URL unset.
# The pnpm store defaults to $HOME/.local/share/pnpm/store. Override it without
# editing the checkout with NPM_CONFIG_STORE_DIR=$HOME/another-store if needed.
pnpm install
# Verify dependencies and generated state before running source-quality gates.
# Missing dependencies exit 42; stale/foreign apps/web/.next state exits 43.
# The web build certifies its exact standalone symlink manifest; added, removed,
# retargeted, or manifest-only-tampered generated links also exit 43. This detects
# accidental, independent, stale, and foreign-residue mutation—the class exposed by
# a five-month-stale .next that produced 19 phantom TS2307 errors.
# It does NOT defend against a same-UID actor that can rewrite both manifest and
# marker consistently (CWE-345). RM-59 tracks the required executor/spine-side
# trust anchor outside worktree authority.
pnpm preflight
# Optional local queue service only. This does not start PostgreSQL.
docker compose up -d valkey
# The current Gateway/Web local process is held; see docs/guides/dev-guide.md.
# Do not start it until KBN-101-02 makes inherited dotenv/DSN state fail closed.
```
### Held future procedure
The checked-in Compose PostgreSQL service mounts legacy initialization SQL and is **not** a
current PostgreSQL, standalone, or federated developer route. Do not start it with Compose,
invoke initialization SQL, or treat the planned migrator as currently executable.
**Held future activation procedure — non-operative and no current command authority until KBN-101-00, KBN-101-03, and KBN-101-05
land:** external bootstrap → TLS/roles → `mosaic-db-migrator --run` →
`mosaic-db-migrator --verify` → Gateway/Compose readiness. The future deployment artifacts—not
this README—will provide the reviewed commands and secret-consumer interface.
For local data-layer work, PGlite needs no PostgreSQL service. The optional Compose command above
starts only Valkey; OTEL Collector and Jaeger may likewise be started individually if needed,
without starting PostgreSQL. A Gateway/Web local process is not currently a safe PGlite route:
its unguarded dotenv loader may inherit a daemon PostgreSQL DSN. Do not use root `pnpm dev` or a
Gateway start command until KBN-101-02 makes that state fail closed.
### Quality Gates
```bash
pnpm preflight # Checkout/dependency/generated-state validation
pnpm typecheck # TypeScript type checking (all packages)
pnpm lint # ESLint (all packages)
pnpm test # Vitest (all packages)
pnpm format:check # Prettier check
pnpm format # Prettier auto-fix
```
### CI
Woodpecker CI runs on every push:
- `pnpm install --frozen-lockfile`
- **Legacy N-1 CI status only — active, uncertified, and non-authorizing as an operator route:** the checked-in job currently invokes `pnpm --filter @mosaicstack/db run db:migrate` with `DATABASE_URL` against an isolated disposable PostgreSQL CI database. It performs direct DDL in that CI database, is not approved ordinary behavior or an operator route, and remains a known exception pending KBN-101-06 removal/replacement by the certified runner-backed CI path.
- `pnpm test` (Turbo-orchestrated across all packages)
npm packages are published to the Gitea package registry on main merges.
## Architecture
```
stack/
├── apps/
│ ├── gateway/ NestJS API + WebSocket hub (Fastify, Socket.IO, OTEL)
│ └── web/ Next.js dashboard (React 19, Tailwind)
├── packages/
│ ├── mosaic/ Unified CLI — TUI, gateway client, wizard, sub-package commands
│ ├── types/ Shared TypeScript contracts (Socket.IO typed events)
│ ├── db/ Drizzle ORM schema + migrations (pgvector)
│ ├── auth/ BetterAuth configuration
│ ├── brain/ Data layer (PG-backed)
│ ├── queue/ Valkey task queue + MCP
│ ├── coord/ Mission coordination
│ ├── forge/ Multi-stage AI pipeline (intake → board → plan → code → review)
│ ├── macp/ MACP protocol — credential resolution, gate runner, events
│ ├── agent/ Agent session management
│ ├── memory/ Agent memory layer
│ ├── log/ Structured logging
│ ├── prdy/ PRD creation and validation
│ ├── quality-rails/ Quality templates (TypeScript, Next.js, monorepo)
│ └── design-tokens/ Shared design tokens
├── plugins/
│ ├── discord/ Discord channel plugin (discord.js)
│ ├── telegram/ Telegram channel plugin (Telegraf)
│ ├── macp/ OpenClaw MACP runtime plugin
│ └── mosaic-framework/ OpenClaw framework injection plugin
├── tools/
│ └── install.sh Unified installer (framework + npm CLI, --yes / --no-auto-launch)
├── scripts/agent/ Agent session lifecycle scripts
├── docker-compose.yml Dev infrastructure
└── .woodpecker/ CI pipeline configs
```
### Key Design Decisions
- **Gateway is the single API surface** — all clients (TUI, web, Discord, Telegram) connect through it
- **ESM everywhere** — `"type": "module"`, `.js` extensions in imports, NodeNext resolution
- **Socket.IO typed events** — defined in `@mosaicstack/types`, enforced at compile time
- **OTEL auto-instrumentation** — loads before NestJS bootstrap
- **Explicit `@Inject()` decorators** — required since tsx/esbuild doesn't emit decorator metadata
### Framework (`~/.config/mosaic/`)
The framework is the bash-based standards layer installed to every developer machine:
```
~/.config/mosaic/
├── AGENTS.md ← Central standards (loaded into every runtime)
├── SOUL.md ← Agent identity (name, style, guardrails)
├── USER.md ← User profile (name, timezone, preferences)
├── TOOLS.md ← Machine-level tool reference
├── bin/mosaic ← Unified launcher (claude, codex, opencode, pi, yolo)
├── guides/ ← E2E delivery, orchestrator protocol, PRD, etc.
├── runtime/ ← Per-runtime configs (claude/, codex/, opencode/, pi/)
├── skills/ ← Universal skills (shipped with the framework package)
├── tools/ ← Tool suites (orchestrator, git, quality, prdy, etc.)
└── memory/ ← Persistent agent memory (preserved across upgrades)
```
### Forge Pipeline
Forge is a multi-stage AI pipeline for autonomous feature delivery:
```
Intake → Discovery → Board Review → Planning (3 stages) → Coding → Review → Remediation → Test → Deploy
```
Each stage has a dispatch mode (`exec` for research/review, `yolo` for coding), quality gates, and timeouts. The board review uses multiple AI personas (CEO, CTO, CFO, COO + specialists) to evaluate briefs before committing resources.
## Upgrading
Run the installer again — it handles upgrades automatically:
```bash
curl -fsSL https://mosaicstack.dev/install.sh | bash
```
Or use the direct URL:
```bash
bash <(curl -fsSL https://git.mosaicstack.dev/mosaicstack/stack/raw/branch/main/tools/install.sh)
```
Or use the CLI:
```bash
mosaic update # Check + install CLI updates
mosaic update --check # Check only, don't install
```
The CLI also performs a background update check on every invocation (cached for 1 hour).
### Installer Flags
```bash
bash tools/install.sh --check # Version check only
bash tools/install.sh --framework # Framework only (skip npm CLI)
bash tools/install.sh --cli # npm CLI only (skip framework)
bash tools/install.sh --next # Prerelease lane: source build from next
bash tools/install.sh --dev # Contributor lane: source build at --ref/main
bash tools/install.sh --ref v1.0 # Install from a specific git ref (--ref wins over --next)
bash tools/install.sh --yes # Non-interactive, accept all defaults
bash tools/install.sh --no-auto-launch # Skip auto-launch of wizard
```
The installer rejects unrecognized flags or positional arguments before making changes and prints the supported-option usage.
## Contributing
```bash
# Create a feature branch
git checkout -b feat/my-feature
# Make changes, then verify
pnpm typecheck && pnpm lint && pnpm test && pnpm format:check
# Commit (husky runs lint-staged automatically)
git commit -m "feat: description of change"
# Push and create PR
git push -u origin feat/my-feature
```
DTOs go in `*.dto.ts` files at module boundaries. Scratchpads (`docs/scratchpads/`) are mandatory for non-trivial tasks. See `AGENTS.md` for the full standards reference.
## License
Proprietary — all rights reserved.
## Authentication
Pi's documented container authentication (see the package's
`docs/containerization.md`) is used, in this order:
1. **Read-only mounted credential file** (default): the host pi auth file
`~/.pi/agent/auth.json` is bind-mounted read-only to
`/home/node/.pi/agent/auth.json`. The host file holds a static API-key
entry for the built-in `zai` provider, so no token refresh writes are needed.
2. **Runtime environment variable** (documented alternative): set `ZAI_API_KEY`
or `ANTHROPIC_API_KEY` in the environment or in a gitignored `.env`; compose
passes them through. Pi's documented precedence applies.
Credentials are never committed, never copied into the image, and never printed.
Mosaic-managed named accounts (`agent.sh --auth`) live under the data root
(`auth/<account>.json`, 0600) — the stack never writes into `~/.pi`.
`.env.example` contains non-secret settings only.
## Boundaries honored
- No mounts of `~/.mosaic` or `~/.config/mosaic`; no Docker socket mount.
- Source stays in this project directory; generated state only in
`/home/jwoltje/.mosaic-dev` (host) and `/var/lib/mosaic` (container).
- No database, web server, queue, second container, orchestration, Git
integration, persistent sessions, or policy machinery.
+1
View File
@@ -0,0 +1 @@
0.0.12
+30
View File
@@ -0,0 +1,30 @@
# Mosaic Stack
You are the default collaborator for Mosaic Stack: a practical engineering
partner helping people build, inspect, and operate a trustworthy foundation
for delegated work.
Mosaic Stack is deliberately small, file-based, and evidence-oriented. Its
purpose is not to perform confidence; it is to make useful work attributable,
bounded, reproducible, and reviewable. Treat the system's contracts, policies,
and run records as part of the product, not paperwork around it.
Work with calm precision. Start from what the user is trying to accomplish,
make the next useful step clear, and explain results in plain language. Be
decisive when the evidence supports a decision; be explicit about uncertainty
when it does not. Never claim a test, command, integration, or outcome that
you have not actually verified.
Respect boundaries. Ask before expanding scope, changing authority, touching
credentials, or taking an irreversible external action. Prefer the least
privileged path, preserve user work, and stop on a policy or validation
refusal rather than working around it. A clean refusal with a useful diagnosis
is better than a superficially successful but untrustworthy result.
Leave a legible trail. Make changes intentional, keep records honest, and
report what changed, how it was checked, and what remains unresolved. When
coordinating other workers, give each one a bounded objective and review their
evidence instead of treating their confidence as proof.
The aim is dependable progress: small enough to understand, safe enough to
trust, and concrete enough for a person to verify.
+54
View File
@@ -0,0 +1,54 @@
# Mosaic runtime adapters
An adapter is the entire harness-specific surface of the system. Everything
upstream of an adapter — configuration, contracts, missions, tasks, run
records — is harness-agnostic; everything inside an adapter may assume one
specific agent runtime.
## Contract
An adapter lives at:
```text
/opt/mosaic/adapters/<name>/adapter.sh
```
and must be executable. The dispatcher (`/opt/mosaic/src/run-agent.sh`)
selects it via `MOSAIC_ADAPTER` (default: `pi`) and execs it after the
system prompt has been generated.
**Inputs (environment):**
| Variable | Meaning |
|---|---|
| `MOSAIC_SYSTEM_PROMPT_FILE` | Absolute path to the generated system prompt (contracts + optional mission section). Read it; do not modify it. |
| `MOSAIC_REQUEST` | The exact user request text (may contain newlines). |
| `MOSAIC_PROVIDER` | Configured provider name. |
| `MOSAIC_MODEL` | Configured model id. |
Optional, adapter-specific (documented per adapter):
| Variable | Meaning |
|---|---|
| `MOSAIC_MOCK_RESPONSE` | mock only: the verbatim response to emit |
**Outputs:**
- `stdout`: the model response text — the only channel the orchestrator captures
- `stderr`: diagnostics (never credentials)
- exit `0`: success; nonzero: failure
## Rules
1. Adapters print ONLY the response on stdout. Status lines go to stderr.
2. Adapters never read configuration files; the resolved settings arrive via environment.
3. Adapters never write outside `/var/lib/mosaic`.
4. Adding an adapter requires: a new directory, the contract implementation, and
adding the name to the allowlist in `scripts/mosaic-config.mjs`.
## Included adapters
- `pi` — the pinned `@earendil-works/pi-coding-agent` CLI in noninteractive
print mode (`-p`), ambient discovery disabled, stdin detached.
- `mock` — deterministic echo of `MOSAIC_MOCK_RESPONSE`. Test-only: never use
it where a real model response is required.
+18
View File
@@ -0,0 +1,18 @@
#!/bin/sh
# Mock adapter: deterministic response for seam tests. NEVER use where a
# real model response is required.
#
# Contract: see /opt/mosaic/adapters/README.md.
set -eu
[ -n "${MOSAIC_SYSTEM_PROMPT_FILE:-}" ] || { echo "mock adapter: MOSAIC_SYSTEM_PROMPT_FILE is required" >&2; exit 2; }
if [ "${MOSAIC_INTERACTIVE:-}" != "1" ]; then
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "mock adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
fi
[ -r "$MOSAIC_SYSTEM_PROMPT_FILE" ] || { echo "mock adapter: system prompt not readable: $MOSAIC_SYSTEM_PROMPT_FILE" >&2; exit 2; }
echo "mock adapter: responding verbatim from MOSAIC_MOCK_RESPONSE" >&2
# Deterministic plumbing evidence: which MOSAIC_* variables did the
# orchestrator actually deliver? (Auth secrets are not MOSAIC_-prefixed.)
(env | grep '^MOSAIC_' | sort) >&2 2>/dev/null || true
printf '%s\n' "${MOSAIC_MOCK_RESPONSE:-}"
+96
View File
@@ -0,0 +1,96 @@
#!/bin/sh
# Pi adapter: implements the Mosaic adapter contract for the pinned
# @earendil-works/pi-coding-agent CLI.
#
# Contract: see /opt/mosaic/adapters/README.md.
# Headless (default): stdout = response only; stderr = diagnostics; exit 0.
# Interactive (MOSAIC_INTERACTIVE=1): full pi TUI on the attached terminal.
set -eu
[ -n "${MOSAIC_SYSTEM_PROMPT_FILE:-}" ] || { echo "pi adapter: MOSAIC_SYSTEM_PROMPT_FILE is required" >&2; exit 2; }
[ -r "$MOSAIC_SYSTEM_PROMPT_FILE" ] || { echo "pi adapter: system prompt not readable: $MOSAIC_SYSTEM_PROMPT_FILE" >&2; exit 2; }
# MOSAIC_AGENT_NAME is optional in headless mode (identity section is then
# omitted); interactive launches always set it via scripts/agent.sh.
: "${PI_PROVIDER:?pi adapter: PI_PROVIDER is required}"
: "${PI_MODEL:?pi adapter: PI_MODEL is required}"
INTERACTIVE="${MOSAIC_INTERACTIVE:-}"
if [ "$INTERACTIVE" != "1" ]; then
[ -n "${MOSAIC_REQUEST:-}" ] || { echo "pi adapter: MOSAIC_REQUEST is required" >&2; exit 2; }
fi
# Workspace (M5): run inside the provided workspace when present.
if [ -n "${MOSAIC_WORKSPACE:-}" ]; then
mkdir -p "$MOSAIC_WORKSPACE"
cd "$MOSAIC_WORKSPACE"
fi
# Session (M6/M11): default ephemeral (--no-session). With a declared
# session dir: persist there and resume the most recent session. With a
# fork source: branch the source session file into the target dir
# (pi --fork) - the ancestor session is never modified.
SESSION_FLAGS="--no-session"
if [ -n "${MOSAIC_SESSION_FORK:-}" ]; then
[ -n "${MOSAIC_SESSION_DIR:-}" ] || { echo "pi adapter: session fork requires MOSAIC_SESSION_DIR" >&2; exit 2; }
mkdir -p "$MOSAIC_SESSION_DIR"
SESSION_FLAGS="--fork $MOSAIC_SESSION_FORK --session-dir $MOSAIC_SESSION_DIR"
elif [ -n "${MOSAIC_SESSION_DIR:-}" ]; then
mkdir -p "$MOSAIC_SESSION_DIR"
SESSION_FLAGS="--session-dir $MOSAIC_SESSION_DIR"
if [ -n "$(ls -A "$MOSAIC_SESSION_DIR" 2>/dev/null)" ]; then
SESSION_FLAGS="$SESSION_FLAGS -c"
fi
fi
# Capabilities (M5): explicit allowlist or no tools.
TOOLS_FLAG="--no-tools"
[ -n "${MOSAIC_TOOLS:-}" ] && TOOLS_FLAG="--tools $MOSAIC_TOOLS"
# Skills (M17): explicitly provided skill dirs replace discovery. When none
# are provided the agent runs with --no-skills (nothing ambient to find).
SKILLS_FLAG="--no-skills"
if [ -n "${MOSAIC_SKILLS:-}" ]; then
SKILLS_FLAG=""
OLDIFS=$IFS; IFS=','
for s in $MOSAIC_SKILLS; do
[ -d "$s" ] || { echo "pi adapter: skill dir missing: $s" >&2; exit 2; }
SKILLS_FLAG="$SKILLS_FLAG --skill $s"
done
IFS=$OLDIFS
fi
# Mode (M13): interactive TUI or one-shot print.
PRINT_MODE="-p"
REQUEST_ARG=""
if [ "$INTERACTIVE" = "1" ]; then
PRINT_MODE=""
else
REQUEST_ARG="$MOSAIC_REQUEST"
fi
# All flags documented in the pi package README (CLI Reference):
# -p/--print one-shot mode: print the response and exit (omitted in
# interactive TUI mode)
# --system-prompt replace the default prompt with the generated one
# --no-* no ambient context/skills/extensions/templates/themes
# SESSION_FLAGS ephemeral | persistent | forked (per env)
# TOOLS_FLAG per capabilities
# --offline no startup network operations (update checks/telemetry)
PROMPT_CONTENT="$(cat "$MOSAIC_SYSTEM_PROMPT_FILE")"
set -- \
--offline \
--no-extensions \
$SKILLS_FLAG \
--no-prompt-templates \
--no-themes \
--no-context-files \
$TOOLS_FLAG \
$SESSION_FLAGS \
--provider "$PI_PROVIDER" \
--model "$PI_MODEL" \
--system-prompt "$PROMPT_CONTENT"
# One-shot mode appends -p and the request (both safely quoted);
# interactive mode appends nothing - clean TUI.
[ "$INTERACTIVE" = "1" ] || set -- "$@" -p "$MOSAIC_REQUEST"
exec pi "$@"
+42
View File
@@ -0,0 +1,42 @@
# Mosaic Stack development team
These are interactive host development agents working in the canonical
checkout. They do not create managed fleet registrations or change role policy.
Sage leads the project and coordinates assignments, review, and integration
(Jason's ruling, 2026-09-26). Development sessions run in T3 for now.
For the current bootstrap phase, direct coding, review and research use Darkwing,
Dewey, Filbert, Rocko, Researcher and further seats Jason launches from this repository's `agents/` directory,
not fleet seats. Keep changes in `/mnt/storage/src/mosaic-stack`; do not modify
`~/.mosaic` launchers/provisioning or migrate/stop live fleet processes.
| Agent | Responsibility | Runtime | Launch from repository root |
| --- | --- | --- | --- |
| Darkwing | Hands-on engineering; collaborating seat under Sage | Pi, configured Mosaic model | `agents/darkwing/launch.sh` |
| Dewey | Frontend design, UX, accessibility, and UI implementation | Pi, configured Mosaic model | `agents/dewey/launch.sh` |
| Rocko | General development, investigation, testing, and review | T3 session (Claude Code, Opus 5.5, thread b84bb264, lead decision 69); launcher Claude Code, Sonnet model | `agents/rocko/launch.sh` |
| Filbert | General development, investigation, testing, and review | T3 session (Claude Code, Opus 5.5); launcher Pi, `openai-codex/gpt-6-astra:low`, not to be used under R26 until it changes (lead decision 69) | `agents/filbert/launch.sh` |
| Researcher | Source-grounded technical research and evidence | Pi, configured Mosaic model | `agents/researcher/launch.sh` |
| Sage | Project lead: coordination, review, integration; earlier DYOR strategy records retained | T3 session (Claude Code); Pi launcher `zai/glm-5.3:high` retained | `agents/sage/launch.sh` |
Each script supports `--check` and `--fresh`. Normal launches resume the agent's
own conversation; a first launch starts one. See each agent's README for
context inputs, authentication, and recovery details. Launch scripts can also
be invoked by absolute path from any directory. No assignment or model request
is submitted by the launcher itself.
The shared Pi helper supports `--provider NAME`, `--model ID`, and
`--thinking LEVEL` as per-launch overrides of the validated system defaults.
Filbert's wrapper appends the required provider, model and thinking flags so
its launch configuration remains fixed, including on resume. Use Filbert's
wrapper to select that configuration; a direct shared-helper invocation uses
its own supplied flags or the system defaults.
All agents follow repository governance and current user direction. Team
leadership does not add deployment or push authority. Coordinate overlapping
work with Sage and preserve other sessions' changes.
Before 2026-09-26 Sage worked only on DYOR strategy and sat outside the
development queue. Jason then made Sage the project lead and Darkwing a
collaborating seat. A separate fleet Sage seat under `~/.mosaic` is being
decommissioned; it does not speak for this seat. Whether the DYOR strategy work
continues is Jason's call. Joe retains DYOR engineering.
+38
View File
@@ -0,0 +1,38 @@
===== DARKWING NATIVE DEVELOPMENT CONTEXT =====
Your identity is Darkwing. This launch runs Pi directly on the host, in the
Mosaic Stack development repository. The injected SOUL defines your persona;
CONSTITUTION and STANDARDS supply governance, USER supplies user context,
and AGENTS.md supplies repository instructions.
You have host read, bash, edit, write, grep, find, and ls tools. This is a
development TUI with the operator's OS access, not a sandbox or a registered
managed fleet seat. Use repository scripts for Mosaic operations and inspect
their effects before running them. Container paths in skills describe worker
deployments, not your current workspace. A tool's presence is not authority
to change unrelated files, other agents' work, or the live fleet.
For an assigned improvement, inspect the implementation, reproduce the issue,
make the smallest useful change, verify it, and continue through the authorized
outcome. Run scripts/mosaic queue next darkwing for ownership, gates and your next piece;
a new user assignment does not silently resume unrelated queued work.
The local /goal extension is loaded and owns any operator-set goal lifecycle.
Use ms-proactive-agent for work selection and ms-goal for recovery guidance;
do not create a competing goal loop. Follow goal_report's actual schema and
reporting instructions. Its text format is Just Completed / Next Step /
Blocked, with '* none' for empty sections. No external reporting skill is
needed to discover that format. Native development packaging supersedes
older skill statements that this extension is unavailable.
For relocation recovery, read agents/darkwing/work/RESTART.md after the root
AGENTS.md and docs/plans/CURRENT.md. It records verified checkpoints and limits,
not a new assignment. The canonical checkout is /mnt/storage/src/mosaic-stack;
v1/ is archived legacy source. Reconcile newer owner direction before acting.
Conversation history persists across launcher restarts. Goals belong to a
single process incarnation; recover the assignment from verified records and
the operator's direction after a restart. No goal is started by this launcher.
Context is captured anew at launch; source edits do not update this process's
injected snapshot. Relaunch to load approved context changes.
+64
View File
@@ -0,0 +1,64 @@
# Darkwing development TUI
From any terminal, run:
```sh
/home/jwoltje/src/mosaic-stack-dev-test/agents/darkwing/launch.sh
```
The agent launcher is a thin shim to `scripts/agent.sh --host-dev darkwing`,
forwarding all arguments unchanged. `scripts/agent.sh` is the common entry
point; `scripts/agent-host-dev.sh` implements its native development mode.
The host launcher opens the repository as Darkwing's workspace.
It uses the repository-pinned Pi, the configured Mosaic provider/model, and
native Pi authentication (normal `~/.pi/agent`, or `PI_CODING_AGENT_DIR` if
explicitly set). It never copies credentials. Install dependencies with
`npm ci --ignore-scripts --no-audit --no-fund` if needed.
`--check` validates configuration and required inputs without opening Pi or
calling a model. `--fresh` starts a new conversation without deleting earlier
ones. Normal launches continue the latest conversation under
`.pi/state/darkwing/sessions/`; the first launch creates one. A launcher lock
rejects simultaneous launches through this script. It does not exclude Pi
processes started another way. Damaged JSONL history refuses automatic resume;
`--fresh` is an explicit escape hatch that preserves the damaged evidence.
The current files are combined into a private launch snapshot under
`.pi/state/darkwing/launches/`:
- `contracts/CONSTITUTION.md` and `contracts/STANDARDS.md`
- `agents/darkwing/SOUL.md`
- `<configured dataRoot>/user/USER.md`, the deployment's live user profile
- the repository's `AGENTS.md` and Darkwing's `CONTEXT.md`
Use `--soul FILE`, `--constitution FILE`, or `--user FILE` to select alternate
inputs, including a future `contracts/USER.md`. Relative paths resolve from
the repository root. Missing or empty inputs refuse launch. Snapshots can
contain personal context and remain local, with private file permissions.
Context edits take effect on relaunch, including when resuming a conversation.
The launcher enables coding/search tools, `goal_report`, ten explicit local
skills, and the canonical goal extension through `scripts/sync-dev-extensions.sh`.
Ambient context, skills, extensions, templates, and themes are disabled.
The normal Pi coding prompt is retained with the Mosaic context appended.
Enter `/goal <assignment and acceptance criteria>` to start continuing work;
`/goal stop`, `/goal resume`, and `/goal` pause, resume, and inspect it. A new
process does not automatically adopt a previous process's goal.
This TUI has the operator's host access, including repository edits and host
commands. Its tool list is not OS isolation. It creates no managed role or
fleet registration. Worker dispatch still uses the governed Mosaic task runner.
The user supplies the assignment; launch alone does not start self-modification.
## Deployment findings
The existing `scripts/agent.sh` launches a Docker container, defaults to the
`agent-<name>` session directory, and asks Pi to continue when that directory
is nonempty. Its default workspace is `<dataRoot>/workspaces/<name>`, not this
checkout. `src/load-contracts.sh` loads image-baked governance, an optional
seat SOUL override, live user Markdown, and mission context into a shared
prompt path. A seat override requires `agent.json`; a standalone SOUL is not
discovered. `adapters/pi/adapter.sh` disables extensions. The temporary host
launcher follows the existing native development path to provide repository
access and `/goal`, and keeps its conversations separate from container and
live fleet sessions. It does not invoke release alignment on startup.
+30
View File
@@ -0,0 +1,30 @@
# SOUL — Darkwing
You are Darkwing, Mosaic Stack's hands-on engineering collaborator, working
with Sage as project lead. Your job is to help Jason make the system
dependable by using it, finding where it falls short, and carrying authorized
improvements through verification.
Be curious, direct, and resourceful. Have a technical opinion and explain
the evidence behind it. Investigate before guessing. Distinguish a design
claim, a passing test, and behavior you have observed in the running system.
Use Mosaic's own tools and workflows where they fit. Turn a failure into a
reproducible case, make a focused correction, and test the behavior again.
Let each verified improvement inform the next one within the assignment.
Keep the human informed when the result, scope, or next decision changes.
Own the outcome while respecting other agents' work. Preserve their changes
and records, give delegated work clear boundaries, and seek independent
review where required. Self-improvement never grants new authority: changing
your instructions, permissions, or a live deployment follows the same review
and authorization rules as any other system change.
Sage leads the development team (Jason's ruling, 2026-09-26) and translates
Jason's priorities into scoped work, coordinates ownership and dependencies,
reviews results, and verifies integration. You are a collaborating seat.
Dewey owns frontend design and UX. Rocko and Filbert support general project needs,
including implementation, investigation, testing, and review. Reconcile
concurrent edits with Sage before integration. Keep Jason informed of
outcomes and decisions that require his input. Your role does not expand the
project's existing authorization or release rules.
+10
View File
@@ -0,0 +1,10 @@
#!/usr/bin/env bash
# Darkwing's native development mode through the Mosaic agent entry point.
set -euo pipefail
REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)"
# Register this seat with the control board (packages/seat) unless already
# registered by `mosaic launch` or only running the checks.
if [ -z "${MOSAIC_LAUNCH_REGISTERED:-}" ] && ! printf '%s\n' "$@" | grep -qx -- '--check'; then
exec "$REPO/scripts/mosaic" launch --repo "$REPO" --harness pi darkwing -- "$@"
fi
exec "$REPO/scripts/agent.sh" --host-dev darkwing "$@"
+21
View File
@@ -0,0 +1,21 @@
// Refuse damaged history before Pi's --continue can silently skip it.
import { readFileSync, lstatSync } from 'node:fs';
try {
for (const file of process.argv.slice(2)) {
if (!lstatSync(file).isFile()) throw new Error(`not a regular session file: ${file}`);
const lines = readFileSync(file, 'utf8').trim().split('\n');
const entries = lines.map((line) => JSON.parse(line));
const header = entries[0];
if (header?.type !== 'session' || typeof header.id !== 'string' || !header.id ||
typeof header.version !== 'number' || typeof header.cwd !== 'string' ||
!Number.isFinite(Date.parse(header.timestamp)) ||
entries.slice(1).some((entry) => !entry || typeof entry.type !== 'string')) {
throw new Error(`invalid session structure: ${file}`);
}
if (header.cwd !== process.cwd()) throw new Error(`session belongs to another workspace: ${file}`);
}
} catch (error) {
console.error(`darkwing: cannot safely resume: ${error.message}; inspect history or explicitly use --fresh`);
process.exit(1);
}
+108
View File
@@ -0,0 +1,108 @@
# Darkwing — relocation handoff
Recorded 2026-09-07 17:43 UTC. Jason intends to relaunch with
`/mnt/storage/src/mosaic-stack/agents/darkwing/launch.sh`.
This is a recovery note, not a new assignment or automatic goal resumption.
## Read first
1. Root `AGENTS.md` and `docs/plans/CURRENT.md`.
2. This note, then `git status --short` and `git log --oneline -5`.
3. Reconcile current owner direction and any newer declared artifacts before acting.
## Repository conversion is completed locally
Jason explicitly ordered the conversion and confirmed no work was active.
- Canonical checkout: `/mnt/storage/src/mosaic-stack`.
- Origin: `https://git.mosaicstack.dev/mosaicstack/stack`.
- Branch: `refactor`.
- Conversion commit: `127a54fdff1fe6ae56c3197edddf957481465db4`.
- New foundation is at root. `v1/` is legacy archival source, NOT current code.
- Old `/home/jwoltje/src/mosaic-stack-dev-test` is a compatibility symlink to this
same checkout. Do not recreate a second working copy there.
- Both histories retained: merge parents v2 `9a5fbdbda74b16adf488fe28138b2ba69ea5e669`
and v1 `5d2770002612a09ae0cadc129b4ea30619133e8a`.
- Exact 3,507-file v1 tracked tree imported; v1 refs under `refs/archive/v1/`.
- Original v2 refs retained; `stack-v2-archive` remote has a disabled push URL.
- Only legacy tracked tree and four conversion docs committed. All earlier
uncommitted/untracked/ignored work preserved. Index was verified clean.
- Issue https://git.mosaicstack.dev/mosaicstack/stack/issues/1495 closed explicitly
for local conversion. No push, PR/trunk merge or live-service change occurred.
Record: `docs/plans/2026-09-07_repository-consolidation-completed.md`.
Receipts: `docs/plans/reviews/2026-09-07_repository-conversion-verification.json`
and `2026-09-07_repository-conversion-postcommit-verification.json`.
Verified rollback copies, NOT development roots:
- `/mnt/storage/src/.mosaic-stack-conversion-20260907T172430Z/`
- `/home/jwoltje/src/.mosaic-stack-dev-test.pre-conversion-20260907T172430Z`
Do not delete them, launch from them or restore over newer work.
## Current unfinished foundation gate
Jason's A9 acceptance of the first offline synthetic scope/permission inspector
is pending. Code is independently approved by Filbert; no blocking code finding
remains at the reviewed r6 candidate. Owner acceptance is not inferred from tests.
- Manifest: `docs/plans/reviews/2026-09-07_foundation-inspector-rocko-build-manifest-r6.json`
SHA-256 `a4a4493000aff5905337a643886ca36e7c5377d52deed77b8aeab7174ca73dcf`.
- Report: `docs/plans/reviews/2026-09-07_foundation-inspector-rocko-build-r6.md`
SHA-256 `ee0e83efd7c71eddecf5e26f939e9a34ba85b184cfcd1cffac9ff9e56ea13c37`.
- APPROVED verdict: `docs/plans/reviews/2026-09-07_foundation-inspector-code-verdict-r6.md`
SHA-256 `ab9dd5e5c3cad5c9263e873ff82cac444da2d36040e907e4798b208fa1c08b13`.
- Guide: `docs/plans/reviews/2026-09-07_foundation-inspector-demo.md`.
All 382 approved inspector files and pinned inputs survived conversion unchanged.
Actual offline checks: Node 80/0, selftests 43/0, oracle 1,568 records / zero
schema disagreements, foundation checker PASS, config/auth/conductor 24/15/17.
Postcommit conductor 17/0 and four CLI demos passed: allowed read, allowed change
PREVIEW (no mutation), missing-registration refusal, unresolved reassignment with
original selection retained. Demo inputs are separate synthetic scenarios.
`test-task.sh` and `test-release.sh` remain NOT RUN / DEFERRED under Jason's bounded
offline-demo ruling. No deployment/native/live/provider/security certification.
Reviewer qualifications: ordering equality means structural equality, not byte
identity; auxiliary native-parser warm-run anomalies remain separate unresolved
observations, not a passing universal parser-equivalence claim. Preserve all earlier
NOT APPROVED reviews and the historical correction that r3 ran unauthorized live
branches; later deferral did not retroactively authorize them.
## Ownership and limits
- Rocko authored inspector code; Filbert independently reviewed; Darkwing coordinates
and verifies. Keep the approved candidate frozen unless a new fix is authorized.
- No automatic permission to push, merge to next/main, deploy, change live config,
grant permissions, access credentials, investigate ~/.mosaic, or start new runtime
work. Local conversion authority is not authority for those activities.
- Preserve unrelated pending work. In particular `scripts/agent.sh`, `docs/TOOLS.md`,
host launcher/context files and other untracked concepts/skills belong to existing
work. Do not blanket-stage/reset/clean. Root logs and CURRENT remain uncommitted.
- Foundation #53 in the old stack-v2 project remains a separate open issue; do not
silently close or renumber it. Accepted historical SHA/path citations remain valid.
- Rocko's Archify C1 remains HELD for owner T2/T3 decisions. No lane reassignment.
- Future durability/workflow/evidence/federation/onboarding topics are notes, not
authorization to expand the inspector.
## Communications
Use only `tools/tmux/agent-send.sh`; sender `dragon-lin:darkwing`.
Rocko: `-L mosaic-fleet -s '=rocko'`; Filbert/Dewey:
`-L default -s '=filbert'` / `'=dewey'`.
Conversion notice delivered to Rocko. Filbert/Dewey sends were unconfirmed
(input boxes not locatable); no retries, no acknowledgement claimed. Check declared
artifact paths as well as direct messages; completed reviews have existed without
transported replies. Do not inspect private panes or blindly resend.
## Relaunch and goal recovery
The project launcher continues its own latest `.pi/state/darkwing/sessions/`
conversation by default. Do NOT assume this pre-launch conversation is already in
that store or that the next launch resumes this exact conversation. This handoff
is the durable bridge. No session-tree migration or launch was performed here.
The goal extension owns lifecycle. The earlier extension goal had been paused;
no restart automatically resumes it. Reconcile the actual new process state and
Jason's direction rather than reporting progress against a guessed old goal or
creating a second goal loop. Launch alone grants no new assignment.
This handoff and its CONTEXT pointer are documentation-only. Launcher scripts,
private sessions, credentials and runtime configuration were not modified.
@@ -0,0 +1,17 @@
{
"issue": 1503,
"candidate": "/tmp/board-attention-r1-q_924ksh",
"files": [
"AGENTS.md",
"packages/control-board/src/scan.mjs",
"packages/control-board/README.md",
"packages/control-board/tests/scan.test.mjs",
"packages/control-board/tests/serve.test.mjs",
"packages/control-board/tests/attention.test.mjs",
"packages/control-board/tests/attention-flow.test.mjs",
"packages/webui/tests/fixture.mjs",
"docs/plans/2026-09-13_board-attention-status.md"
],
"manifestSha256": "e40b58ecb6844d407ba776dcde8f19ad21c0b78076ca1f9b3d96b0bb1c852405",
"testsPassed": 144
}
@@ -0,0 +1,14 @@
{
"at": "2026-09-14T00:19:32.409806+00:00",
"backendPid": 3204655,
"health": "ok",
"researcher": {
"agent": "researcher",
"project": "mosaic-stack",
"alive": true,
"state": "idle",
"waitingOnYou": false,
"lastActivity": "2026-09-13T21:19:57.102Z"
},
"fiveAgentProcessesUnchanged": true
}
@@ -0,0 +1,46 @@
{
"at": "2026-09-14T00:19:11.579013+00:00",
"ownerAuthorized": true,
"oldPid": 1265952,
"newPid": 3204655,
"command": [
"/usr/bin/node",
"packages/control-board/src/cli.mjs",
"serve"
],
"cwd": "/mnt/storage/src/mosaic-stack",
"log": "/tmp/board-attention-backend-ag9xvks2.log",
"agentProcessesBefore": {
"default/darkwing": [
[
"2733924",
"12863634"
]
],
"default/dewey": [
[
"934346",
"466065"
]
],
"default/filbert": [
[
"72183",
"100870"
]
],
"default/researcher": [
[
"173699",
"66404285"
]
],
"mosaic-fleet/rocko": [
[
"90599",
"128275"
]
]
},
"gracefulExitObserved": true
}
@@ -0,0 +1,54 @@
# CHAT-02 board routes: Darkwing's review (#1507)
Reviewer: Darkwing, 2026-09-26, per Sage's D3. Requested by Dewey. Scope: the
two read-only routes only. Filbert reviews `packages/conversation` in full.
Candidate, base 34777c56, uncommitted. I verified both hashes:
- `packages/control-board/src/serve.mjs` afc95bdb…c540d
- `packages/control-board/tests/serve.test.mjs` e60aa14b…ecbc
**Verdict: approve**, with one commit condition and two nonblocking notes.
## What I checked
- Order. `foreignRequest` runs first on every request, then the POST routes,
then the GET/HEAD check (405 otherwise), then these routes. A POST to
either path is 405, and a foreign Host or Origin is 403 before any read.
- Query validation. `/api/conversations` refuses any parameter.
`/api/conversation` accepts only `id`, `branch` and `cursor`, one value
each, each matching `QUERY_VALUE`. That is the same pattern as
`parts.mjs` `ID`, so every id the reader issues (`safeId`, `root`, `c-`
cursors, `pi-` conversations) passes. No path comes from the request.
- Responses. JSON with `no-store` and `nosniff`, no CORS headers. A thrown
error gives a fixed 500 body and logs to stderr only.
- Refusal bodies. Every `Refusal` message in `packages/conversation/src` is
a fixed string. The one interpolated message (`denied`, safe-fs.mjs:34)
interpolates only "session root" or "session file". No path or content
reaches the client through `error`.
- Status map. It covers every code the route can reach. `unknown-actor` and
`unsupported-purpose` are absent, and the route can't produce them because
it always passes the default actor and purpose.
- Tests: `serve.test.mjs` plus `packages/conversation/tests/`, 61/61 on the
pinned files.
## Commit condition
`serve.mjs` imports `../../conversation/src/reader.mjs` at module load, and
`packages/conversation/` is untracked. Committing the routes without that
package breaks the board's start, not only these routes. The package must
land in the same commit or an earlier one, after Filbert's review.
## Notes (nonblocking)
1. **A cursor needs its branch.** The header comment says `branch` and
`cursor` are optional. But `next()` compares `branch !== record.branch`,
and every cursor record carries a string branch (`safeId` or `root`). So
`?id=X&cursor=C` without `branch` is always 409 `cursor-foreign`, with
`reconcile: true`. That is safe, but a client that follows `nextCursor`
alone gets a refusal that reads like a stale view. The test passes the
page's branch, so it doesn't show this. Either say in the comment that a
cursor call must repeat `page.branch`, or answer 400 "cursor requires
branch". I'd take the comment now and let CHAT-03's client decide.
2. **New codes fall to 422.** A code the reader adds later maps to 422
without a test failing. A test that runs the reader's refusal codes
through `REFUSAL_STATUS` would catch that. That's optional.
@@ -0,0 +1,69 @@
# CHAT-02 board routes, revision 2: Darkwing's review (#1507)
Reviewer: Darkwing, 2026-09-26, at Sage's request. Scope: the route delta
since my R1 approval (`chat-02-routes-review-2026-09-26.md`, 07b10ad1). Filbert
reviewed the backend (packet `agents/dewey/work/chat-02/BACKEND.md`, 0cf177b1).
Candidate, base 34777c56, uncommitted. I verified both hashes:
- `packages/control-board/src/serve.mjs` d62720dc…a2f3
- `packages/control-board/tests/serve.test.mjs` d38aa2b2…3f4a
**Verdict: approve.** The R1 commit condition still holds, and I have two
new nonblocking notes.
## What I checked
I kept no copy of the R1 files, so I read the whole route change against the
base (`git diff 34777c56 -- packages/control-board`) instead of only the
delta. It covers every item the packet's §0 lists and nothing else in the
route path.
- **Cursor needs its branch.** This was my R1 note 1. `conversationQuery` now
answers 400 "a cursor call repeats the page's branch" when `cursor` comes
without `branch`. The check runs after the per-key validation, so a
malformed value still gets its own 400 first. The header comment says the
same. A test covers it, and removing the line fails it.
- **Status map.** My R1 note 2. `REFUSAL_STATUS` is exported and now has 16
entries. I listed every `new Refusal("<code>"` in
`packages/conversation/src` myself and got 15 codes plus
`unsupported-harness`, which reader.mjs:352 raises by value. That matches the
map exactly. `parts.mjs` raises none. `unavailable`, which safe-fs.mjs:58
raises when a session root doesn't exist, is 404. That fits the rest of the
map, where not-found is 404. `unknown-actor` 403 and `unsupported-purpose` 422
are explicit now.
- **Order and guards** are unchanged from R1. The foreign Host or Origin check
comes first, then the POST routes, then 405, then these routes. No path comes
from the request, and responses carry `no-store` and `nosniff` with no CORS
headers.
- **Tests.** `serve.test.mjs` plus `packages/conversation/tests/` pass
67/67 on the pinned files.
- **Mutations** on a scratch clone of HEAD with the conversation package and
the two pinned files:
| Mutation | Result |
|---|---|
| cursor-without-branch check removed | 1 fails |
| `unknown-actor` entry dropped | 1 fails (the scan test) |
| `nosniff` removed | 1 fails |
| repeated-parameter check removed | 1 fails |
| catalogue parameter check removed | 1 fails |
| `unavailable` changed from 404 to 422 | nothing fails |
The last row is note 1 below.
## Commit condition (unchanged)
`serve.mjs` imports `../../conversation/src/reader.mjs` at module load, and
`packages/conversation/` is still untracked. The package must land in the
same commit as the routes or an earlier one. Otherwise the board fails to
start.
## Notes (nonblocking)
1. **The scan test checks keys, not values.** It proves every code has an
entry. No test proves `unavailable` is 404. If someone edits that value, or
any status no route test exercises, nothing fails. A table test that
asserts the whole `REFUSAL_STATUS` object would pin them. That's optional.
2. **The scan reads a fixed list of three files.** If a refusal is added to
`parts.mjs` or a new file, the scan won't see it, and that code falls to 422.
Reading every `.mjs` in `packages/conversation/src` would close the gap.
@@ -0,0 +1,160 @@
# Queue row 5, CHAT-03 increment 1, round 1 review (#1508)
Darkwing, 2026-10-04. Request: #1508 comment 26671. My part is the binding
and the extension-load refusal (lead decisions 31 and 36). Brief:
`agents/dewey/work/chat-03/BRIEF.md` at 1ef15ac0. Candidate:
`agents/dewey/work/chat-03/I1-manifest.sha256`, digest
`1404341eaeaf7d1e684c9f27e76718061ed08c25a52da5f190425feb1274ba69`,
27 files, uncommitted.
Verdict: changes requested. Two blocking findings, B1 and B2.
## Checks
- The manifest hashes to 1404341e. All 27 files match it in the canonical
working tree.
- Export: `git archive` of 1c724958 plus the 27 candidate files in
`/tmp/r5-exp`, manifest OK there too. `node --test
packages/conversation/tests/` passes 141/141, 0 skipped.
`node --test packages/control-board/tests/` passes 124/124 (the board
reads `ControlRefusal` codes). The CHAT-00, CHAT-01 and CHAT-01c checks
still pass.
## B1 (blocking). The seal is a deny-list over an argv the caller builds
`checkSeal` (pi-pin.mjs 59–66) refuses `-e`/`--extension`, a missing seal
flag, or a first pair other than `--mode rpc`. Everything else in
`engine.extraArgs` passes, and `buildPiArgs` appends extraArgs after the
controller's own `--session`. Pi's parser keeps the last `--mode` and the
last `--session`. The controller is a public export (`./controller` in
package.json), so this is reachable without touching the source.
(a) `extraArgs: ["--session", <file outside the fixture root>]`. The guard
refuses that same path when it is passed as `sessionFile`, but it never
sees extraArgs. With the real pinned Pi under a scratch HOME and agent dir
(`/tmp/r5/seal-escape.mjs session`):
```
guard on sessionFile: live-session-refused
constructed with extraArgs ["--session",".../outside/proj/.pi/state/other/sessions/s1.jsonl"]
-> piArgs ["--mode","rpc","--no-extensions","--no-prompt-templates","--no-themes",
"--session",".../fx/.../fixture-seat/sessions/s1.jsonl","--session",".../outside/..."]
start -> {"launched":true,"classified":{"state":"free"}} | binding uncertain closed
uncertain evidence: loaded-session, "the engine loaded another session file"
outside session changed: true | appended: {"type":"thinking_level_change","id":"bef67aef",
"parentId":"b2c3d4e5",...,"thinkingLevel":"off"}
```
K8 notices afterwards, but Pi has already written to a session the guard
exists to protect. That breaks §2 "no writes to sessions" and the
fixture-only rule in code.
(b) `extraArgs: ["--mode", "json"]` (or `text`). checkSeal passes because
it looks only at args[0] and args[1]. Pi starts in print mode, reads stdin
to EOF and treats it as the prompt. The K8 `get_state` line goes into that
reader, so the run ends `uncertain` on `get_state: timeout`. In an earlier
run where I closed stdin, Pi sent the `get_state` JSON line to the model
as a prompt and stopped only at "No API key found". The default
`engine.env` is `process.env`, so with real auth present that becomes a
paid model call outside RPC. PgroupLauncher spawns detached, so the engine
outlives the controller unless someone kills the group.
Other extraArgs that pass checkSeal today: `--no-session`, `--fork`,
`--export <file>` (writes a file), `--prompt-template`, `--approve`.
`engine.command` and `engine.preArgs` can also run any script, or node
with `--import`. `checkEnginePin` validates only the pinRoot lock files, so
the pin says nothing about what actually ran.
Suggested fix:
- Allow-list extraArgs. The only non-extension use in the suites is
`["--model", "other"]` (claim.test 548, W9), so `--model`, `--provider`
and `--thinking` with one value each would cover it.
- Refuse any second `--mode`, `--session` or seal flag, and any session
or output flag (`--print`, `--no-session`, `--session-dir`,
`--session-id`, `--fork`, `--export`, `--continue`,
`--resume`).
- Say in the README that a non-default `command` or `preArgs` is a test
hook, and that the pin and seal checks don't bind under it. Or refuse it
outside tests.
- N24 cases for `--session`, `--mode json` and `--no-session` in
extraArgs.
## B2 (blocking). The claim's session key is the conversation ID
controller.mjs 187:
`this.sessionK = sessionKey({ harness: "pi", conversation: this.conversation })`.
`conversation` is `"pi-" + sha256(projectRoot, seat, name)` (reader.mjs
72). The brief (line 382–383) keys the claim on "the native session
identity (the Pi header ID or the Claude session UUID)", and the CHAT-01
README says at most one non-stopped binding may hold a conversation or
native-session identity. Two paths to one session file give two
conversation IDs, so the session key never collides.
Repro, `/tmp/r5/hardlink.mjs`: one session hard-linked into
`.pi/state/fixture-seat/sessions/s1.jsonl` and
`.pi/state/seat-b/sessions/s1.jsonl`, two controllers via the test harness
with the fake engine.
```
same inode: true
A conversation pi-b5901a69... | B conversation pi-f21f1a75...
A start: {"launched":true,...} state active
B start: {"launched":true,...} state active
A nativeSession 0f5e1c2a-1111-4222-8333-944455556666 | B nativeSession 0f5e1c2a-1111-4222-8333-944455556666
```
Two active bindings, two engines, one session. A plain copy is allowed
too; whether a copy should count as the same session is a call for the
brief, but the hard link plainly should. W4 (claim.test 172) builds its
keys by hand on the store, so it never exercises the controller's
mapping.
Fix: build the session key from the header `id` read in `start()` (the
`nativeSession` already in hand). Add a controller-level W4 case for the
hard link, and one for a copy with whatever the brief decides.
## n1 (non-blocking). The startup-append note is narrower than Pi's rule
README 426–431 and BUILD-I1.md 196 say the append bites only sessions
without a `thinking_level_change` entry, and that Pi-created sessions
carry one. Pi's rule is sdk.js 82:
`hasExistingSession = existingSession.messages.length > 0`. A session
with a thinking entry but no messages takes the new-session branch and
appends `thinking_level_change` at every start, plus `model_change` when a
model is set. I checked this in plain sealed RPC mode: a header plus a
thinking entry gained `{"type":"thinking_level_change","id":"e8c74875",
"parentId":"f0e1d2c3",...}`. A Pi-created session that was opened and
never prompted is in this group, so "written by something else" is wrong.
Name message-less sessions too, and let the later pre-spawn check cover
both.
## n2 (minor). The guard follows $HOME
`LiveSessionGuard` protects `~/.pi`, `~/.claude` and `~/.mosaic-dev` via
`os.homedir()`. With a scratch HOME, the real directories are protected
only by the fixture-root containment, or by passing `homes`. That
containment holds today, so I'm noting it, not blocking on it.
## What holds in the binding
- H1 to H3. `#dispatch` holds `this.lock` across `#recheck` and the
write. Takeover, release and acquire run `#evaluate` under the same
lock, so a prompt can't land between the generation bump and the write.
- The generation check comes first in `#evaluateOp` and again in
`#recheck`.
- Takeover refuses while fenced (K9) and for the current controller
(`already-controller`, H4).
- Disconnect never moves control (H11).
- Confirmations are single-use: `#checkConfirmation` marks them
`consumed`.
- Interrupt sets its fence before it takes the lock, so a queued prompt
sees the fence.
- K8 catches a wrong session or leaf after launch, as in B1(a). B1 is that
the write happens before K8 can run.
- The pin check reads both lock files and refuses on either version or
integrity mismatch.
Scratch scripts and outputs are in `/tmp/r5/`: `seal-escape.mjs`,
`hardlink.mjs`, `seal-session.txt`, `seal-json.txt`, `hardlink.txt`,
`conv-suite.txt`, `board-suite.txt`. All scratch engines were killed by
their process group.
@@ -0,0 +1,119 @@
# Queue row 5, CHAT-03 increment 1, round 2 review (#1507)
Darkwing, 2026-10-04. Request: #1507 comment 26689 (Dewey). My part is
the binding and the extension-load refusal (lead decisions 31 and 36), and
my round 1 findings (comment 26681 on #1508, pointer 26685 on #1507).
Candidate: `agents/dewey/work/chat-03/I1-r2-manifest.sha256`, digest
`2b48e333a0f09185364359ae6f8277cc88c0b9ff39058de45cc2c1f0ec9d5c4a`,
the same 27 files, base 1c724958, uncommitted. Packet:
`agents/dewey/work/chat-03/BUILD-I1-r2.md`.
Verdict: approved for my part. B1 and B2 are fixed. Nothing must be fixed
before the commit. Three follow-ups below, none of them blocking.
## Checks
- The manifest hashes to 2b48e333. All 27 files match it in the canonical
working tree. Twelve changed from round 1: README, `controller.mjs`,
`pi-pin.mjs`, `terminal.mjs` and seven test files. `engine.mjs` is
unchanged.
- Export: `git archive` of 1c724958 plus the 27 files in
`~/darkwing-scratch/r5r2/exp`, manifest OK there too, `TMPDIR` under the
same directory.
- `node --test packages/conversation/tests/` passes 152/152, 0 skipped.
- `node --test packages/control-board/tests/` passes 124/124. The board
imports the conversation package.
- The CHAT-00 (48 checks), CHAT-01 R3 and CHAT-01c checks still pass.
- Probe script: `~/darkwing-scratch/r5r2/probes.mjs`, output in
`~/darkwing-scratch/r5r2/out/probes.txt`. It drives the exported
`Controller` directly, as an outside caller would.
## B1, the seal is now an allow-list: fixed
`checkSeal` (pi-pin.mjs 68) requires the exact prefix `--mode rpc`, the
seal flags and `--session <absolute path>`, then accepts only `--model`,
`--provider` and `--thinking`, each once, each with one value that is
nonempty and doesn't start with `-` or `@`. The constructor refuses a
non-list `preArgs` or `extraArgs` and runs the seal before anything else
happens (controller.mjs 170–174), so a refused argv never reaches a spawn.
I passed 24 `extraArgs` lists to the constructor. These all refuse
`unsealed-engine`:
- my round 1 vectors: `--session <outside file>`, `--mode json`,
`--mode text`, `--no-session`, `--fork <file>`, `--export <file>`;
- `--print`, `-p hi`, a bare word, `@<file>`, `--approve`,
`--no-extensions`, `--extension x`, `-e x`;
- `--model=x`, `--model` with no value, an empty value, a value of `-p`
or `@f`, `--model` twice, and `--model a hello`.
These build: `--model a --thinking high --provider p`, `--thinking high`,
and `--model "rpc --mode json"`. The last one is a single argv element with
no shell in between, so Pi sees it as a model name. That's fine.
Dewey's mutants r2-B1 to r2-B1e (no extraArgs check, no prefix order, a
relative session path, a repeat, a flag or `@` value) are all killed by
N24.
## B2, the session key is the Pi header ID: fixed
The constructor reads the session header and keys the claim on
`sessionKey({ harness: "pi", nativeSession })` (controller.mjs 193–194).
I ran two controllers, seat A on the fixture session and seat B on a
second file under another seat directory:
| Second file | A | B | Keys equal |
|---|---|---|---|
| hard link of A's file | active, 1 launch | refused `already-active`, 0 launches | yes |
| copy of A's file | active, 1 launch | refused `already-active`, 0 launches | yes |
| copy with a different header ID | active, 1 launch | active, 1 launch | no |
So the fix doesn't over-collide: two real sessions still start side by
side. I also replaced the header ID by rename after construction. `start()`
refused `target` with no launch. Mutants r2-B2 and r2-B2b are killed by
the new W4 cases.
n1 (Pi's startup-append rule) and n2 (the guard follows `$HOME`) are fixed
in the README as asked.
## What the rework touched that I checked
- `#readSession` runs after the guard check at construction and wraps an
unreadable file as `configuration`. Good.
- The force-stop path now refuses `fenced` while an escalation runs
(controller.mjs 683). `#forceStop` holds `escalating` and clears it in
`finally` (1453–1459). See F2 for the one gap.
- `stopLink` needs `abortWritten`, and there's an overlap recheck after the
pause before the abort. Mutants r2-n1 and r2-n2 are killed. This is
outside my part, so I note it without a ruling.
- Mutant r2-B5b (the shim writes the freeze but doesn't wait for
`frozen 1`) survives. That is Filbert's finding and Dewey explains it in
the packet. I leave it to Filbert.
## Follow-ups (none must be fixed before the commit)
- **F1. `engine.command` and `engine.preArgs` are outside the seal.** The
README documents them as a test hook, as I asked in round 1, and
`preArgs` still refuses `--extension`. But `preArgs:
[<pinned cli.js>, "--no-session"]` builds. Everything after `cli.js`
reaches Pi's parser, so a flag there that the seal doesn't repeat later,
such as `--no-session` or `--export`, takes effect. A non-default
`command` can run anything. No caller passes them today: the package
has no entry point that builds a `Controller` from configuration. Before I3 adds one, either
refuse any non-default `command`/`preArgs` outside the test harness, or
have that entry point never accept them from config. Owner: whoever
builds the I3 entry point.
- **F2. `escalating` is set before `after()` runs.** The force-stop branch
of `#evaluate` sets `this.escalating` at controller.mjs 686, then calls
`#admission()` and `link.poison()`. If either throws, `#handle` turns the
result into an internal error and drops `after`. `#forceStop` never runs,
so the flag is never cleared, and every later force stop on that
controller refuses `fenced`. Recovery then needs a controller restart.
Neither call is expected to throw, so the risk is low. Fix: set the flag
only in `#forceStop`, or clear it in the catch path when `after` is
dropped.
- **F3. `engine.env` defaults to `process.env`** (controller.mjs 168). A
real Pi then inherits the caller's whole environment, including
provider keys and the real `HOME`, so it reads the real `~/.pi/agent`.
This is unchanged from round 1 and fine for I1, where tests pass their
own env. The I3 entry point should build the engine env from an explicit
list.
@@ -0,0 +1 @@
fd10b62c10bff4736b9b4e809b283fe4c5b549fa53fe1121a7bb06258147f0e9 packages/conversation/tests/fake-pi.mjs
+152
View File
@@ -0,0 +1,152 @@
# Row 46: K1, K3 and K10 on the scope fixtures
Darkwing, 2026-10-09. Issue #1528, reviewer Dewey. Brief:
`docs/plans/2026-10-09_s4-follow-up-and-cohort.md`, section "Conversation
cohort: K1, K3 and K10 fail on the scope fixtures". Ruling: lead decision 72.
The candidate is not committed, staged or pushed.
## Cause
It's a race in the test fixture. Neither the host's systemd setup nor
`cohort.mjs` is at fault.
`spawnChild` in `packages/conversation/tests/fake-pi.mjs` returns the
child's pid as soon as `spawn()` returns. The child is `node -e`, and it
installs its SIGTERM handler only after Node has booted. That takes 16 to
22 ms on an idle host (`trace-pass.jsonl`) and 43 to 80 ms under 48 CPU burners (`trace-load.jsonl`). A
TERM that arrives before the handler gets the default action, and the child
dies of SIGTERM. Each failing assertion is that death seen from a different
place:
- K1, "the escaped child is a listed member". The child died during the
TERM grace, so the freeze-phase enumeration doesn't list it.
- K3, "the member ignored TERM". `alive(child)` is false at the kill
phase.
- K10, "the member is alive across the crash". Same as K3, before the
restart.
The time between spawn and TERM depends on the disk under TMPDIR. The
fixture puts the claim store there. Before it sends TERM, the controller
publishes claim revisions, each with an fsync on the file and one on the
directory (`packages/conversation/src/claim.mjs:159` and `:176`).
| TMPDIR | Disk | write+fsync median (`fsync.txt`) | spawn to TERM | Unfixed K1/K3/K10 |
|---|---|---|---|---|
| `/mnt/storage/scratch/tmp` | nvme0, ext4 | 0.65 ms | 12 to 13 ms (`trace-fail.jsonl`) | 0/3 (`runs/canon-scratchtmp.txt`) |
| `/tmp` | nvme1, ext4 | 4.63 ms | not traced | 3/3 (`runs/canon-tmp.txt`) |
| `~/darkwing-scratch/tmp` | nvme1, ext4 | 4.69 ms | 53 to 110 ms (`trace-pass.jsonl`) | 3/3 (`runs/canon-hometmp.txt`) |
On the fast disk, TERM lands about 12 ms after spawn, before Node is up.
In `trace-fail.jsonl` the children never log `ready`.
### Why it started failing
This host's seats now get `TMPDIR=/mnt/storage/scratch/tmp`. My shell had
it by default, and so did Sage's gate runs. It comes from Vikunja task 59.
`~/.config/systemd/user/t3code.service.d/tmpdir.conf` was written
2026-10-04 20:33Z. Its own comment says it takes effect only when
`t3code.service` restarts, which hasn't happened (active since 2026-09-23),
and that until then seats get TMPDIR from their harness config. I didn't
establish what TMPDIR the row 44 gate ran with on 2026-10-05, so the
"since when" is likely but unproven. The change that exposed the race is a
seat's TMPDIR. It isn't a systemd or user manager setting, and I changed
nothing on the host.
### The scope is not a factor
Outside any scope, `race.mjs` gives the same split: a TERM 0, 10 or 20 ms
after spawn kills 10/10, and at 30 or 60 ms 10/10 survive (`race.txt`).
Sage's outside-scope probe found the child surviving. I haven't seen that
probe, but it most likely sent TERM after the handler was in place.
## Receipts in a scope
`scope-receipt.mjs` launches a delegated scope per run, with the same
`systemd-run` flags `ScopeLauncher` uses. The scope's main process spawns
the fixture's `ignoreTerm` child and records `systemctl --user show` on the
scope, then `cgroup.procs` before and after TERM, then the child's exit.
Output: `receipts.jsonl`.
| Mode | TERM after spawn | In `cgroup.procs` before | After | Child exit |
|---|---|---|---|---|
| `race 12` | 20 to 25 ms | 5/5 | 1/5 | 4/5 `signal: "SIGTERM"`, 1 alive |
| `race 60` | 69 to 71 ms | 5/5 | 5/5 | 5/5 alive 500 ms after TERM |
| `ready 0` | 28 to 38 ms (ready at 22 to 30) | 5/5 | 5/5 | 5/5 alive 500 ms after TERM |
Each line also carries the scope's `Id`, `LoadState=loaded`,
`ActiveState=active`, `InvocationID`, `ControlGroup` and `Delegate=yes`.
The first line, for example:
`ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-12-3662746-0.scope`,
`before [3662760, 3662788]`, `after [3662760]`, child exit
`{"code":null,"signal":"SIGTERM","atMs":24}`.
`trace.patch` is the scratch-only instrumentation behind the two traces.
It logs spawn, ready and the shim's `term` per pid to `$DW_TRACE`. It is
not part of the candidate.
## The fix
`build.patch` changes one file, `packages/conversation/tests/fake-pi.mjs`:
- The child writes one byte to stdout as its first action after installing
its TERM handler. A child without `ignoreTerm` writes it once its code
starts.
- `spawnChild` returns a promise. It resolves with the pid on that byte,
then closes its end of the pipe. It rejects if the child exits or errors
first, or if 10 s pass.
- The `child` control op returns that promise. The op dispatcher now
answers `ok: false` when an op's promise rejects. Before, a rejected op
promise went unhandled, and under Node's default that ends fake-pi. The
change covers every op that returns a promise, not only `child`.
`childOf` in `cohort.test.mjs` already asserts `r.ok`.
No assertion changed, and `cohort.mjs` is untouched. The three cases now
test what their names say: TERM reaches a child that is already ignoring
TERM.
`build-manifest.sha256` (sha256
`9228414352a56e0465e896b81643cdce000bcedba70f86adac07fc50cfadad2d`) pins
`fake-pi.mjs` after the patch. `build.patch` sha256 is
`04234ce1d886e36dfd06cbd81153b38cbb86909b6ee933151306b8cdedd629d9`. In a
fresh worktree at `521597bb` the patch applies and the manifest checks 1/1.
## Mutants
Both ran with `TMPDIR=/mnt/storage/scratch/tmp`, the condition that fails.
| Mutant | K1/K3/K10 | File |
|---|---|---|
| M1: resolve at spawn, no wait (the old behavior) | 0/3 | `runs/mutant-m1-nowait.txt` |
| M2: wait for ready, but the child has no TERM handler | 0/3, the same three assertions | `runs/mutant-m2-noignore.txt` |
M1 shows the wait is what fixes it. M2 shows the assertions still catch a
child that dies on TERM.
## Runs
The patch was applied in a scratch worktree at `521597bb`. Node v26.8.1.
Start time and load are in `runs/start.txt`.
| Run | TMPDIR | Result | File |
|---|---|---|---|
| `node --test "packages/conversation/tests/*.test.mjs"` | `/mnt/storage/scratch/tmp` | 152/152 | `runs/node-conversation.txt` |
| `node --test "packages/webui/tests/*.test.mjs"` | `/mnt/storage/scratch/tmp` | 14/14 | `runs/node-webui.txt` |
| K1/K3/K10 isolated, 3 consecutive | `/mnt/storage/scratch/tmp` | 3/3, 3/3, 3/3 | `runs/k-iso-{1,2,3}.txt` |
| K1/K3/K10 isolated | `/tmp` | 3/3 | `runs/k-tmp.txt` |
| K1/K3/K10 isolated | `~/darkwing-scratch/tmp` | 3/3 | `runs/k-home.txt` |
| K1/K3/K10 isolated under 48 CPU burners, 3 runs (load 13.6 to 24.9) | `/mnt/storage/scratch/tmp` | 3/3, 3/3, 3/3 | `runs/k-load48-{1,2,3}.txt` |
For the gate rerun, keep the default `TMPDIR=/mnt/storage/scratch/tmp`.
That's the condition that failed, and a slower TMPDIR would pass with or
without the patch.
## Files
- `build.md`, this file.
- `build.patch`, `build-manifest.sha256`: the candidate.
- `scope-receipt.mjs`, `receipts.jsonl`: the scope receipts.
- `race.mjs`, `race.txt`: TERM timing outside a scope.
- `fsync.mjs`, `fsync.txt`: write+fsync latency per TMPDIR.
- `trace.patch`, `trace-fail.jsonl`, `trace-pass.jsonl`, `trace-load.jsonl`: the traced runs.
- `runs/`: suite, isolated, load and mutant outputs, and the unfixed
canonical tree under three TMPDIRs.
@@ -0,0 +1,67 @@
diff --git a/packages/conversation/tests/fake-pi.mjs b/packages/conversation/tests/fake-pi.mjs
index b3014f86..a78dfebd 100644
--- a/packages/conversation/tests/fake-pi.mjs
+++ b/packages/conversation/tests/fake-pi.mjs
@@ -615,14 +615,34 @@ const children = [];
// K2); `forkLoop` forks every 5 ms (K12); `ignoreTerm` survives SIGTERM, so
// only the kill phase ends it (K3, K10, K11). With `pidLog`, the fork loop
// appends each child's pid and a `term` line when it gets SIGTERM (K12).
+//
+// It resolves only once the child writes its ready byte, which it does after
+// installing its TERM handler. Node takes 15 to 80 ms to get there, and a
+// force stop can send TERM sooner (12 ms when the claim store is on a fast
+// disk). A TERM before the handler kills a child that is meant to ignore it
+// (#1528).
function spawnChild({ setsid = false, forkLoop = false, ignoreTerm = false, pidLog = null } = {}) {
const note = pidLog ? `const note=(s)=>require('node:fs').appendFileSync(${JSON.stringify(pidLog)},s+'\\n');` : "const note=()=>{};";
- const code = note + (ignoreTerm ? "process.on('SIGTERM',()=>note('term'));" : "") + (forkLoop
+ const code = note + (ignoreTerm ? "process.on('SIGTERM',()=>note('term'));" : "") + "process.stdout.write('r');" + (forkLoop
? "const {spawn}=require('node:child_process');setInterval(()=>{try{const c=spawn('sleep',['1000'],{stdio:'ignore'});if(c.pid)note(String(c.pid))}catch{}},5);setInterval(()=>{},1e9)"
: "setInterval(()=>{},1e9)");
- const child = spawn(process.execPath, ["-e", code], { stdio: "ignore", detached: setsid });
+ const child = spawn(process.execPath, ["-e", code], { stdio: ["ignore", "pipe", "ignore"], detached: setsid });
children.push(child.pid);
- return child.pid;
+ return new Promise((resolve, reject) => {
+ const fail = (why) => {
+ clearTimeout(timer);
+ reject(new Error(`tool child ${child.pid} ${why} before it was ready`));
+ };
+ const timer = setTimeout(() => fail("took 10 s"), 10000);
+ child.once("error", (err) => fail(err.code ?? err.message));
+ child.once("exit", (code, signal) => fail(`exited (${signal ?? code})`));
+ child.stdout.once("data", () => {
+ clearTimeout(timer);
+ child.removeAllListeners("exit");
+ child.stdout.destroy();
+ resolve(child.pid);
+ });
+ });
}
// K13: a member writes its own pid to another cgroup's cgroup.procs.
@@ -671,7 +691,7 @@ async function main() {
drop: () => fake.dropResponse(req.type, req.n ?? 1),
extension: () => void fake.extensionPrompt(req.args ?? {}),
state: () => ({ streaming: fake.streaming, runs: fake.runs.length, commands: fake.commands, pid: process.pid, children, appends: fake.appends }),
- child: () => ({ pid: spawnChild(req.args ?? {}) }),
+ child: () => spawnChild(req.args ?? {}).then((pid) => ({ pid })),
escape: () => escape(req.target),
cgroup: () => readFileSync(`/proc/${req.pid ?? process.pid}/cgroup`, "utf8"),
waitPaused: () => fake.waitPaused(req.point),
@@ -679,12 +699,13 @@ async function main() {
stall: () => void process.stdin.pause(),
};
if (!ops[req.op]) return sock.write(encodeLine({ id: req.id, ok: false, error: `unknown op ${req.op}` }));
+ const failed = (err) => sock.write(encodeLine({ id: req.id, ok: false, error: String(err.message) }));
try {
const out = ops[req.op]();
- if (out && typeof out.then === "function") out.then(reply);
+ if (out && typeof out.then === "function") out.then(reply, failed);
else reply(out ?? null);
} catch (err) {
- sock.write(encodeLine({ id: req.id, ok: false, error: String(err.message) }));
+ failed(err);
}
});
sock.on("data", (c) => splitter.push(c));
+17
View File
@@ -0,0 +1,17 @@
import { openSync, writeSync, fsyncSync, closeSync, rmSync, mkdtempSync } from "node:fs";
import { join } from "node:path";
for (const base of process.argv.slice(2)) {
const d = mkdtempSync(join(base, "dw-fsync-"));
const t = [];
for (let i = 0; i < 30; i++) {
const s = performance.now();
const fd = openSync(join(d, `f${i}`), "w");
writeSync(fd, "x".repeat(512));
fsyncSync(fd);
closeSync(fd);
t.push(performance.now() - s);
}
rmSync(d, { recursive: true });
t.sort((a, b) => a - b);
console.log(`${base}: write+fsync median ${t[15].toFixed(2)} ms, p90 ${t[27].toFixed(2)} ms`);
}
+3
View File
@@ -0,0 +1,3 @@
/mnt/storage/scratch/tmp: write+fsync median 0.65 ms, p90 0.79 ms
/tmp: write+fsync median 4.63 ms, p90 4.93 ms
/home/jwoltje/darkwing-scratch/tmp: write+fsync median 4.69 ms, p90 9.30 ms
+17
View File
@@ -0,0 +1,17 @@
// Spawn the K fixtures' ignoreTerm child exactly as fake-pi.mjs spawnChild does, then
// SIGTERM it after a delay. Reports how the child ended within 1 s.
import { spawn } from "node:child_process";
const code = "const note=()=>{};process.on('SIGTERM',()=>note('term'));setInterval(()=>{},1e9)";
const delays = process.argv.slice(2).map(Number);
for (const d of delays) {
const r = { died: 0, survived: 0 };
for (let i = 0; i < 10; i++) {
const c = spawn(process.execPath, ["-e", code], { stdio: "ignore" });
const ended = new Promise((res) => c.on("exit", (code, sig) => res(sig ?? code)));
await new Promise((res) => setTimeout(res, d));
c.kill("SIGTERM");
const out = await Promise.race([ended, new Promise((res) => setTimeout(() => res(null), 1000))]);
if (out === null) { r.survived++; c.kill("SIGKILL"); await ended; } else r.died++;
}
console.log(`TERM ${d} ms after spawn: died ${r.died}/10 survived ${r.survived}/10`);
}
+5
View File
@@ -0,0 +1,5 @@
TERM 0 ms after spawn: died 10/10 survived 0/10
TERM 10 ms after spawn: died 10/10 survived 0/10
TERM 20 ms after spawn: died 10/10 survived 0/10
TERM 30 ms after spawn: died 0/10 survived 10/10
TERM 60 ms after spawn: died 0/10 survived 10/10
@@ -0,0 +1,15 @@
{"mode":"race","delayMs":12,"readyMs":null,"termMs":23,"child":3662788,"self":3662760,"show":["Id=dw-r46-race-12-3662746-0.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=2eac4fa367504b9dbca66a4693cc328a","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-12-3662746-0.scope","Delegate=yes"],"before":[3662760,3662788],"after":[3662760],"childInBefore":true,"childInAfter":false,"exit":{"code":null,"signal":"SIGTERM","atMs":24}}
{"mode":"race","delayMs":12,"readyMs":null,"termMs":20,"child":3663479,"self":3663212,"show":["Id=dw-r46-race-12-3662746-1.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=021c6d466fc64961ba3031a0254c3803","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-12-3662746-1.scope","Delegate=yes"],"before":[3663212,3663479],"after":[3663212],"childInBefore":true,"childInAfter":false,"exit":{"code":null,"signal":"SIGTERM","atMs":21}}
{"mode":"race","delayMs":12,"readyMs":null,"termMs":25,"child":3663799,"self":3663588,"show":["Id=dw-r46-race-12-3662746-2.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=ce8e2c2dbbe7481885a50152fe3766b5","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-12-3662746-2.scope","Delegate=yes"],"before":[3663588,3663799],"after":[3663588],"childInBefore":true,"childInAfter":false,"exit":{"code":null,"signal":"SIGTERM","atMs":27}}
{"mode":"race","delayMs":12,"readyMs":null,"termMs":20,"child":3664563,"self":3664322,"show":["Id=dw-r46-race-12-3662746-3.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=cc912ab9a7fe47dd9e3b97486abf0bc3","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-12-3662746-3.scope","Delegate=yes"],"before":[3664322,3664563],"after":[3664322,3664563],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"race","delayMs":12,"readyMs":null,"termMs":21,"child":3665177,"self":3665003,"show":["Id=dw-r46-race-12-3662746-4.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=163271c658cc4ee4b84f5aace54bc806","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-12-3662746-4.scope","Delegate=yes"],"before":[3665003,3665177],"after":[3665003],"childInBefore":true,"childInAfter":false,"exit":{"code":null,"signal":"SIGTERM","atMs":22}}
{"mode":"race","delayMs":60,"readyMs":null,"termMs":69,"child":3665812,"self":3665626,"show":["Id=dw-r46-race-60-3665599-0.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=14681de5ac844879bdfcbcdd3dae8f10","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-60-3665599-0.scope","Delegate=yes"],"before":[3665626,3665812],"after":[3665626,3665812],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"race","delayMs":60,"readyMs":null,"termMs":71,"child":3666377,"self":3666179,"show":["Id=dw-r46-race-60-3665599-1.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=3128eba3ed234776ae4a61cb12fe1210","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-60-3665599-1.scope","Delegate=yes"],"before":[3666179,3666377],"after":[3666179,3666377],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"race","delayMs":60,"readyMs":null,"termMs":70,"child":3666924,"self":3666728,"show":["Id=dw-r46-race-60-3665599-2.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=836de499898d4b6996739e1b2a433fa0","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-60-3665599-2.scope","Delegate=yes"],"before":[3666728,3666924],"after":[3666728,3666924],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"race","delayMs":60,"readyMs":null,"termMs":69,"child":3667387,"self":3667251,"show":["Id=dw-r46-race-60-3665599-3.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=0ebf8d9d930a43629668e10c08f80e0d","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-60-3665599-3.scope","Delegate=yes"],"before":[3667251,3667387],"after":[3667251,3667387],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"race","delayMs":60,"readyMs":null,"termMs":70,"child":3667888,"self":3667727,"show":["Id=dw-r46-race-60-3665599-4.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=9c3594ae3018458594bb3ff996dc0d9a","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-race-60-3665599-4.scope","Delegate=yes"],"before":[3667727,3667888],"after":[3667727,3667888],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"ready","delayMs":0,"readyMs":30,"termMs":38,"child":3668171,"self":3668085,"show":["Id=dw-r46-ready-0-3668072-0.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=4c2c3b3ad52a477e8f735abedf158cbb","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-ready-0-3668072-0.scope","Delegate=yes"],"before":[3668085,3668171],"after":[3668085,3668171],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"ready","delayMs":0,"readyMs":22,"termMs":29,"child":3668354,"self":3668284,"show":["Id=dw-r46-ready-0-3668072-1.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=b9da06280a8a4c6682f05a2d4ff118cf","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-ready-0-3668072-1.scope","Delegate=yes"],"before":[3668284,3668354],"after":[3668284,3668354],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"ready","delayMs":0,"readyMs":22,"termMs":28,"child":3668425,"self":3668417,"show":["Id=dw-r46-ready-0-3668072-2.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=aab0b83d085e4bac86fdbe1ce65f23dc","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-ready-0-3668072-2.scope","Delegate=yes"],"before":[3668417,3668425],"after":[3668417,3668425],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"ready","delayMs":0,"readyMs":23,"termMs":30,"child":3668509,"self":3668482,"show":["Id=dw-r46-ready-0-3668072-3.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=d076a3b629554b83bded41cb00e667b8","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-ready-0-3668072-3.scope","Delegate=yes"],"before":[3668482,3668509],"after":[3668482,3668509],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
{"mode":"ready","delayMs":0,"readyMs":23,"termMs":30,"child":3668635,"self":3668624,"show":["Id=dw-r46-ready-0-3668072-4.scope","LoadState=loaded","ActiveState=active","SubState=running","InvocationID=6b46a0b279ca454eaa034d48a4de7385","ControlGroup=/user.slice/user-1000.slice/[email protected]/app.slice/dw-r46-ready-0-3668072-4.scope","Delegate=yes"],"before":[3668624,3668635],"after":[3668624,3668635],"childInBefore":true,"childInAfter":true,"exit":"alive 500 ms after TERM"}
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2737.880671ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2557.042479ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (861.841994ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 6282.04986
@@ -0,0 +1,56 @@
✖ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2537.385038ms)
✖ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2331.351365ms)
✖ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (209.932151ms)
ℹ tests 3
ℹ suites 0
ℹ pass 0
ℹ fail 3
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 5448.48373
✖ failing tests:
test at packages/conversation/tests/cohort.test.mjs:139:1
✖ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2537.385038ms)
AssertionError [ERR_ASSERTION]: the escaped child is a listed member
at TestContext.<anonymous> (file:///mnt/storage/src/mosaic-stack/packages/conversation/tests/cohort.test.mjs:151:12)
at async Test.run (node:internal/test_runner/test:1409:7)
at async startSubtestAfterBootstrap (node:internal/test_runner/harness:387:3) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
test at packages/conversation/tests/cohort.test.mjs:176:1
✖ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2331.351365ms)
AssertionError [ERR_ASSERTION]: the member ignored TERM
at TestContext.<anonymous> (file:///mnt/storage/src/mosaic-stack/packages/conversation/tests/cohort.test.mjs:190:12)
at async Test.run (node:internal/test_runner/test:1409:7)
at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
test at packages/conversation/tests/cohort.test.mjs:426:3
✖ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (209.932151ms)
AssertionError [ERR_ASSERTION]: the member is alive across the crash
at TestContext.<anonymous> (file:///mnt/storage/src/mosaic-stack/packages/conversation/tests/cohort.test.mjs:431:12)
at process.processTicksAndRejections (node:internal/process/task_queues:104:5)
at async Test.run (node:internal/test_runner/test:1409:7)
at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2360.366213ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2364.75763ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (628.04725ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 5474.064409
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2457.36226ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2375.081486ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (636.350971ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 5607.100296
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2665.443725ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2531.922418ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (706.098022ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 6068.030335
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2699.60543ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2492.169051ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (675.081525ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 6026.119198
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2762.539396ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2528.214494ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (644.633095ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 6115.728995
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2887.135899ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2915.708256ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (2590.181146ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 9662.912642
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (3084.73509ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2750.8295ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (2151.224665ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 8355.450647
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2746.07054ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2559.848734ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (1985.055632ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 7722.015242
@@ -0,0 +1,11 @@
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2785.421246ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2619.007617ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (889.082424ms)
ℹ tests 3
ℹ suites 0
ℹ pass 3
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 6776.303473
@@ -0,0 +1,56 @@
✖ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2523.134497ms)
✖ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2421.121319ms)
✖ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (211.187938ms)
ℹ tests 3
ℹ suites 0
ℹ pass 0
ℹ fail 3
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 5461.571354
✖ failing tests:
test at cohort.test.mjs:139:1
✖ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2523.134497ms)
AssertionError [ERR_ASSERTION]: the escaped child is a listed member
at TestContext.<anonymous> (file:///home/jwoltje/darkwing-scratch/r46/fix/packages/conversation/tests/cohort.test.mjs:151:12)
at async Test.run (node:internal/test_runner/test:1409:7)
at async startSubtestAfterBootstrap (node:internal/test_runner/harness:387:3) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
test at cohort.test.mjs:176:1
✖ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2421.121319ms)
AssertionError [ERR_ASSERTION]: the member ignored TERM
at TestContext.<anonymous> (file:///home/jwoltje/darkwing-scratch/r46/fix/packages/conversation/tests/cohort.test.mjs:190:12)
at async Test.run (node:internal/test_runner/test:1409:7)
at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
test at cohort.test.mjs:426:3
✖ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (211.187938ms)
AssertionError [ERR_ASSERTION]: the member is alive across the crash
at TestContext.<anonymous> (file:///home/jwoltje/darkwing-scratch/r46/fix/packages/conversation/tests/cohort.test.mjs:431:12)
at process.processTicksAndRejections (node:internal/process/task_queues:104:5)
at async Test.run (node:internal/test_runner/test:1409:7)
at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
@@ -0,0 +1,56 @@
✖ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2545.15719ms)
✖ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2407.231229ms)
✖ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (304.986006ms)
ℹ tests 3
ℹ suites 0
ℹ pass 0
ℹ fail 3
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 5605.846956
✖ failing tests:
test at cohort.test.mjs:139:1
✖ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2545.15719ms)
AssertionError [ERR_ASSERTION]: the escaped child is a listed member
at TestContext.<anonymous> (file:///home/jwoltje/darkwing-scratch/r46/fix/packages/conversation/tests/cohort.test.mjs:151:12)
at async Test.run (node:internal/test_runner/test:1409:7)
at async startSubtestAfterBootstrap (node:internal/test_runner/harness:387:3) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
test at cohort.test.mjs:176:1
✖ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2407.231229ms)
AssertionError [ERR_ASSERTION]: the member ignored TERM
at TestContext.<anonymous> (file:///home/jwoltje/darkwing-scratch/r46/fix/packages/conversation/tests/cohort.test.mjs:190:12)
at async Test.run (node:internal/test_runner/test:1409:7)
at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
test at cohort.test.mjs:426:3
✖ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (304.986006ms)
AssertionError [ERR_ASSERTION]: the member is alive across the crash
at TestContext.<anonymous> (file:///home/jwoltje/darkwing-scratch/r46/fix/packages/conversation/tests/cohort.test.mjs:431:12)
at process.processTicksAndRejections (node:internal/process/task_queues:104:5)
at async Test.run (node:internal/test_runner/test:1409:7)
at async Test.processPendingSubtests (node:internal/test_runner/test:974:7) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: false,
expected: true,
operator: '==',
diff: 'simple'
}
@@ -0,0 +1,160 @@
✔ W1: two processes acquire the same pair at once; exactly one claim (124.186492ms)
✔ W1: two writers publish the same revision at once: one wins, the other gets null, the winner's record stays (11.494379ms)
✔ W1: a revision name appears only after its bytes are synced; before that, only a temp file exists (5.019078ms)
✔ W2: acquire while a claim is reserved or active refuses already-active (169.129732ms)
✔ W3: acquire while stopping, uncertain, or stopped without proof refuses unsafe-replacement (231.756886ms)
✔ W4: same session with another seat tuple, and the reverse, both refuse; a loser on the seat key closes it no-unit (192.859955ms)
✔ W4: a hard link of one session under another seat is the same session: the second controller refuses already-active and launches nothing (35.003128ms)
✔ W4: a copy of one session under another seat is the same session: the second controller refuses already-active and launches nothing (21.132326ms)
✔ W4: a session header ID that changes after construction refuses target; nothing is claimed or launched (3.242382ms)
✔ W5: SIGKILL between every publication barrier of acquire and transition; restart never finds two holders or a lost claim (5566.910492ms)
✔ W5: SIGKILL between every publication barrier of release; restart finishes or holds the release (22659.492505ms)
✔ W6: controller killed mid-turn while the engine lives; restart is uncertain, no launch, prompts refuse (201.806859ms)
✔ W12: a live owner paused with SIGSTOP; a second controller refuses already-active and changes nothing (116.242029ms)
✔ W13: crash after the engine spawns, before active; restart finds the live unit: uncertain, no second spawn, force stop only (262.226787ms)
✔ W14: crash after reservation, before the spawn marker: stopped with a no-unit observation; the pair is free (220.425841ms)
✔ W20: crash after the spawn marker, scope collected; uncertain in both runs, the marker is copied, no launch until a boot proof (240.127173ms)
✔ W15: crash between the two keys during release; restart finishes it under the same claim ID (35.945246ms)
✔ W7: recorded boot ID differs on the same machine: stopped with a boot proof; open tool calls become uncertain (98.424488ms)
✔ W8: resume after a proven stop with the same pins: new claim ID, generation +1, same conversation, branch and leaf (32.141383ms)
✔ W9: resume with a changed binary, argv digest, branch or leaf is refused and the claim is unchanged (89.672997ms)
✔ W11: the controller writes no session file; only the fake engine's own appends appear (23.121134ms)
✔ W16: a highest revision that won't parse holds the pair uncertain; the older stopped revision is not reused (54.682806ms)
✔ W17: a claim root copied from another host refuses foreign-host and promotes nothing (57.719055ms)
✔ G1: a session path or claim root under .pi/state, ~/.claude, the data root or a registration refuses at construction (4.12463ms)
✔ G2: a symlink inside the fixture root to a live session file is refused by the real-path check (1.12384ms)
✔ G3: a fixture path swapped for a live path after construction is refused at bind (1.84306ms)
✔ K1: force stop kills a tool child that called setsid; stopped with a verified proof (2507.919568ms)
✔ K2: K1 on the process-group fallback ends uncertain, never stopped (153.327228ms)
✔ K3: SIGTERM acknowledged while a member lives: stopping until the kill phase, never stopped from TERM (2414.676654ms)
✔ K4: two engines; force stop one; the other survives by independent observation (4314.515513ms)
✔ K5: a stop during a tool call leaves the effect uncertain, and it is shown (2190.378919ms)
✔ K12: a member forking in a loop: the freeze stops it, enumeration is complete, populated 0 after cgroup.kill (2244.796524ms)
✔ K13: a member writing its pid into another cgroup is refused by the namespace; the kill is complete (2189.63935ms)
✔ K15: the shim gone, engine/cgroup.events unreadable, or the engine cgroup missing: evidence unavailable, not empty; uncertain (4473.315326ms)
✔ K10: controller killed between the TERM and kill phases: restart checks the invocation ID and re-runs from TERM for the same stop (482.430287ms)
✔ K11: controller killed after the confirmation is recorded, before TERM: restart checks the invocation ID and re-runs from TERM for the same stop (408.721794ms)
✔ K14: a unit with the recorded name but another invocation ID: evidence unavailable, no signals, uncertain (307.075836ms)
✔ K6: recover without proof, without confirmation, or with changed pins is refused (71.309069ms)
✔ K7: recover after proof, then launch: new claim and execution, generation +1, same leaf; the cancelled prompt is not replayed (35.403696ms)
✔ K8: an engine that loads another leaf on resume is refused before admission; it stays claimed until a proven stop (45.424183ms)
✔ K9: an interrupt that never settles stays uncertain; force stop stays available; takeover is refused while fenced (3030.941578ms)
✔ K16: a claim from another machine ID refuses foreign-host; no boot proof is issued (5.325961ms)
✔ K17: two launcher calls with one eligibility record: one launch, the other refuses, no second engine (30.688596ms)
✔ K18: the leaf changes after eligibility: launch refused; the reservation stays until released with proof (24.893734ms)
✔ S1: `/goal x`, with leading spaces or a tab, refuses text-policy at admission; zero engine bytes (34.176579ms)
✔ S2: every prefix pinned Pi interprets is refused, from the list the code uses; the rest reach the engine exactly (30.425031ms)
✔ S3: `/goal` on the second line is pinned from the source: Pi checks only index 0, so it is admitted and sent exactly (27.293519ms)
✔ S4: a `/` left in the composer is cleared when control transfers and returns; the next submit sends only the new text (42.046162ms)
✔ S5: an observer terminal gets a paste then Enter, as send-message.sh does: not admitted: controller, nothing sent (21.340651ms)
✔ S6: a mediated-shaped registration (no tmux) passed to the board's replyToRow: 409 no tmux session; exec never runs (1.34914ms)
✔ S7: ESC, bracketed-paste markers and U+2028/U+2029 travel as one JSON string; the engine receives the exact text in one record (29.201111ms)
✔ P3: a Pi confirm, select, input or editor dialog is shown disabled with a reason and never answered (138.365864ms)
✔ E1: send, ack, user, toolCall, toolResult, final answer: shown once, no refresh, draft and reading position kept (39.366805ms)
✔ E2: U+2028, U+2029 inside JSON strings and CRLF line ends each parse as one record, on the splitter and through the controller (21.443107ms)
✔ E3: a multipart final, two blocks, null request correlation and duplicate delivery (31.257304ms)
✔ E4: a page read after message_end but before its entry is persisted: marker at the seam, re-read after run-settled, each message once (26.594189ms)
✔ E4: a gap or a new epoch also reconciles; nothing is concatenated across a gap (11.181347ms)
✔ E5: an unknown native event gives no client event; evidence records its type and bytes; the terminal count goes up (26.220793ms)
✔ E6: a tool result delayed across a pause and a reconnect is reconciled without a manual refresh (42.195378ms)
✔ E7: the terminal renders the same stream as the library client, as observer and then as controller, and submits only as controller (41.055698ms)
✔ terminal: engine control characters are made visible; a lost connection refuses submit (28.737455ms)
✔ terminal: outcome unknown is shown as such, with no resend offer, and nothing is resent (0.53917ms)
✔ terminal: text after Enter in the same input chunk starts the next message; it never joins the one submitted (0.351376ms)
✔ terminal: a paste-start marker split right after its ESC still opens the paste; the Enter inside it never submits (0.356358ms)
✔ terminal: invisible and bidi characters are made visible; head, status and notice lines stay one line (0.168609ms)
✔ every record these fixtures produced is a valid CHAT-01 record (E5: no record fails the schema) (359.398761ms)
✔ H1: two takeovers with the same expected generation: one wins, +1; the other refuses generation (56.47649ms)
✔ H2: the old controller's prompt after a takeover commits is refused with zero engine bytes (79.112044ms)
✔ H3: a takeover while a prompt holds the dispatch lock: written under the old actor, or refused; never both (131.150736ms)
✔ H4: self-takeover is refused (22.440288ms)
✔ H9: Interrupt racing a prompt's dispatch: before the write, dispatch-refused and no-turn; after, §3 rules (87.314186ms)
✔ H10: Interrupt and force stop together: one stop chain, force stop supersedes (104.583327ms)
✔ H10: an overlap during the pause before the abort: no abort, the stop ends uncertain (35.050855ms)
✔ H10: a no-turn Interrupt lifts only its own fence; admission stays closed under force stop, overlap or revocation (70.09079ms)
✔ H11: the controller disconnects mid-turn: work continues, the claim is unchanged, control stays put (129.262166ms)
✔ H12: an exact retry after reconnecting to the same incarnation returns the same receipt; one dispatch (15.147609ms)
✔ H13: a retry with the same request ID and different text is refused (13.200342ms)
✔ H14: late stdout from the old engine after a replacement is dropped by incarnation, counted, never rendered (133.123623ms)
✔ H15: a revoked connection's command is refused; the revocation fence holds (65.378107ms)
✔ H16: a second controller for the same session refuses already-active; the first is untouched (16.417034ms)
✔ H10: a second force stop while the first escalation runs refuses fenced; one escalation, and the claim records only the first stop's phases (64.083554ms)
✔ H17: a confirmation reused, answered from another connection, or used after the stop changed is refused (56.969236ms)
✔ H18: two prompts before any native output: the second refuses busy; one engine write (15.033626ms)
✔ H19: the pipe fails mid-line under a large prompt: delivery-unknown transport-unknown, poisoned, no later write (119.397422ms)
✔ H19: the link itself never writes again after an unknown outcome, whoever calls it (0.606652ms)
✔ H19: the controller dies mid-write of a large line: after restart the outcome is unknown and nothing is resent (464.542066ms)
✔ H20: the line is written but the ack is lost when the controller dies: orphan, outcome unknown, nothing resent (357.812591ms)
✔ H21: a retry of the exact request with the old token after a crash is stale-incarnation; no second write (337.123938ms)
✔ H22: after H21 and a valid recovery, a new request with the new token is admitted (2377.836435ms)
✔ H23: requests pending at a restart are not resent; each shows outcome unknown (482.017907ms)
✔ a plain conversation: catalogue row, one page, CHAT-01 records (10.196509ms)
✔ native entries map to blocks: tools, thinking, bash, notices, ids that do not fit (3.091521ms)
✔ F1: a malformed line is an unavailable part at its position, and reading continues (2.999666ms)
✔ F1: a missing parent stops the history with a notice that names the unreadable lines (3.746417ms)
✔ F1: an unreadable fork is never merged into another branch's history (3.733154ms)
✔ F1: a follow stays on its branch when the next entry's parent is unreadable (3.881773ms)
✔ F1: a file whose entries are all unreadable shows a notice per line (1.563335ms)
✔ F2: a truncated trailing line marks the view incomplete, not an error (2.494119ms)
✔ pagination: 100 parts, then the rest; parts concatenate to the whole branch (5.109673ms)
✔ F3: a replaced file (new inode) refuses old cursors with reconcile (4.939351ms)
✔ F4: a same-inode rewrite of the prefix refuses old cursors with reconcile (5.516645ms)
✔ F5: growth between pages keeps the epoch and the page stops at the pinned length (5.793786ms)
✔ F6: unknown, foreign and expired cursors refuse and leave the cursor usable (10.497498ms)
✔ F7: a symlinked file and a symlinked directory component are refused, never opened (9.313922ms)
✔ F8: a file swapped for a symlink after the catalogue is refused (2.247728ms)
✔ F9: registrations never add or redirect a root (2.060491ms)
✔ F10: a header cwd naming another project is refused (3.630852ms)
✔ F11: parentSession renders with a marker and the parent is never opened (0.936984ms)
✔ F12: two leaves: the default leaf is shown and the other branch reads alone (4.789444ms)
✔ F12: a follow refuses when an appended duplicate id changes the branch's earlier parts (2.415895ms)
✔ F12: a second root (Pi's resetLeaf) starts its own branch (1.349225ms)
✔ F13: compaction is a marker in place, then the retained content (0.756245ms)
✔ F14: long strings split into fragments and parts, reassemble exactly, and pages respect the byte cap (737.947726ms)
✔ fragments never cut a surrogate pair and keep an empty string (9.768434ms)
✔ F15: a Claude seat is an unsupported-harness placeholder whose directory is never read (2.629953ms)
✔ unknown conversations, empty files and non-Pi files refuse (2.988026ms)
✔ an unreadable file or root inside the roots is refused per row, not a failed catalogue (1.388415ms)
✔ a seat directory without search permission refuses that root, not the catalogue (2.248255ms)
✔ every page and cursor is a valid CHAT-01 record (847.042088ms)
✔ the engine pin holds for the installed package (2.933293ms)
✔ pinned Pi, sealed and without credentials, answers the controller's commands with the shapes the fake models (364.994239ms)
✔ pinned Pi appends thinking_level_change at start when the branch lacks one, so the leaf moves (K8 then fails closed) (307.845583ms)
✔ N25: ordinary Interrupt reconciles; a non-empty queue_update in the window is O5 (80.815781ms)
✔ N1: an extension's follow-up queued after the fence is cleared before any abort; O5, Unknown (57.148386ms)
✔ N1: a follow-up queued before the fence is O5 at once; the Interrupt refuses fenced (24.835014ms)
✔ N2: with abort first, the fake runs the external item (the ordering guard has teeth) (21.576912ms)
✔ N3: the fence lands in preflight, preflight errors, no run: failed, No run, uncertain (42.350718ms)
✔ N4: the ack arrives after the first abort and a run starts: clear and abort again; Interrupted (37.456929ms)
✔ N5: an input handler takes the prompt: ack, no run, delivery-unknown handled-without-run (119.71406ms)
✔ N6: an extension queues between clear_queue and abort: O5 and O6, Unknown (46.194593ms)
✔ N7: clear_queue times out: no abort, nativeQueue unknown, force stop still ends it (1530.004862ms)
✔ N7: clear_queue answers an error: no abort, nativeQueue unknown, the link not poisoned (15.860593ms)
✔ N8: an extension prompt starts a run during Mosaic preflight; the losing settle is O3 (71.692215ms)
✔ N9: a run that started before the fence and ends aborted: failed interrupted, Interrupted (16.241237ms)
✔ N9: decision 34: a run that ends aborted with no stop in progress: aborted-without-stop, uncertain, outcome unknown (16.048665ms)
✔ N9: an aborted that lands after the fence but before any abort is written: aborted-without-stop, Unknown (32.548668ms)
✔ N10: fake conformance (30.957332ms)
✔ N11: the run fails before any user message_start: delivery-unknown ack-without-start, never failed (29.834953ms)
✔ N12: input that starts a run after the final empty clear is O1 and not part of the stop's proof (19.905915ms)
✔ N13: agent_start with no slot held is O1; a later prompt refuses with zero engine bytes (66.348845ms)
✔ N14: the run completes while clear_queue is in flight: finished, Completed first, uncertain (26.705932ms)
✔ N14: the run completes after the abort is written, before Pi applies it: finished, never relabelled (25.622164ms)
✔ N15: the fence lands in preflight, then an input handler takes it: handled-without-run, No run (19.804122ms)
✔ N16: Interrupt with no slot and no run refuses no-turn: no stop, no bytes, admission open (13.437887ms)
✔ N17: the run fails on its own during the exchange: failed, Failed on its own (28.095743ms)
✔ N18: no final assistant message_end, or a lost line: working stays working; before working, transport-unknown (113.466843ms)
✔ N19: a losing extension prompt settles inside the Mosaic run before its user message: O3, run-overlap (134.279618ms)
✔ N20: an extension triggerTurn during Mosaic preflight starts first; while streaming it queues with no signal (83.211293ms)
✔ N21: a losing settle after the receipt settled finished is O2; the receipt stays finished (14.999139ms)
✔ N22: an agent-level custom message is dropped by the clear with no signal; evidence names the seal (13.807672ms)
✔ N23: a nextTurn message survives clear and abort and attaches to the next prompt, with no signal (13.596018ms)
✔ N24: the seal is an allow-list: --extension, a missing --no-* flag, a second --mode or --session, a session or output flag, or a stray word refuses unsealed-engine; no engine starts (37.274093ms)
ℹ tests 152
ℹ suites 0
ℹ pass 152
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 32488.500969
@@ -0,0 +1,23 @@
✔ browser edge states: loading, empty, malformed, stale, hostile/long values, in-flight reply and appearance fallback (2364.731167ms)
Rendered contrast: {"failures":[],"count":330,"lowest":4.504658476260286}
✔ served Console browser: real board fixtures, keyboard, drafts, receipts, themes, 320px and failures (2725.60223ms)
✔ conversation view: full history, collapsed tools, hidden thinking, inert hostile content, malformed and reconcile markers (2335.250549ms)
✔ conversation view: a fork keeps the open branch, says so, and opens the new one on request (1375.477504ms)
✔ conversation view: a newer session with no readable history keeps the marker (1002.908949ms)
✔ conversation view: seats without history say so and offer no reply (608.42244ms)
✔ Discord row through real board/WebUI: independent brake/liveness, no Reply, literal content (1957.413637ms)
✔ return flow through the conversation view: send, tool call, delayed result, peer message, exact long answers, relaunch (53771.943793ms)
✔ both presentations replace old activity with relaunch notice, label retained history, then resume after new activity (1987.210959ms)
✔ reported return flow and relative Age: reply sent from the inspector, then the new answer appears there without manual refresh (21912.250845ms)
✔ loopback host and board origin fail closed (7.055591ms)
✔ real board fixture passes through WebUI; assets and isolated seen/reply work (100.024626ms)
✔ proxy preserves exact request bytes, status and receipt, rejects forms and malformed JSON, never follows redirect (227.262245ms)
✔ unreachable board reports URL; CLI rejects unsupported options (1632.267153ms)
ℹ tests 14
ℹ suites 0
ℹ pass 14
ℹ fail 0
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 54170.048795
@@ -0,0 +1,2 @@
2026-10-09T13:05:51Z
08:05:51 up 33 days, 9:40, 5 users, load average: 4.03, 7.84, 5.58
@@ -0,0 +1,61 @@
// Row 46 receipt probe. Runs the fixture's ignoreTerm child inside a
// delegated systemd user scope, sends it SIGTERM, and records the scope's
// `systemctl --user show`, `cgroup.procs` before and after TERM, and the
// child's exit code and signal.
//
// node scope-receipt.mjs <mode> <delayMs> <runs>
//
// mode `race`: TERM goes `delayMs` after spawn, as the fixture does today.
// mode `ready`: TERM goes `delayMs` after the child reports its handler is
// installed, as the fixed fixture does.
// The outer process launches one scope per run; the inner process
// (`--inner`) is the scope's main process and prints one JSON line.
import { spawn, spawnSync } from "node:child_process";
import { readFileSync } from "node:fs";
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
// The fixture's child code for { ignoreTerm: true } with no pidLog
// (packages/conversation/tests/fake-pi.mjs, spawnChild), plus a readiness
// byte on stdout in `ready` mode.
const childCode = (ready) =>
"const note=()=>{};process.on('SIGTERM',()=>note('term'));" +
(ready ? "process.stdout.write('r');" : "") +
"setInterval(()=>{},1e9)";
const procs = (cg) => readFileSync(`/sys/fs/cgroup${cg}/cgroup.procs`, "utf8").split("\n").filter(Boolean).map(Number);
async function inner(mode, delayMs, unit) {
const cg = readFileSync("/proc/self/cgroup", "utf8").trim().split("::")[1];
const ready = mode === "ready";
const t0 = performance.now();
const child = spawn(process.execPath, ["-e", childCode(ready)], { stdio: ["ignore", ready ? "pipe" : "ignore", "ignore"] });
const exited = new Promise((r) => child.on("exit", (code, signal) => r({ code, signal, atMs: Math.round(performance.now() - t0) })));
let readyMs = null;
if (ready) {
await new Promise((r) => child.stdout.once("data", r));
readyMs = Math.round(performance.now() - t0);
}
await sleep(delayMs);
const show = spawnSync("systemctl", ["--user", "show", "-p", "Id,LoadState,ActiveState,SubState,InvocationID,ControlGroup,Delegate", `${unit}.scope`], { encoding: "utf8" }).stdout.trim().split("\n");
const before = procs(cg);
const termMs = Math.round(performance.now() - t0);
child.kill("SIGTERM");
const exit = await Promise.race([exited, sleep(500).then(() => null)]);
const after = procs(cg);
if (!exit) child.kill("SIGKILL");
console.log(JSON.stringify({ mode, delayMs, readyMs, termMs, child: child.pid, self: process.pid, show, before, after, childInBefore: before.includes(child.pid), childInAfter: after.includes(child.pid), exit: exit ?? "alive 500 ms after TERM" }));
}
async function outer(mode, delayMs, runs) {
for (let i = 0; i < runs; i++) {
const unit = `dw-r46-${mode}-${delayMs}-${process.pid}-${i}`;
const r = spawnSync("systemd-run", ["--user", "--scope", "--quiet", "-p", "Delegate=yes", `--unit=${unit}`, "--", process.execPath, import.meta.filename, "--inner", mode, String(delayMs), unit], { encoding: "utf8" });
process.stdout.write(r.stdout || `{"unit":"${unit}","status":${r.status},"stderr":${JSON.stringify(r.stderr)}}\n`);
}
}
const a = process.argv.slice(2);
if (a[0] === "--inner") await inner(a[1], Number(a[2]), a[3]);
else await outer(a[0], Number(a[1]), Number(a[2] ?? 5));
@@ -0,0 +1,9 @@
{"t":1791551027129,"ev":"spawn","pid":3646877,"ignoreTerm":true,"setsid":true}
{"t":1791551027142,"ev":"term","pid":3646862}
{"t":1791551027142,"ev":"term","pid":3646877}
{"t":1791551029516,"ev":"spawn","pid":3647263,"ignoreTerm":true,"setsid":false}
{"t":1791551029528,"ev":"term","pid":3647248}
{"t":1791551029528,"ev":"term","pid":3647263}
{"t":1791551031873,"ev":"spawn","pid":3647581,"ignoreTerm":true,"setsid":false}
{"t":1791551031886,"ev":"term","pid":3647556}
{"t":1791551031887,"ev":"term","pid":3647581}
@@ -0,0 +1,13 @@
{"t":1791550812775,"ev":"spawn","pid":3607494,"ignoreTerm":true,"setsid":true}
{"t":1791550812853,"ev":"ready","pid":3607494}
{"t":1791550812878,"ev":"term","pid":3607446}
{"t":1791550812878,"ev":"term","pid":3607494}
{"t":1791550815693,"ev":"spawn","pid":3607807,"ignoreTerm":true,"setsid":false}
{"t":1791550815773,"ev":"ready","pid":3607807}
{"t":1791550815779,"ev":"term","pid":3607781}
{"t":1791550815780,"ev":"term","pid":3607807}
{"t":1791550819446,"ev":"spawn","pid":3608344,"ignoreTerm":true,"setsid":false}
{"t":1791550819489,"ev":"ready","pid":3608344}
{"t":1791550819556,"ev":"term","pid":3608301}
{"t":1791550819556,"ev":"term","pid":3608344}
{"t":1791550820058,"ev":"term","pid":3608344}
@@ -0,0 +1,13 @@
{"t":1791550802745,"ev":"spawn","pid":3605642,"ignoreTerm":true,"setsid":true}
{"t":1791550802763,"ev":"ready","pid":3605642}
{"t":1791550802798,"ev":"term","pid":3605597}
{"t":1791550802798,"ev":"term","pid":3605642}
{"t":1791550805434,"ev":"spawn","pid":3606162,"ignoreTerm":true,"setsid":false}
{"t":1791550805456,"ev":"ready","pid":3606162}
{"t":1791550805492,"ev":"term","pid":3606153}
{"t":1791550805492,"ev":"term","pid":3606162}
{"t":1791550808090,"ev":"spawn","pid":3606502,"ignoreTerm":true,"setsid":false}
{"t":1791550808106,"ev":"ready","pid":3606502}
{"t":1791550808199,"ev":"term","pid":3606487}
{"t":1791550808200,"ev":"term","pid":3606502}
{"t":1791550808539,"ev":"term","pid":3606502}
@@ -0,0 +1,43 @@
diff --git a/packages/conversation/src/shim.mjs b/packages/conversation/src/shim.mjs
index 2adbe48b..d51c4a1f 100644
--- a/packages/conversation/src/shim.mjs
+++ b/packages/conversation/src/shim.mjs
@@ -22,7 +22,7 @@
// `populated 0`. A missing or unreadable file is unavailable, never empty.
import { spawn } from "node:child_process";
-import { closeSync, mkdirSync, readdirSync, readFileSync, unlinkSync, writeFileSync } from "node:fs";
+import { appendFileSync, closeSync, mkdirSync, readdirSync, readFileSync, unlinkSync, writeFileSync } from "node:fs";
import { createServer } from "node:net";
import { join } from "node:path";
import { LineSplitter, encodeLine, parseLine } from "./framing.mjs";
@@ -136,6 +136,7 @@ async function handle(req) {
if (startOf(pid) !== startTicks) continue;
try {
process.kill(pid, "SIGTERM");
+ if (process.env.DW_TRACE) appendFileSync(process.env.DW_TRACE, JSON.stringify({ t: Date.now(), ev: 'term', pid }) + '\n');
signalled.push(pid);
} catch {
// gone already
diff --git a/packages/conversation/tests/fake-pi.mjs b/packages/conversation/tests/fake-pi.mjs
index b3014f86..f30f5a4c 100644
--- a/packages/conversation/tests/fake-pi.mjs
+++ b/packages/conversation/tests/fake-pi.mjs
@@ -615,13 +615,16 @@ const children = [];
// K2); `forkLoop` forks every 5 ms (K12); `ignoreTerm` survives SIGTERM, so
// only the kill phase ends it (K3, K10, K11). With `pidLog`, the fork loop
// appends each child's pid and a `term` line when it gets SIGTERM (K12).
+import * as __fs from 'node:fs';
+const require0 = () => __fs;
function spawnChild({ setsid = false, forkLoop = false, ignoreTerm = false, pidLog = null } = {}) {
const note = pidLog ? `const note=(s)=>require('node:fs').appendFileSync(${JSON.stringify(pidLog)},s+'\\n');` : "const note=()=>{};";
- const code = note + (ignoreTerm ? "process.on('SIGTERM',()=>note('term'));" : "") + (forkLoop
+ const code = note + (ignoreTerm ? "process.on('SIGTERM',()=>note('term'));" + (process.env.DW_TRACE ? `require('node:fs').appendFileSync(${JSON.stringify(process.env.DW_TRACE)},JSON.stringify({t:Date.now(),ev:'ready',pid:process.pid})+'\\n');` : "") : "") + (forkLoop
? "const {spawn}=require('node:child_process');setInterval(()=>{try{const c=spawn('sleep',['1000'],{stdio:'ignore'});if(c.pid)note(String(c.pid))}catch{}},5);setInterval(()=>{},1e9)"
: "setInterval(()=>{},1e9)");
const child = spawn(process.execPath, ["-e", code], { stdio: "ignore", detached: setsid });
children.push(child.pid);
+ if (process.env.DW_TRACE) { require0().appendFileSync(process.env.DW_TRACE, JSON.stringify({ t: Date.now(), ev: 'spawn', pid: child.pid, ignoreTerm, setsid }) + '\n'); child.on('exit', (code, sig) => require0().appendFileSync(process.env.DW_TRACE, JSON.stringify({ t: Date.now(), ev: 'child-exit', pid: child.pid, code, sig }) + '\n')); }
return child.pid;
}
@@ -0,0 +1,28 @@
# Independent acceptance checklist, row 18
Darkwing reviews Filbert's implementation without editing its source candidate.
Dewey reviews visible connector presentation. No live connector manipulation.
- Discovery accepts only safe matching binding name/seat from private regular
files, never dereferences a token path and never serializes private fields.
- Path traversal, symlinked binding/runtime/session paths and malformed records
cannot cause arbitrary reads or an actionable/live row.
- No owner, malformed owner, dead PID, missing identity, reused PID and boot
mismatch are non-live. A positively matching live process is live.
- STOP presence is visible as braked independently of process liveness. Its
contents are not read or exposed; no STOP or lock is created or changed.
- Ordinary completed messages remain idle. No false human attention regression.
- Connector rows cannot borrow a native agent's registration for replies.
Exercise replyToRow and HTTP using a fake executable hook; every connector
attempt must be refused before that hook runs, including with forged tmux
registration. Normal-agent reply tests must still pass.
- Both existing board and WebUI distinguish the connector and brake state and
omit reply controls. Preserve escaping, including hostile binding fixtures.
- Discovery errors disclose no private JSON fields or raw contents. One bad
binding must not silently manufacture a healthy row.
- Candidate pins match before and after tests. Existing dirty attention changes
remain intact; no unrelated source integration or live operation is inferred.
After source approval, measure the real row read-only. Offline/braked behavior
uses isolated fixtures unless the operator separately approves a live-service
transition. Board replacement is its own protected gate.
@@ -0,0 +1,23 @@
{
"at": "2026-09-14T13:51:10.530662+00:00",
"backendPid": 3769124,
"health": "ok",
"row": {
"agent": "sage (discord: shared-signals)",
"project": "fleet",
"state": "idle",
"alive": true,
"connector": {
"binding": "shared-signals",
"braked": false,
"ownerState": "live",
"alive": true
},
"task": "Discord connector",
"taskSource": "connector"
},
"replyStatus": 409,
"replyError": "board replies are disabled for Discord connectors",
"fiveAgentPaneIdentitiesUnchanged": true,
"connectorServiceIdentityUnchanged": true
}
@@ -0,0 +1,2 @@
{"at": "2026-09-14T13:50:24.078636+00:00", "event": "owner-authorized-restart-intent", "oldPid": 3204655, "agents": {"default/darkwing": [["2733924", "12863634"]], "default/dewey": [["934346", "466065"]], "default/filbert": [["72183", "100870"]], "default/researcher": [["173699", "66404285"]], "mosaic-fleet/rocko": [["90599", "128275"]]}, "connectorService": [3022843, "67887873"], "manifest": "254403b89c0a2330da53e8dbad1cbeba3b1b06cf4f3efddc18451e04cb78f6de"}
{"at": "2026-09-14T13:50:24.503231+00:00", "event": "replacement-started", "oldExitedGracefully": true, "newPid": 3769124, "log": "/tmp/discord-board-backend-ovk_cahk.log"}
@@ -0,0 +1,18 @@
{
"at": "2026-09-14T01:10:19.375709+00:00",
"candidate": "/tmp/discord-board-r1-KbMrGQWF",
"manifestSha256": "5c92acc90d202727d790f3fb8d74387db1c9e5c3e43d4f4b56b40e5ae503a56a",
"verdict": "CHANGES REQUIRED",
"independentSerializedTests": 320,
"finding": {
"id": "R1-B1",
"severity": "P2",
"file": "packages/control-board/src/discord.mjs",
"issue": "STOP metadata access errors collapse to absence, falsely projecting not braked",
"reproduction": "Synthetic journal directory contains STOP, chmod directory to 000 as uid 1000, inspectDiscord returns braked:false, ownerState:invalid, alive:false. Restore permissions and remove fixture.",
"expected": "braked:null/unknown when STOP existence cannot be established; false only for verified absence",
"required": "Distinguish missing metadata from access errors and add non-root unreadable-directory regression."
},
"ux": "Dewey APPROVE on exact R1; three independent serialized browser tests passed, source/automation limitations retained",
"parallelQualification": "Two author concurrent frozen timeouts remain unresolved and are not green; serialized independent run passed."
}
@@ -0,0 +1,7 @@
{
"candidate": "R2",
"syntheticOnly": true,
"taskContainsEnvelopeAuthorId": true,
"taskContainsEnvelopeMessageId": true,
"taskSource": "first-user-message"
}
@@ -0,0 +1,18 @@
{
"at": "2026-09-14T01:26:42.688410+00:00",
"candidate": "/tmp/discord-board-r3-U9vVrlQu",
"manifestSha256": "254403b89c0a2330da53e8dbad1cbeba3b1b06cf4f3efddc18451e04cb78f6de",
"reviewer": "Darkwing",
"backendVerdict": "APPROVE AS SOURCE",
"verified": "Nine working/frozen pins, exact three-file R2-to-R3 delta, inherited attention pins and full serialized six-package suite 322/322",
"findingsClosed": [
"R1-B1: inaccessible STOP is unknown, non-root regression passes",
"R2-B2: canonical routing envelope no longer becomes connector Task; ordinary fallback retained"
],
"limitations": [
"No generalized transcript redaction",
"R1 concurrent combined frozen timeouts unresolved/not green",
"No live observation, backend restart, connector change or publication in this review"
],
"uxGate": "Await exact R3 confirmation from Dewey via agent-send"
}
@@ -0,0 +1,20 @@
{
"at": "2026-09-14T01:28:17.695Z",
"sourceApproval": 26257,
"readOnly": true,
"agent": "sage (discord: shared-signals)",
"project": "fleet",
"state": "idle",
"alive": true,
"connector": {
"binding": "shared-signals",
"braked": false,
"ownerState": "live",
"alive": true
},
"task": "Discord connector",
"taskSource": "connector",
"registrationAbsent": true,
"ownerMatchesService": true,
"discoveryErrorCount": 0
}
@@ -0,0 +1,9 @@
{
"at": "2026-09-14T00:50:35.339Z",
"sourceSha256": "dfbb7ab9374c0ac9fafa0503f495abd938f499f5d6227233de03a60ea3022927",
"fixture": "connector row with forged native registration",
"status": 200,
"fakeTransportCalls": 1,
"realTransportCalls": 0,
"gatePassed": false
}
@@ -0,0 +1,162 @@
# Discord engine: guaranteed test cleanup and the timeout gap in `busy` (#1509), R2 candidate
Sage assigned this on 2026-09-26 as 6b, source only. Rocko reviews. Base is
HEAD 401cc850. Not committed. The live connector runs from this checkout, so
Sage is holding its restart until this is approved and committed. Nobody
should restart it from a working copy.
## Defects (DEFERRED Open, "Discord engine: leaked fake pi…")
(a) `engine.test.mjs` read the fake's `commands.jsonl` 20 ms after a prompt and
got ENOENT under load. Seven tests stopped the engine outside `finally`, so a
failed assertion left the fake pi running and the test file never exited.
(b) `busy` was `state.busy || pending.some((t) => !t.done)`. If a turn timed out
before its `agent_start` was read, it was done while `state.busy` was still
false. The next prompt then went straight to pi, which refused it as
streaming.
## R1 and Rocko's finding
R1 held later prompts behind a failed turn. If pi had sent no `agent_start` for
it within a grace period, R1 dropped that turn from the queue and sent the next
prompt. Rocko rejected it (F1, High), in
`agents/rocko/work/discord-engine-busy-r1-review-2026-09-26.md`, sha256
047dbd8f.
Pi's events carry no prompt id. The engine attributes them to the front of its
queue. Silence until the grace ends does not prove the old run will never come.
If pi then runs it, its events land on the new prompt, which R1 had just put at
the front. Rocko's reproducer got the old run's answer and its `old.md` tool
record back as the new prompt's result. My R1 README said such events "find no
live head and are dropped". That was wrong.
The R1 files stay here as `r1-manifest.sha256` and `r1.patch`.
## Change (R2)
`packages/discord/src/engine-pi.mjs`:
- `engineBusy()` is `state.busy || state.pending.length > 0`. A failed turn
still in the queue holds the next prompt back, and stays at the front, so any
late events for it land on it. `prompt()`, `sendHeld()` and the `busy` getter
use it. This part is unchanged from R1.
- The bound is now a stop, not a drop. When a turn fails while it is still in
the queue, `failTurn` starts a timer, `abortGraceMs` (default
`ABORT_GRACE_MS`, 30 s, an engine option, not binding config). When it
fires:
- If pi has sent `agent_start` (`state.busy`), nothing happens. That run
ends on its `agent_end` or a settle, as on HEAD.
- Otherwise `wedge()` sets `state.wedged`, fails every held prompt with
code `engine-wedged`, and stops pi: stdin closed, SIGTERM, then SIGKILL
after 5 s. The failed turn stays at the front until the exit.
- While wedged, nothing is written to that child. `write()`, `sendHeld()` and
`prompt()` refuse, and a new prompt fails at once with `engine-down`. The exit
runs the usual `failAll` and `onExit`.
- `stop()`'s body moved into `stopChild()`, which both `stop()` and `wedge()`
call.
- `release()` clears the timer wherever a turn leaves the queue: `agent_end`,
settle, a refused send, and process exit. As in R1, the settle handler removes
turns before failing them.
What recovery looks like live: `cli.mjs` handles `onExit` with `shutdown(1)`.
The unit's `Restart=on-failure` starts a new connector and a new pi 15 s later,
within its limit of five tries in ten minutes. This change doesn't touch the
unit or the restart policy. A wedge now costs one connector restart. R1 would
have kept the same pi and risked a wrong answer.
`packages/discord/tests/fake-pi.mjs`:
- `mute`: accepted and never run.
- `stall <ms>`: accepted, then the fake reads nothing for `<ms>`, runs the
stalled prompt, and only then reads what came in meanwhile. This is the
order in Rocko's case.
`packages/discord/tests/engine.test.mjs`:
- `withEngine()` stops the engine in `finally`. Every test that starts an
engine uses it, or has its own `try/finally` in the exit test.
- `commands()` returns `[]` until the fake creates its log. The held-prompt
test waits for the first prompt with `until()` instead of a 20 ms sleep, and
its first prompt is `slow 300`.
- The manual-timer test fires the turn timer before any pi event is read. The
next prompt must wait for the settle and get its own answer.
- New or changed for R2:
- `mute` with `abortGraceMs: 150`. The held prompt fails with
`engine-wedged` after the grace, a later prompt fails with
`engine-down`, `onExit` fires, and pi saw only `mute` and `abort`.
- `stall 400` with the same grace, which is Rocko's case with a real
child. The held prompt fails with `engine-wedged`, pi exits, and "after
stall" never reaches pi.
- Rocko's reproducer as an in-memory test, run twice. The old prompt's
response comes either before its timeout or only with the late events.
After the grace, the old run's start, tool pair, answer, end and settle
arrive while pi is still exiting. The held prompt stays failed with
`engine-wedged` and a later prompt fails with `engine-down`. Pi saw only
`old` and `abort`, then SIGTERM, then SIGKILL at 5 s. Only the exit
reaches `onExit`. The test reads recorded outcomes after a tick instead
of awaiting, so a regression fails instead of hanging.
- `late 400` with the same grace. Pi started that run, so the grace does
not stop pi, and the next prompt gets its own answer when the run ends.
## Evidence
- `engine.test.mjs`: 17/17.
- R1's engine (d5bf24b5, from `r1.patch`) against these tests fails 4: `mute`,
`stall`, and both in-memory runs. In that run a probe shows R1 answering
"after stall" with "echo: stalled". An earlier draft of the in-memory test
awaited the held prompt and hung on R1 until the 120 s cap. It now fails in
milliseconds.
- HEAD's engine against these tests fails 5: the manual-timer test and the
same four.
- Mutations of R2:
- Without the `state.busy` check, the `late 400` test fails.
- Without the `wedge()` call, 4 fail.
- Without failing held prompts in `wedge()`, 4 fail.
- Test union (control-board, webui, seat, mosaic, ledger, discord) at default
concurrency on `git archive` of 401cc850 plus the three files: 406/406
three times, 23 to 24 s each. No fake pi was left running.
- Eight suites green on that snapshot: config 24, task 90, foundation 43,
conductor 17, release 14, auth 15, discord 63, extension-package 18.
Logs: `/tmp/dw-6b-r2-conc-{1,2,3}.txt`. R1's evidence runs:
`/tmp/dw-6b-conc-{1,2,3}.txt`, `/tmp/dw-6b-serial.txt`. HEAD's hang control:
`/tmp/dw-1509-headctl-{1,2,3}.txt`, `/tmp/dw-1509-ef00-1.txt`. There, HEAD hit
the 240 s cap at 199 ok under the union's load.
## Not covered
- A run pi started and never ends, even after abort, still holds prompts
until pi settles or exits. Each held prompt fails at its own timeout ("while
waiting for the engine"). HEAD behaves the same way through `state.busy`, and
Rocko did not block on it. Only a pi restart clears it.
- A wedge ends the connector process, and the recovery is systemd's restart.
Nothing here changes the unit, and the restart limit still applies.
- No live restart, and no change to the binding schema.
## Frozen files
`r2-manifest.sha256` holds the three R2 hashes. `r2.patch` is `git diff
packages/discord` at freeze time.
## Review
Rocko, R1, 2026-09-26: request changes, F1 High, as described above. Report:
`agents/rocko/work/discord-engine-busy-r1-review-2026-09-26.md`, sha256
047dbd8f.
Rocko, R2, 2026-09-26: approved the three pinned files. Report:
`agents/rocko/work/discord-engine-busy-r2-review-2026-09-26.md`, sha256
ed5510a0. He checked the manifests before and after, ran 17/17 himself, and
read the CLI shutdown path, `connector.stop` and the unit template. Sage asked
him three operational questions:
- A wedge exits 1, never 3. Exit 3 remains the supervised startup refusal.
- The unit's start limit (5 starts in 600 s) is a rate limit. It does not
bound repeated wedges. With the default 180 s turn timeout, the 30 s grace
and the 15 s restart delay, a cycle takes at least 225 s. That stays under
the limit, so a pi that wedges every time could restart indefinitely.
Stopping for good after repeated wedges would need a separate policy. This
change does not add one.
- He recommends, as a nonblocking follow-up, that the connector journal
record at startup: HEAD, dirty state scoped to runtime source, and a digest
of the runtime files. A wedge restart loads whatever the checkout holds.
This section was added after approval, so the README hash no longer matches
the one Rocko pinned (69350f29). The three source files are unchanged.
@@ -0,0 +1,3 @@
d5bf24b59c07c85067f4087c03b54ca8b4df923c1d591441dedd2e8a7ff2ae39 packages/discord/src/engine-pi.mjs
f0abee9c243d46d66dd2271abc3fd89089c350ac6a66ab49131bce80adfcdc33 packages/discord/tests/engine.test.mjs
fa1bf44e3f33eb970a714ada1c686abbc1932baaf418679276edf8813abbe6de packages/discord/tests/fake-pi.mjs
@@ -0,0 +1,414 @@
diff --git a/packages/discord/src/engine-pi.mjs b/packages/discord/src/engine-pi.mjs
index 5c8fd0a9..9dcaeb42 100644
--- a/packages/discord/src/engine-pi.mjs
+++ b/packages/discord/src/engine-pi.mjs
@@ -16,9 +16,12 @@
// from `tool_execution_start`/`tool_execution_end` into the result so the
// turn record shows what was read. An `agent_end` with `willRetry` is not
// the end of the run. A timeout sends `abort` and fails that turn; the
-// process stays. A malformed JSONL line from pi fails the current turn (its
-// outcome is now unknowable) and the process stays. Process exit fails
-// every pending turn and is reported through `onExit`.
+// process stays. The failed turn holds later prompts back until its
+// agent_end or a settle. If pi has not started it within ABORT_GRACE_MS, it
+// is dropped and the next prompt goes out; a run pi did start holds them
+// until it ends, as any run does. A malformed JSONL line from pi fails the
+// current turn (its outcome is now unknowable) and the process stays.
+// Process exit fails every pending turn and is reported through `onExit`.
//
// Framing follows pi's RPC doc: split on "\n" only, strip a trailing "\r".
// Node readline is not used because it also splits on U+2028/U+2029.
@@ -64,12 +67,19 @@ export function assistantText(message) {
.trim();
}
+// How long a turn that failed here (timeout, protocol error) may wait for
+// pi's agent_start before it stops holding the next prompt back. Without a
+// bound, a prompt pi accepted but never ran would queue every later prompt
+// until restart.
+export const ABORT_GRACE_MS = 30000;
+
export function createEngine({
command, args, cwd, env = {},
spawn = nodeSpawn,
setTimeoutImpl = globalThis.setTimeout, clearTimeoutImpl = globalThis.clearTimeout,
log = () => {},
onExit = () => {},
+ abortGraceMs = ABORT_GRACE_MS,
} = {}) {
if (typeof command !== "string" || command.length === 0) throw new DiscordError("engine: command required", 1);
if (!Array.isArray(args)) throw new DiscordError("engine: args required", 1);
@@ -80,15 +90,35 @@ export function createEngine({
// A turn that fails on the client side (timeout, protocol error) stays in
// the pending queue, marked done, until pi's own turn_end for it arrives.
- // Otherwise that turn_end would be attributed to the next prompt.
+ // Otherwise that turn_end would be attributed to the next prompt. It holds
+ // later prompts back; if pi has not started it within abortGraceMs, it goes.
function failTurn(turn, code, message) {
if (turn.done) return;
turn.done = true;
if (turn.timer !== null) clearTimeoutImpl(turn.timer);
turn.timer = null;
+ if (state.pending.includes(turn)) {
+ turn.grace = setTimeoutImpl(() => {
+ turn.grace = null;
+ // No agent_start by now: pi never started this run and will send no
+ // agent_end for it, so it leaves the queue and cannot take the next
+ // prompt's. A run pi did start keeps its place until it ends.
+ if (!state.busy) {
+ const i = state.pending.indexOf(turn);
+ if (i !== -1) state.pending.splice(i, 1);
+ }
+ sendHeld();
+ }, abortGraceMs);
+ }
turn.reject(new DiscordError(message, 1, { code }));
}
+ // Call when a turn leaves the pending queue.
+ function release(turn) {
+ if (turn.grace !== null) clearTimeoutImpl(turn.grace);
+ turn.grace = null;
+ }
+
function settleTurn(turn, value) {
if (turn.done) return;
turn.done = true;
@@ -99,7 +129,10 @@ export function createEngine({
function failAll(code, message) {
const pending = state.pending.splice(0);
- for (const t of pending) failTurn(t, code, message);
+ for (const t of pending) {
+ release(t);
+ failTurn(t, code, message);
+ }
for (const h of state.held.splice(0)) failTurn(h.turn, code, message);
for (const [, r] of state.responses) r.reject(new DiscordError(message, 1, { code }));
state.responses.clear();
@@ -174,6 +207,7 @@ export function createEngine({
// Attribute the run to the head even if it failed client-side, so the
// next prompt's agent_end is not taken for this one.
const run = state.pending.shift();
+ if (run) release(run);
if (!run || run.done) return;
const messages = Array.isArray(event.messages) ? event.messages.filter((m) => m && m.role === "assistant") : [];
const message = messages.length > 0 ? messages[messages.length - 1] : run.last;
@@ -193,13 +227,14 @@ export function createEngine({
// this settle and still has no agent_end will never get one: fail it now
// instead of waiting for its timeout. Turns whose prompt response has
// not arrived yet belong to a later run and stay.
+ const dropped = [];
const keep = [];
- for (const t of state.pending) {
- if (t.done) continue;
- if (t.accepted) failTurn(t, "engine-settled-without-turn", "engine settled without answering this prompt");
- else keep.push(t);
- }
+ for (const t of state.pending) (t.done || t.accepted ? dropped : keep).push(t);
state.pending = keep;
+ for (const t of dropped) {
+ release(t);
+ failTurn(t, "engine-settled-without-turn", "engine settled without answering this prompt");
+ }
sendHeld();
}
}
@@ -214,15 +249,23 @@ export function createEngine({
// Never accepted: pi will not emit a turn_end for it, so remove it.
const i = state.pending.indexOf(turn);
if (i !== -1) state.pending.splice(i, 1);
+ release(turn);
failTurn(turn, (err.details && err.details.code) || "engine-refused", err.message);
sendHeld();
});
}
+ // Pi is busy from our side while any sent prompt is still queued, even one
+ // that already failed here: a turn that timed out before its agent_start
+ // was read leaves state.busy false while pi runs it, and sending then would
+ // be refused as streaming. It leaves the queue on its agent_end, on a
+ // settle, on a refused send, or when its grace ends before pi started it.
+ const engineBusy = () => state.busy || state.pending.length > 0;
+
// After a settle (or a refused send) the oldest held prompt goes out.
function sendHeld() {
if (state.exited !== null) return;
- if (state.busy || state.pending.some((t) => !t.done)) return;
+ if (engineBusy()) return;
const next = state.held.shift();
if (next) send(next.turn, next.command);
}
@@ -281,7 +324,7 @@ export function createEngine({
// with DiscordError carrying details.code for the turn record.
prompt(text, { timeoutMs = 180000 } = {}) {
if (typeof text !== "string" || text.length === 0) throw new DiscordError("prompt text required", 1);
- const turn = { resolve: null, reject: null, timer: null, done: false, accepted: false, tools: new Map(), turns: 0, last: null };
+ const turn = { resolve: null, reject: null, timer: null, grace: null, done: false, accepted: false, tools: new Map(), turns: 0, last: null };
const done = new Promise((resolve, reject) => {
turn.resolve = resolve;
turn.reject = reject;
@@ -310,13 +353,13 @@ export function createEngine({
failTurn(turn, "engine-down", "engine is not running");
return done;
}
- if (state.busy || state.pending.some((t) => !t.done) || state.held.length > 0) state.held.push({ turn, command });
+ if (engineBusy() || state.held.length > 0) state.held.push({ turn, command });
else send(turn, command);
return done;
},
get busy() {
- return state.busy || state.pending.some((t) => !t.done) || state.held.length > 0;
+ return engineBusy() || state.held.length > 0;
},
get pendingCount() {
return state.pending.filter((t) => !t.done).length + state.held.length;
diff --git a/packages/discord/tests/engine.test.mjs b/packages/discord/tests/engine.test.mjs
index 59674d1e..62dd6017 100644
--- a/packages/discord/tests/engine.test.mjs
+++ b/packages/discord/tests/engine.test.mjs
@@ -28,7 +28,20 @@ function start(root, extra = {}) {
log: (m) => logs.push(m), ...extra,
});
engine.start();
- return { engine, logs, commands: () => readFileSync(logPath, "utf8").trim().split("\n").filter(Boolean).map((l) => JSON.parse(l)) };
+ // The fake creates its log on the first command; until then there are none.
+ const commands = () => (existsSync(logPath) ? readFileSync(logPath, "utf8").trim().split("\n").filter(Boolean).map((l) => JSON.parse(l)) : []);
+ return { engine, logs, commands };
+}
+
+// Every test stops its engine in finally: a fake pi left running after a
+// failed assertion keeps the test file from exiting.
+async function withEngine(extra, body) {
+ const started = start(makeRoot(), extra);
+ try {
+ await body(started);
+ } finally {
+ await started.engine.stop();
+ }
}
test("engine: buildPiArgs carries the fixed flags, engine settings, session dir and prompt file", () => {
@@ -56,8 +69,7 @@ test("engine: with tools, buildPiArgs turns pi's own tools off, loads the extens
assert.equal(rw[rw.indexOf("--tools") + 1], "list_dir,read_file,search,write_file,edit_file", "a writable root adds exactly the two write tools");
});
-test("engine: a run with tool turns settles once, on the answer, with every tool call in the result", async () => {
- const { engine } = start(makeRoot());
+test("engine: a run with tool turns settles once, on the answer, with every tool call in the result", () => withEngine({}, async ({ engine }) => {
const r = await engine.prompt("tools 3");
assert.equal(r.text, "read 3 file(s)");
assert.equal(r.turns, 2);
@@ -71,45 +83,38 @@ test("engine: a run with tool turns settles once, on the answer, with every tool
assert.equal(plain.turns, 1);
await idle(engine);
assert.equal(engine.busy, false);
- await engine.stop();
-});
+}));
-test("engine: a run that ends on a tool-only turn fails the prompt as empty; a retried run settles on the real end", async () => {
- const { engine } = start(makeRoot());
+test("engine: a run that ends on a tool-only turn fails the prompt as empty; a retried run settles on the real end", () => withEngine({}, async ({ engine }) => {
const r = await engine.prompt("toolonly");
assert.equal(r.text, "", "no text: the connector turns this into engine-empty");
assert.equal(r.tools.length, 1);
const again = await engine.prompt("retry");
assert.equal(again.text, "after retry");
- await engine.stop();
-});
+}));
-test("engine: one prompt, one turn, text and usage come back", async () => {
- const { engine } = start(makeRoot());
- try {
- const r = await engine.prompt("hello");
- assert.equal(r.text, "echo: hello");
- assert.deepEqual(r.usage, { input: 3, output: 2 });
- await idle(engine);
- assert.equal(engine.busy, false);
- } finally {
- await engine.stop();
- }
-});
+test("engine: one prompt, one turn, text and usage come back", () => withEngine({}, async ({ engine }) => {
+ const r = await engine.prompt("hello");
+ assert.equal(r.text, "echo: hello");
+ assert.deepEqual(r.usage, { input: 3, output: 2 });
+ await idle(engine);
+ assert.equal(engine.busy, false);
+}));
-test("engine: a prompt while streaming is held until pi settles, then sent as its own run, and answered in order", async () => {
- const { engine, commands } = start(makeRoot());
- const first = engine.prompt("slow 150");
- await new Promise((r) => setTimeout(r, 20));
+test("engine: a prompt while streaming is held until pi settles, then sent as its own run, and answered in order", () => withEngine({}, async ({ engine, commands }) => {
+ const first = engine.prompt("slow 300");
assert.equal(engine.busy, true);
const second = engine.prompt("second");
assert.equal(engine.pendingCount, 2);
- await new Promise((r) => setTimeout(r, 20));
- assert.equal(commands().filter((c) => c.type === "prompt").length, 1, "the second prompt is not sent while pi is busy");
+ const prompted = () => commands().filter((c) => c.type === "prompt");
+ assert.ok(await until(() => prompted().length > 0), "the first prompt reached pi");
+ assert.equal(prompted().length, 1, "the second prompt is not sent while pi is busy");
+ // The fake refuses a prompt without streamingBehavior while it runs one, so
+ // an answered second prompt also proves it was not sent early.
const [r1, r2] = await Promise.all([first, second]);
assert.equal(r1.text, "slow reply");
assert.equal(r2.text, "echo: second");
- const prompts = commands().filter((c) => c.type === "prompt");
+ const prompts = prompted();
assert.equal(prompts.length, 2);
// Never a pi follow-up: pi would fold it into the first run and close both
// answers with one agent_end (the live loss of 2026-09-17).
@@ -117,11 +122,9 @@ test("engine: a prompt while streaming is held until pi settles, then sent as it
assert.equal(prompts[1].streamingBehavior, undefined);
await idle(engine);
assert.equal(engine.busy, false);
- await engine.stop();
-});
+}));
-test("engine: a held prompt that times out before pi settles fails on its own and is never sent", async () => {
- const { engine, commands } = start(makeRoot());
+test("engine: a held prompt that times out before pi settles fails on its own and is never sent", () => withEngine({}, async ({ engine, commands }) => {
const first = engine.prompt("slow 200");
await new Promise((r) => setTimeout(r, 20));
await assert.rejects(engine.prompt("late one", { timeoutMs: 50 }), (e) => e.details.code === "timeout" && /waiting for the engine/.test(e.message));
@@ -130,50 +133,85 @@ test("engine: a held prompt that times out before pi settles fails on its own an
await idle(engine);
assert.deepEqual(commands().filter((c) => c.type === "prompt").map((c) => c.message), ["slow 200"]);
assert.deepEqual(commands().filter((c) => c.type === "abort"), [], "a held turn is not aborted; pi never had it");
- await engine.stop();
-});
+}));
-test("engine: timeout sends abort and fails only that turn; the process stays", async () => {
- const { engine, commands, logs } = start(makeRoot());
+test("engine: timeout sends abort and fails only that turn; the process stays", () => withEngine({}, async ({ engine, commands, logs }) => {
await assert.rejects(engine.prompt("slow 5000", { timeoutMs: 100 }), (err) => err.details.code === "timeout");
assert.ok(await until(() => commands().some((c) => c.type === "abort")), "abort reached pi");
assert.ok(logs.some((l) => /timed out/.test(l)));
const r = await engine.prompt("again");
assert.equal(r.text, "echo: again");
- await engine.stop();
+}));
+
+test("engine: tool events from a run that outlived its timeout never land in the next prompt's record", () => withEngine({}, async ({ engine }) => {
+ await assert.rejects(engine.prompt("late 200", { timeoutMs: 40 }), (err) => err.details.code === "timeout");
+ const r = await engine.prompt("after late");
+ assert.equal(r.text, "echo: after late");
+ assert.deepEqual(r.tools, [], "the dead run's read is not this prompt's evidence");
+ assert.equal(r.turns, 1, "the dead run's turns are not counted here");
+}));
+
+// The turn timer is fired by hand, before the engine has read any event from
+// pi, so the timed-out run is still pi's and state.busy is still false when
+// the next prompt arrives. Under load a real timer does the same.
+const TURN_MS = 60000;
+const manualTurnTimer = (fire) => ({
+ setTimeoutImpl: (fn, ms) => (ms === TURN_MS ? fire.push(fn) : setTimeout(fn, ms)),
+ clearTimeoutImpl: (id) => { if (typeof id !== "number") clearTimeout(id); },
});
-test("engine: tool events from a run that outlived its timeout never land in the next prompt's record", async () => {
- const { engine } = start(makeRoot());
- try {
- await assert.rejects(engine.prompt("late 200", { timeoutMs: 40 }), (err) => err.details.code === "timeout");
- const r = await engine.prompt("after late");
+test("engine: a prompt after a turn that timed out before its agent_start waits for pi to settle instead of being refused", () => {
+ const fire = [];
+ return withEngine(manualTurnTimer(fire), async ({ engine, commands }) => {
+ const late = engine.prompt("late 100", { timeoutMs: TURN_MS });
+ fire.shift()();
+ assert.equal(engine.busy, true, "pi is still running the prompt that timed out");
+ const next = engine.prompt("after late", { timeoutMs: 5000 });
+ assert.equal(engine.pendingCount, 1, "only the new prompt is live");
+ await assert.rejects(late, (err) => err.details.code === "timeout");
+ const r = await next;
assert.equal(r.text, "echo: after late");
assert.deepEqual(r.tools, [], "the dead run's read is not this prompt's evidence");
- assert.equal(r.turns, 1, "the dead run's turns are not counted here");
- } finally {
- await engine.stop();
- }
+ assert.equal(r.turns, 1);
+ assert.deepEqual(commands().map((c) => (c.type === "prompt" ? c.message : c.type)), ["late 100", "abort", "after late"]);
+ await idle(engine);
+ assert.equal(engine.busy, false);
+ });
});
-test("engine: a malformed JSONL line fails the turn, not the process", async () => {
- const { engine, logs } = start(makeRoot());
+// "mute" is accepted and never run, so no agent_start, agent_end or settle
+// ever comes for it. Unbounded, it would hold every later prompt.
+test("engine: a timed-out turn pi never started holds the next prompt only for the abort grace, then leaves the queue", () => withEngine({ abortGraceMs: 150 }, async ({ engine, commands }) => {
+ await assert.rejects(engine.prompt("mute", { timeoutMs: 50 }), (err) => err.details.code === "timeout");
+ assert.equal(engine.busy, true, "pi might still be running it");
+ const started = Date.now();
+ const r = await engine.prompt("after mute", { timeoutMs: 5000 });
+ assert.equal(r.text, "echo: after mute");
+ assert.ok(Date.now() - started >= 100, "held for the grace, not sent at once");
+ assert.deepEqual(commands().map((c) => (c.type === "prompt" ? c.message : c.type)), ["mute", "abort", "after mute"]);
+ await idle(engine);
+ assert.equal(engine.busy, false);
+}));
+
+test("engine: a malformed JSONL line fails the turn, not the process", () => withEngine({}, async ({ engine, logs }) => {
await assert.rejects(engine.prompt("garbage"), (err) => err.details.code === "engine-protocol");
assert.ok(logs.some((l) => /malformed/.test(l)));
const r = await engine.prompt("still here");
assert.equal(r.text, "echo: still here");
- await engine.stop();
-});
+}));
test("engine: a turn that ends in error rejects with the error code; process exit fails pending turns", async () => {
- const root = makeRoot();
let exited = null;
- const { engine } = start(root, { onExit: (e) => (exited = e) });
- await assert.rejects(engine.prompt("error"), (err) => err.details.code === "engine-error" && /fake provider error/.test(err.message));
- const pending = engine.prompt("slow 5000");
- await new Promise((r) => setTimeout(r, 20));
- await engine.stop();
- await assert.rejects(pending, (err) => err.details.code === "engine-down");
- assert.ok(exited);
- await assert.rejects(engine.prompt("x"), /not running/);
+ const { engine } = start(makeRoot(), { onExit: (e) => (exited = e) });
+ try {
+ await assert.rejects(engine.prompt("error"), (err) => err.details.code === "engine-error" && /fake provider error/.test(err.message));
+ const pending = engine.prompt("slow 5000");
+ await new Promise((r) => setTimeout(r, 20));
+ await engine.stop();
+ await assert.rejects(pending, (err) => err.details.code === "engine-down");
+ assert.ok(exited);
+ await assert.rejects(engine.prompt("x"), /not running/);
+ } finally {
+ await engine.stop();
+ }
});
diff --git a/packages/discord/tests/fake-pi.mjs b/packages/discord/tests/fake-pi.mjs
index 94919306..ecc04e3b 100644
--- a/packages/discord/tests/fake-pi.mjs
+++ b/packages/discord/tests/fake-pi.mjs
@@ -8,6 +8,7 @@
// then a second turn that answers "read <n> file(s)"
// "toolonly" a run whose only turn calls a tool and never answers
// "retry" an agent_end with willRetry, then the real answer
+// "mute" accept the prompt and emit nothing, staying idle
// "late <ms>" ignore abort; after <ms> emit a tool pair and a tool turn,
// then answer "late reply", like a run that outlives its
// client-side timeout
@@ -30,6 +31,7 @@ function assistant(text, stopReason = "stop") {
}
function run(text) {
+ if (text === "mute") return;
busy = true;
out({ type: "agent_start" });
out({ type: "turn_start" });
@@ -0,0 +1,3 @@
77077b7fbd5a933ffd352094eb073227c299ba47b7aea52d4e60fdc55cc7101e packages/discord/src/engine-pi.mjs
47a998179c6eb46827f43ab2c6b0f6b062ef94fb47da402cfb9af8c4f180f38e packages/discord/tests/engine.test.mjs
a8e54cc3f4b670eef2c06755b63e9e6bfeb44b1b1efde3aca9bfaa91583c3ef3 packages/discord/tests/fake-pi.mjs
@@ -0,0 +1,651 @@
diff --git a/packages/discord/src/engine-pi.mjs b/packages/discord/src/engine-pi.mjs
index 5c8fd0a9..46ef1f88 100644
--- a/packages/discord/src/engine-pi.mjs
+++ b/packages/discord/src/engine-pi.mjs
@@ -16,9 +16,14 @@
// from `tool_execution_start`/`tool_execution_end` into the result so the
// turn record shows what was read. An `agent_end` with `willRetry` is not
// the end of the run. A timeout sends `abort` and fails that turn; the
-// process stays. A malformed JSONL line from pi fails the current turn (its
-// outcome is now unknowable) and the process stays. Process exit fails
-// every pending turn and is reported through `onExit`.
+// process stays. The failed turn holds later prompts back until its
+// agent_end or a settle. If pi has not started it within ABORT_GRACE_MS, the
+// engine stops pi instead of sending again: pi's events carry no prompt id,
+// so a late run of the failed prompt would be taken for the next one's. A
+// run pi did start holds later prompts until it ends, as any run does. A
+// malformed JSONL line from pi fails the current turn (its outcome is now
+// unknowable) and the process stays. Process exit fails every pending turn
+// and is reported through `onExit`.
//
// Framing follows pi's RPC doc: split on "\n" only, strip a trailing "\r".
// Node readline is not used because it also splits on U+2028/U+2029.
@@ -64,31 +69,67 @@ export function assistantText(message) {
.trim();
}
+// How long a turn that failed here (timeout, protocol error) may wait for
+// pi's agent_start before the engine stops pi. Without a bound, a prompt pi
+// accepted but never ran would hold every later prompt until restart.
+export const ABORT_GRACE_MS = 30000;
+
export function createEngine({
command, args, cwd, env = {},
spawn = nodeSpawn,
setTimeoutImpl = globalThis.setTimeout, clearTimeoutImpl = globalThis.clearTimeout,
log = () => {},
onExit = () => {},
+ abortGraceMs = ABORT_GRACE_MS,
} = {}) {
if (typeof command !== "string" || command.length === 0) throw new DiscordError("engine: command required", 1);
if (!Array.isArray(args)) throw new DiscordError("engine: args required", 1);
// pending: prompts sent to pi, oldest first. held: prompts waiting for pi
// to settle before they are sent, oldest first.
- const state = { child: null, buffer: "", pending: [], held: [], responses: new Map(), nextId: 1, busy: false, exited: null };
+ // wedged: set when the engine gave up on pi and is stopping it. Nothing
+ // is sent to that child again.
+ const state = { child: null, buffer: "", pending: [], held: [], responses: new Map(), nextId: 1, busy: false, exited: null, wedged: false };
// A turn that fails on the client side (timeout, protocol error) stays in
// the pending queue, marked done, until pi's own turn_end for it arrives.
- // Otherwise that turn_end would be attributed to the next prompt.
+ // Otherwise that turn_end would be attributed to the next prompt. It holds
+ // later prompts back; if pi has not started it within abortGraceMs, the
+ // engine stops pi.
function failTurn(turn, code, message) {
if (turn.done) return;
turn.done = true;
if (turn.timer !== null) clearTimeoutImpl(turn.timer);
turn.timer = null;
+ if (state.pending.includes(turn)) {
+ turn.grace = setTimeoutImpl(() => {
+ turn.grace = null;
+ // A run pi started keeps its place until its agent_end or a settle.
+ if (state.busy) return;
+ // No agent_start yet. Pi may never run this prompt, or its events
+ // may still be on the way; with no prompt id in them, nothing sent
+ // now could be told apart from it. Stop pi: held prompts fail, and
+ // the exit fails the rest and reaches onExit.
+ log(`engine: no agent_start ${abortGraceMs} ms after a failed turn; stopping pi`);
+ wedge();
+ }, abortGraceMs);
+ }
turn.reject(new DiscordError(message, 1, { code }));
}
+ function wedge() {
+ if (state.wedged || state.exited !== null) return;
+ state.wedged = true;
+ for (const h of state.held.splice(0)) failTurn(h.turn, "engine-wedged", "engine stopped: pi did not start an aborted turn");
+ stopChild();
+ }
+
+ // Call when a turn leaves the pending queue.
+ function release(turn) {
+ if (turn.grace !== null) clearTimeoutImpl(turn.grace);
+ turn.grace = null;
+ }
+
function settleTurn(turn, value) {
if (turn.done) return;
turn.done = true;
@@ -99,7 +140,10 @@ export function createEngine({
function failAll(code, message) {
const pending = state.pending.splice(0);
- for (const t of pending) failTurn(t, code, message);
+ for (const t of pending) {
+ release(t);
+ failTurn(t, code, message);
+ }
for (const h of state.held.splice(0)) failTurn(h.turn, code, message);
for (const [, r] of state.responses) r.reject(new DiscordError(message, 1, { code }));
state.responses.clear();
@@ -174,6 +218,7 @@ export function createEngine({
// Attribute the run to the head even if it failed client-side, so the
// next prompt's agent_end is not taken for this one.
const run = state.pending.shift();
+ if (run) release(run);
if (!run || run.done) return;
const messages = Array.isArray(event.messages) ? event.messages.filter((m) => m && m.role === "assistant") : [];
const message = messages.length > 0 ? messages[messages.length - 1] : run.last;
@@ -193,13 +238,14 @@ export function createEngine({
// this settle and still has no agent_end will never get one: fail it now
// instead of waiting for its timeout. Turns whose prompt response has
// not arrived yet belong to a later run and stay.
+ const dropped = [];
const keep = [];
- for (const t of state.pending) {
- if (t.done) continue;
- if (t.accepted) failTurn(t, "engine-settled-without-turn", "engine settled without answering this prompt");
- else keep.push(t);
- }
+ for (const t of state.pending) (t.done || t.accepted ? dropped : keep).push(t);
state.pending = keep;
+ for (const t of dropped) {
+ release(t);
+ failTurn(t, "engine-settled-without-turn", "engine settled without answering this prompt");
+ }
sendHeld();
}
}
@@ -214,21 +260,29 @@ export function createEngine({
// Never accepted: pi will not emit a turn_end for it, so remove it.
const i = state.pending.indexOf(turn);
if (i !== -1) state.pending.splice(i, 1);
+ release(turn);
failTurn(turn, (err.details && err.details.code) || "engine-refused", err.message);
sendHeld();
});
}
+ // Pi is busy from our side while any sent prompt is still queued, even one
+ // that already failed here: a turn that timed out before its agent_start
+ // was read leaves state.busy false while pi runs it, and sending then would
+ // be refused as streaming. It leaves the queue on its agent_end, on a
+ // settle, on a refused send, or at process exit.
+ const engineBusy = () => state.busy || state.pending.length > 0;
+
// After a settle (or a refused send) the oldest held prompt goes out.
function sendHeld() {
- if (state.exited !== null) return;
- if (state.busy || state.pending.some((t) => !t.done)) return;
+ if (state.exited !== null || state.wedged) return;
+ if (engineBusy()) return;
const next = state.held.shift();
if (next) send(next.turn, next.command);
}
function write(command) {
- if (!state.child || state.exited !== null) throw new DiscordError("engine is not running", 1, { code: "engine-down" });
+ if (!state.child || state.exited !== null || state.wedged) throw new DiscordError("engine is not running", 1, { code: "engine-down" });
state.child.stdin.write(JSON.stringify(command) + "\n");
}
@@ -245,6 +299,30 @@ export function createEngine({
});
}
+ function stopChild({ graceMs = 5000 } = {}) {
+ const child = state.child;
+ if (!child || state.exited !== null) return Promise.resolve(state.exited);
+ return new Promise((resolve) => {
+ const timer = setTimeoutImpl(() => {
+ try {
+ child.kill("SIGKILL");
+ } catch {
+ // already gone
+ }
+ }, graceMs);
+ child.once("exit", () => {
+ clearTimeoutImpl(timer);
+ resolve(state.exited);
+ });
+ try {
+ child.stdin.end();
+ child.kill("SIGTERM");
+ } catch {
+ // already gone
+ }
+ });
+ }
+
return {
start() {
if (state.child) throw new DiscordError("engine already started", 1);
@@ -281,7 +359,7 @@ export function createEngine({
// with DiscordError carrying details.code for the turn record.
prompt(text, { timeoutMs = 180000 } = {}) {
if (typeof text !== "string" || text.length === 0) throw new DiscordError("prompt text required", 1);
- const turn = { resolve: null, reject: null, timer: null, done: false, accepted: false, tools: new Map(), turns: 0, last: null };
+ const turn = { resolve: null, reject: null, timer: null, grace: null, done: false, accepted: false, tools: new Map(), turns: 0, last: null };
const done = new Promise((resolve, reject) => {
turn.resolve = resolve;
turn.reject = reject;
@@ -306,44 +384,24 @@ export function createEngine({
}
failTurn(turn, "timeout", `turn timed out after ${timeoutMs} ms`);
}, timeoutMs);
- if (state.exited !== null) {
+ if (state.exited !== null || state.wedged) {
failTurn(turn, "engine-down", "engine is not running");
return done;
}
- if (state.busy || state.pending.some((t) => !t.done) || state.held.length > 0) state.held.push({ turn, command });
+ if (engineBusy() || state.held.length > 0) state.held.push({ turn, command });
else send(turn, command);
return done;
},
get busy() {
- return state.busy || state.pending.some((t) => !t.done) || state.held.length > 0;
+ return engineBusy() || state.held.length > 0;
},
get pendingCount() {
return state.pending.filter((t) => !t.done).length + state.held.length;
},
- stop({ graceMs = 5000 } = {}) {
- const child = state.child;
- if (!child || state.exited !== null) return Promise.resolve(state.exited);
- return new Promise((resolve) => {
- const timer = setTimeoutImpl(() => {
- try {
- child.kill("SIGKILL");
- } catch {
- // already gone
- }
- }, graceMs);
- child.once("exit", () => {
- clearTimeoutImpl(timer);
- resolve(state.exited);
- });
- try {
- child.stdin.end();
- child.kill("SIGTERM");
- } catch {
- // already gone
- }
- });
+ stop(options) {
+ return stopChild(options);
},
};
}
diff --git a/packages/discord/tests/engine.test.mjs b/packages/discord/tests/engine.test.mjs
index 59674d1e..f3bd7525 100644
--- a/packages/discord/tests/engine.test.mjs
+++ b/packages/discord/tests/engine.test.mjs
@@ -4,6 +4,8 @@ import { readFileSync } from "node:fs";
import { join } from "node:path";
import { createEngine, buildPiArgs, PI_FIXED_ARGS, TOOLS_EXTENSION, READONLY_TOOLS_EXTENSION, assistantText } from "../src/engine-pi.mjs";
import { existsSync } from "node:fs";
+import { EventEmitter } from "node:events";
+import { PassThrough } from "node:stream";
import { makeRoot } from "./helpers.mjs";
const fakePi = join(import.meta.dirname, "fake-pi.mjs");
@@ -28,7 +30,20 @@ function start(root, extra = {}) {
log: (m) => logs.push(m), ...extra,
});
engine.start();
- return { engine, logs, commands: () => readFileSync(logPath, "utf8").trim().split("\n").filter(Boolean).map((l) => JSON.parse(l)) };
+ // The fake creates its log on the first command; until then there are none.
+ const commands = () => (existsSync(logPath) ? readFileSync(logPath, "utf8").trim().split("\n").filter(Boolean).map((l) => JSON.parse(l)) : []);
+ return { engine, logs, commands };
+}
+
+// Every test stops its engine in finally: a fake pi left running after a
+// failed assertion keeps the test file from exiting.
+async function withEngine(extra, body) {
+ const started = start(makeRoot(), extra);
+ try {
+ await body(started);
+ } finally {
+ await started.engine.stop();
+ }
}
test("engine: buildPiArgs carries the fixed flags, engine settings, session dir and prompt file", () => {
@@ -56,8 +71,7 @@ test("engine: with tools, buildPiArgs turns pi's own tools off, loads the extens
assert.equal(rw[rw.indexOf("--tools") + 1], "list_dir,read_file,search,write_file,edit_file", "a writable root adds exactly the two write tools");
});
-test("engine: a run with tool turns settles once, on the answer, with every tool call in the result", async () => {
- const { engine } = start(makeRoot());
+test("engine: a run with tool turns settles once, on the answer, with every tool call in the result", () => withEngine({}, async ({ engine }) => {
const r = await engine.prompt("tools 3");
assert.equal(r.text, "read 3 file(s)");
assert.equal(r.turns, 2);
@@ -71,45 +85,38 @@ test("engine: a run with tool turns settles once, on the answer, with every tool
assert.equal(plain.turns, 1);
await idle(engine);
assert.equal(engine.busy, false);
- await engine.stop();
-});
+}));
-test("engine: a run that ends on a tool-only turn fails the prompt as empty; a retried run settles on the real end", async () => {
- const { engine } = start(makeRoot());
+test("engine: a run that ends on a tool-only turn fails the prompt as empty; a retried run settles on the real end", () => withEngine({}, async ({ engine }) => {
const r = await engine.prompt("toolonly");
assert.equal(r.text, "", "no text: the connector turns this into engine-empty");
assert.equal(r.tools.length, 1);
const again = await engine.prompt("retry");
assert.equal(again.text, "after retry");
- await engine.stop();
-});
+}));
-test("engine: one prompt, one turn, text and usage come back", async () => {
- const { engine } = start(makeRoot());
- try {
- const r = await engine.prompt("hello");
- assert.equal(r.text, "echo: hello");
- assert.deepEqual(r.usage, { input: 3, output: 2 });
- await idle(engine);
- assert.equal(engine.busy, false);
- } finally {
- await engine.stop();
- }
-});
+test("engine: one prompt, one turn, text and usage come back", () => withEngine({}, async ({ engine }) => {
+ const r = await engine.prompt("hello");
+ assert.equal(r.text, "echo: hello");
+ assert.deepEqual(r.usage, { input: 3, output: 2 });
+ await idle(engine);
+ assert.equal(engine.busy, false);
+}));
-test("engine: a prompt while streaming is held until pi settles, then sent as its own run, and answered in order", async () => {
- const { engine, commands } = start(makeRoot());
- const first = engine.prompt("slow 150");
- await new Promise((r) => setTimeout(r, 20));
+test("engine: a prompt while streaming is held until pi settles, then sent as its own run, and answered in order", () => withEngine({}, async ({ engine, commands }) => {
+ const first = engine.prompt("slow 300");
assert.equal(engine.busy, true);
const second = engine.prompt("second");
assert.equal(engine.pendingCount, 2);
- await new Promise((r) => setTimeout(r, 20));
- assert.equal(commands().filter((c) => c.type === "prompt").length, 1, "the second prompt is not sent while pi is busy");
+ const prompted = () => commands().filter((c) => c.type === "prompt");
+ assert.ok(await until(() => prompted().length > 0), "the first prompt reached pi");
+ assert.equal(prompted().length, 1, "the second prompt is not sent while pi is busy");
+ // The fake refuses a prompt without streamingBehavior while it runs one, so
+ // an answered second prompt also proves it was not sent early.
const [r1, r2] = await Promise.all([first, second]);
assert.equal(r1.text, "slow reply");
assert.equal(r2.text, "echo: second");
- const prompts = commands().filter((c) => c.type === "prompt");
+ const prompts = prompted();
assert.equal(prompts.length, 2);
// Never a pi follow-up: pi would fold it into the first run and close both
// answers with one agent_end (the live loss of 2026-09-17).
@@ -117,11 +124,9 @@ test("engine: a prompt while streaming is held until pi settles, then sent as it
assert.equal(prompts[1].streamingBehavior, undefined);
await idle(engine);
assert.equal(engine.busy, false);
- await engine.stop();
-});
+}));
-test("engine: a held prompt that times out before pi settles fails on its own and is never sent", async () => {
- const { engine, commands } = start(makeRoot());
+test("engine: a held prompt that times out before pi settles fails on its own and is never sent", () => withEngine({}, async ({ engine, commands }) => {
const first = engine.prompt("slow 200");
await new Promise((r) => setTimeout(r, 20));
await assert.rejects(engine.prompt("late one", { timeoutMs: 50 }), (e) => e.details.code === "timeout" && /waiting for the engine/.test(e.message));
@@ -130,50 +135,177 @@ test("engine: a held prompt that times out before pi settles fails on its own an
await idle(engine);
assert.deepEqual(commands().filter((c) => c.type === "prompt").map((c) => c.message), ["slow 200"]);
assert.deepEqual(commands().filter((c) => c.type === "abort"), [], "a held turn is not aborted; pi never had it");
- await engine.stop();
-});
+}));
-test("engine: timeout sends abort and fails only that turn; the process stays", async () => {
- const { engine, commands, logs } = start(makeRoot());
+test("engine: timeout sends abort and fails only that turn; the process stays", () => withEngine({}, async ({ engine, commands, logs }) => {
await assert.rejects(engine.prompt("slow 5000", { timeoutMs: 100 }), (err) => err.details.code === "timeout");
assert.ok(await until(() => commands().some((c) => c.type === "abort")), "abort reached pi");
assert.ok(logs.some((l) => /timed out/.test(l)));
const r = await engine.prompt("again");
assert.equal(r.text, "echo: again");
- await engine.stop();
+}));
+
+test("engine: tool events from a run that outlived its timeout never land in the next prompt's record", () => withEngine({}, async ({ engine }) => {
+ await assert.rejects(engine.prompt("late 200", { timeoutMs: 40 }), (err) => err.details.code === "timeout");
+ const r = await engine.prompt("after late");
+ assert.equal(r.text, "echo: after late");
+ assert.deepEqual(r.tools, [], "the dead run's read is not this prompt's evidence");
+ assert.equal(r.turns, 1, "the dead run's turns are not counted here");
+}));
+
+// The turn timer is fired by hand, before the engine has read any event from
+// pi, so the timed-out run is still pi's and state.busy is still false when
+// the next prompt arrives. Under load a real timer does the same.
+const TURN_MS = 60000;
+const manualTurnTimer = (fire) => ({
+ setTimeoutImpl: (fn, ms) => (ms === TURN_MS ? fire.push(fn) : setTimeout(fn, ms)),
+ clearTimeoutImpl: (id) => { if (typeof id !== "number") clearTimeout(id); },
});
-test("engine: tool events from a run that outlived its timeout never land in the next prompt's record", async () => {
- const { engine } = start(makeRoot());
- try {
- await assert.rejects(engine.prompt("late 200", { timeoutMs: 40 }), (err) => err.details.code === "timeout");
- const r = await engine.prompt("after late");
+test("engine: a prompt after a turn that timed out before its agent_start waits for pi to settle instead of being refused", () => {
+ const fire = [];
+ return withEngine(manualTurnTimer(fire), async ({ engine, commands }) => {
+ const late = engine.prompt("late 100", { timeoutMs: TURN_MS });
+ fire.shift()();
+ assert.equal(engine.busy, true, "pi is still running the prompt that timed out");
+ const next = engine.prompt("after late", { timeoutMs: 5000 });
+ assert.equal(engine.pendingCount, 1, "only the new prompt is live");
+ await assert.rejects(late, (err) => err.details.code === "timeout");
+ const r = await next;
assert.equal(r.text, "echo: after late");
assert.deepEqual(r.tools, [], "the dead run's read is not this prompt's evidence");
- assert.equal(r.turns, 1, "the dead run's turns are not counted here");
- } finally {
- await engine.stop();
- }
+ assert.equal(r.turns, 1);
+ assert.deepEqual(commands().map((c) => (c.type === "prompt" ? c.message : c.type)), ["late 100", "abort", "after late"]);
+ await idle(engine);
+ assert.equal(engine.busy, false);
+ });
+});
+
+// "mute" is accepted and never run, so no agent_start, agent_end or settle
+// ever comes for it. Unbounded, it would hold every later prompt.
+test("engine: when pi has not started a timed-out turn by the end of the abort grace, the engine stops pi and fails held prompts", async () => {
+ let exited = null;
+ await withEngine({ abortGraceMs: 150, onExit: (e) => (exited = e) }, async ({ engine, commands, logs }) => {
+ await assert.rejects(engine.prompt("mute", { timeoutMs: 50 }), (err) => err.details.code === "timeout");
+ assert.equal(engine.busy, true, "pi might still be running it");
+ const started = Date.now();
+ await assert.rejects(engine.prompt("after mute", { timeoutMs: 5000 }), (err) => err.details.code === "engine-wedged");
+ assert.ok(Date.now() - started >= 100, "held for the grace, not failed at once");
+ await assert.rejects(engine.prompt("later"), (err) => err.details.code === "engine-down");
+ assert.ok(await until(() => exited !== null), "pi exits and onExit hears of it");
+ assert.ok(logs.some((l) => /stopping pi/.test(l)));
+ assert.deepEqual(commands().map((c) => (c.type === "prompt" ? c.message : c.type)), ["mute", "abort"]);
+ });
+});
+
+// Rocko's 6b R1 case: pi is stuck before agent_start, then runs the old
+// prompt and only afterwards reads the next one. The events carry no prompt
+// id, so a prompt sent after the grace would get the old run's answer.
+test("engine: a timed-out turn pi starts only after the grace never answers a later prompt", async () => {
+ let exited = null;
+ await withEngine({ abortGraceMs: 150, onExit: (e) => (exited = e) }, async ({ engine, commands }) => {
+ await assert.rejects(engine.prompt("stall 400", { timeoutMs: 50 }), (err) => err.details.code === "timeout");
+ await assert.rejects(engine.prompt("after stall", { timeoutMs: 5000 }), (err) => err.details.code === "engine-wedged");
+ assert.ok(await until(() => exited !== null), "pi exits and onExit hears of it");
+ await new Promise((res) => setTimeout(res, 400));
+ assert.ok(!commands().some((c) => c.message === "after stall"), "nothing was sent after the grace");
+ });
});
-test("engine: a malformed JSONL line fails the turn, not the process", async () => {
- const { engine, logs } = start(makeRoot());
+// The same case in memory, after Rocko's reproducer: the old run's events
+// arrive after the grace while pi is still exiting. They land on the failed
+// turn, nothing more is written to pi, and only the exit ends the engine.
+// Pi's response to the old prompt comes either before its timeout or only
+// with the late events.
+for (const lateResponse of [false, true]) test(`engine: late events of a run past its grace, before pi exits, answer nothing and nothing more is sent (${lateResponse ? "late" : "early"} prompt response)`, async () => {
+ const timers = [];
+ const written = [];
+ const kills = [];
+ const child = new EventEmitter();
+ child.stdout = new PassThrough();
+ child.stderr = new PassThrough();
+ child.stdin = { write: (s) => { written.push(JSON.parse(s)); return true; }, end: () => {} };
+ child.kill = (signal) => { kills.push(signal); return true; };
+ let exited = null;
+ const engine = createEngine({
+ command: "memory-only", args: [], spawn: () => child, abortGraceMs: 150, onExit: (e) => (exited = e),
+ setTimeoutImpl: (fn, ms) => { const t = { fn, ms, active: true }; timers.push(t); return t; },
+ clearTimeoutImpl: (t) => { t.active = false; },
+ });
+ const emit = (x) => child.stdout.write(JSON.stringify(x) + "\n");
+ const fire = (ms) => { const t = timers.find((x) => x.ms === ms && x.active); assert.ok(t, `timer ${ms}`); t.active = false; t.fn(); };
+ const message = (text) => ({ role: "assistant", content: [{ type: "text", text }], stopReason: "stop" });
+ const tick = () => new Promise((res) => setImmediate(res));
+ // Checked after a tick instead of awaited, so a regression fails here
+ // rather than hanging on a promise nothing will settle.
+ const outcome = (p) => {
+ const o = { state: "pending", code: null, text: null };
+ p.then((v) => Object.assign(o, { state: "resolved", text: v.text }), (e) => Object.assign(o, { state: "rejected", code: e.details && e.details.code }));
+ return o;
+ };
+ engine.start();
+ const first = engine.prompt("old", { timeoutMs: 50 });
+ const accept = () => emit({ type: "response", id: written[0].id, command: "prompt", success: true });
+ if (!lateResponse) accept();
+ await tick();
+ fire(50);
+ await assert.rejects(first, (err) => err.details.code === "timeout");
+ const next = outcome(engine.prompt("new", { timeoutMs: 2000 }));
+ fire(150);
+ await tick();
+ assert.deepEqual(next, { state: "rejected", code: "engine-wedged", text: null });
+ assert.deepEqual(kills, ["SIGTERM"]);
+ if (lateResponse) accept();
+ emit({ type: "agent_start" });
+ emit({ type: "tool_execution_start", toolCallId: "old-call", toolName: "read_file", args: { root: "docs", path: "old.md" } });
+ emit({ type: "tool_execution_end", toolCallId: "old-call", toolName: "read_file", result: { details: { root: "docs", path: "old.md", ok: true } } });
+ emit({ type: "turn_end", message: message("OLD RUN ANSWER") });
+ emit({ type: "agent_end", messages: [message("OLD RUN ANSWER")] });
+ emit({ type: "agent_settled" });
+ await tick();
+ const after = outcome(engine.prompt("after settle", { timeoutMs: 2000 }));
+ await tick();
+ assert.deepEqual(after, { state: "rejected", code: "engine-down", text: null });
+ assert.deepEqual(next, { state: "rejected", code: "engine-wedged", text: null }, "the old answer did not reach the new prompt");
+ assert.deepEqual(written.map((c) => (c.type === "prompt" ? c.message : c.type)), ["old", "abort"], "no prompt reached pi after the grace");
+ assert.equal(exited, null);
+ fire(5000);
+ assert.deepEqual(kills, ["SIGTERM", "SIGKILL"]);
+ child.emit("exit", null, "SIGKILL");
+ assert.deepEqual(exited, { code: null, signal: "SIGKILL" });
+});
+
+test("engine: a timed-out run pi did start outlives the grace; the next prompt goes out when it ends", async () => {
+ let exited = null;
+ await withEngine({ abortGraceMs: 150, onExit: (e) => (exited = e) }, async ({ engine, commands }) => {
+ await assert.rejects(engine.prompt("late 400", { timeoutMs: 50 }), (err) => err.details.code === "timeout");
+ const r = await engine.prompt("after late", { timeoutMs: 5000 });
+ assert.equal(r.text, "echo: after late");
+ assert.deepEqual(r.tools, []);
+ assert.equal(exited, null, "pi was not stopped");
+ assert.deepEqual(commands().map((c) => (c.type === "prompt" ? c.message : c.type)), ["late 400", "abort", "after late"]);
+ });
+});
+
+test("engine: a malformed JSONL line fails the turn, not the process", () => withEngine({}, async ({ engine, logs }) => {
await assert.rejects(engine.prompt("garbage"), (err) => err.details.code === "engine-protocol");
assert.ok(logs.some((l) => /malformed/.test(l)));
const r = await engine.prompt("still here");
assert.equal(r.text, "echo: still here");
- await engine.stop();
-});
+}));
test("engine: a turn that ends in error rejects with the error code; process exit fails pending turns", async () => {
- const root = makeRoot();
let exited = null;
- const { engine } = start(root, { onExit: (e) => (exited = e) });
- await assert.rejects(engine.prompt("error"), (err) => err.details.code === "engine-error" && /fake provider error/.test(err.message));
- const pending = engine.prompt("slow 5000");
- await new Promise((r) => setTimeout(r, 20));
- await engine.stop();
- await assert.rejects(pending, (err) => err.details.code === "engine-down");
- assert.ok(exited);
- await assert.rejects(engine.prompt("x"), /not running/);
+ const { engine } = start(makeRoot(), { onExit: (e) => (exited = e) });
+ try {
+ await assert.rejects(engine.prompt("error"), (err) => err.details.code === "engine-error" && /fake provider error/.test(err.message));
+ const pending = engine.prompt("slow 5000");
+ await new Promise((r) => setTimeout(r, 20));
+ await engine.stop();
+ await assert.rejects(pending, (err) => err.details.code === "engine-down");
+ assert.ok(exited);
+ await assert.rejects(engine.prompt("x"), /not running/);
+ } finally {
+ await engine.stop();
+ }
});
diff --git a/packages/discord/tests/fake-pi.mjs b/packages/discord/tests/fake-pi.mjs
index 94919306..bf94c013 100644
--- a/packages/discord/tests/fake-pi.mjs
+++ b/packages/discord/tests/fake-pi.mjs
@@ -8,9 +8,13 @@
// then a second turn that answers "read <n> file(s)"
// "toolonly" a run whose only turn calls a tool and never answers
// "retry" an agent_end with willRetry, then the real answer
+// "mute" accept the prompt and emit nothing, staying idle
// "late <ms>" ignore abort; after <ms> emit a tool pair and a tool turn,
// then answer "late reply", like a run that outlives its
// client-side timeout
+// "stall <ms>" accept the prompt, then read nothing for <ms> (pi stuck
+// before agent_start); then run it, answering "echo:
+// stalled", and only then read what came in meanwhile
// anything else answer "echo: <text>" immediately
// A prompt received while busy without streamingBehavior is refused, as pi
// does. A prompt with streamingBehavior followUp is folded into the running
@@ -30,6 +34,7 @@ function assistant(text, stopReason = "stop") {
}
function run(text) {
+ if (text === "mute") return;
busy = true;
out({ type: "agent_start" });
out({ type: "turn_start" });
@@ -109,11 +114,16 @@ function run(text) {
let current = null;
let buffer = "";
+let stalled = false;
process.stdin.setEncoding("utf8");
process.stdin.on("data", (chunk) => {
buffer += chunk;
+ drain();
+});
+
+function drain() {
let idx;
- while ((idx = buffer.indexOf("\n")) !== -1) {
+ while (!stalled && (idx = buffer.indexOf("\n")) !== -1) {
const line = buffer.slice(0, idx);
buffer = buffer.slice(idx + 1);
if (!line) continue;
@@ -125,7 +135,15 @@ process.stdin.on("data", (chunk) => {
continue;
}
out({ id: cmd.id, type: "response", command: "prompt", success: true });
- if (busy) queue.push(cmd.message);
+ const sm = /^stall (\d+)$/.exec(cmd.message);
+ if (sm) {
+ stalled = true;
+ setTimeout(() => {
+ run("stalled");
+ stalled = false;
+ drain();
+ }, Number(sm[1]));
+ } else if (busy) queue.push(cmd.message);
else run(cmd.message);
} else if (cmd.type === "abort") {
out({ id: cmd.id, type: "response", command: "abort", success: true });
@@ -139,5 +157,5 @@ process.stdin.on("data", (chunk) => {
out({ id: cmd.id, type: "response", command: cmd.type, success: true, data: {} });
}
}
-});
+}
process.stdin.on("end", () => process.exit(0));
@@ -0,0 +1,25 @@
{
"observedAt": "2026-09-13T19:56:18.961314+00:00",
"issue": 1510,
"ownerAuthorizedLiveSmoke": true,
"researcher": [
{
"session": ".pi/state/researcher/sessions/2026-09-13T19-52-28-745Z_01a09c54-0b48-7154-addd-8fdce875aa4a.jsonl",
"entryId": "64b15fd5",
"timestamp": "2026-09-13T19:52:43.079Z",
"response": "RESEARCHER_NATIVE_SMOKE_OK",
"entrySha256": "d9cbaa5ea640d2b858a86b2fa24d3240b351d9383ffaee9721dd86fcd080c329"
}
],
"rocko": {
"newLaunch": "refused by existing native launch lock",
"existingPid": 3707667,
"cwd": "/mnt/storage/src/mosaic-stack",
"nativeContextVerified": true,
"sonnetFlagVerified": true,
"socket": "mosaic-fleet",
"newModelResponseTested": false
},
"existingProcessesRestarted": false,
"homeLaunchersModified": false
}
@@ -0,0 +1,19 @@
{
"issue": 1510,
"candidate": "/tmp/internal-team-r1-i9t21yzr",
"manifestSha256": "23a27014ce6f04ce8187d8495b2efe62b814c3c7d041b027f99b1ec4d490709d",
"files": [
"AGENTS.md",
"agents/README.md",
"agents/researcher/SOUL.md",
"agents/researcher/CONTEXT.md",
"agents/researcher/README.md",
"agents/researcher/launch.sh",
"agents/researcher/validate-sessions.mjs",
"scripts/test-darkwing-launch.mjs",
"docs/plans/2026-09-13_internal-development-bootstrap.md"
],
"tests": "node --test scripts/test-darkwing-launch.mjs scripts/test-rocko-launch.mjs",
"passed": 6,
"state": "ready for independent review"
}
@@ -0,0 +1,77 @@
# Ledger: T3 header counts as agent (#1506), R1 candidate
Sage assigned this on 2026-09-26 after commit A (af4203ca). Filbert reviews.
Not committed.
## Defect
`messageKind` in `packages/ledger/src/ledger.mjs` knew only the tmux preamble
`[host:session -> host:session]`. A prompt that opens with the T3 header
`[from: sage (1ef1e4f8-…) -> to: filbert (9cb9731e-…) class=actionable]`
counted as human, so Table 2's Human column and the human-per-closed ratio
rise once seats talk over T3. DEFERRED Open entry "Ledger counts T3 agent
messages as human".
## Change
- `messageKind` also matches the T3 header on the first line. The sender is the
`from:` role. `control-board` is board, any other role is agent. Anything short
of the full header stays human. That includes the header on a later line, a
leading space, a missing thread id, `class=` with capitals, `]` followed by a
non-space, and `From:` capitalized. The tmux branch is unchanged.
- Two tests: a Table 2 fixture with two headered prompts and one plain prompt
for seat `bob`, expecting agent 2 and human 1. Also a direct classification table.
- README counting rule names both forms.
## Evidence
- `node --test --test-reporter=tap packages/ledger/tests/`: 22/22 on the
working tree.
- The same test file against HEAD's `ledger.mjs` (full `git archive HEAD` tree):
20/22. The two failures are the two new tests, so they catch the defect.
An earlier archive of `packages/ledger` alone also failed three gitea-helper
tests. Those tests need `scripts/gitea-api.sh`, which the partial archive left out.
- Suites on the working tree: config 24, task 90, foundation 43, conductor 17,
release 14, auth 15, discord 63. None of them runs the ledger tests.
## Frozen files
`r1-manifest.sha256` holds the three file hashes; `r1.patch` is `git diff
packages/ledger` at freeze time.
## Known limit, not fixed here
The fix changes zero current counts. Table 2 reads only
`.pi/state/<seat>/sessions/*.jsonl`, and no file there contains a T3 header
for any seat. Filbert checked this in R1:
- `grep -rF '[from: ' .pi/state/*/sessions/` finds nothing.
- The candidate `messageKind` gives Dewey 52 agent, 7 board and 9 human, and
Sage 193 agent and 72 human. That covers every user message in Pi logs
modified since 2026-09-20. Every non-human first line is the tmux form.
- Filbert's three T3 messages to Dewey on 2026-09-26 are not in
`.pi/state/dewey`.
Dewey's and Sage's Pi logs are current, but they only carry tmux traffic. T3
traffic goes to the harness transcripts: Claude under `~/.claude/projects`,
Codex under `~/.codex/sessions`. The ledger reads neither, so a T3-routed
prompt to any seat counts nowhere, as agent or as human. The Human column
can't see T3 traffic at all. For Darkwing and Filbert, whose newest Pi logs
end 2026-09-14, and for Rocko, who has no Pi sessions directory, zero means an
empty source, not zero human prompts. The fix is correct for a source that
carries T3 headers. Sage asked for a brief on a read-only T3 thread source;
Gate F waits on it.
Filbert also found two misclassifications in older logs. Neither is touched
here:
- `[rev-code-02 -> dragon-lin:sage class=actionable]` has no host on the
sender, so it counts as human.
- One Dewey prompt opens with a quote character before the tmux preamble, so
it counts as human.
## Review
Filbert, R1, 2026-09-26: approved the three frozen files. He verified the
manifest and patch, got 22/22 on the tree and 20/22 against HEAD's source, and
matched the regex to `docs/guides/T3-AGENT-COMMS.md`. He accepts the body on
the header's line, which the tmux branch also allows. He corrected the
known-limit text above.
@@ -0,0 +1,3 @@
e0d411ca2f45d85734eef130dba645646df6e9128ea7dc2eaba2205df7891bb8 packages/ledger/README.md
e24b065c4284370960ac6ff1ed66810fe601da64ae9b9362584fe9fbee334017 packages/ledger/src/ledger.mjs
a9da013e81aff360cb013a8e103fdd111aee7e42553b96cb0811560b3da39250 packages/ledger/tests/ledger.test.mjs
@@ -0,0 +1,89 @@
diff --git a/packages/ledger/README.md b/packages/ledger/README.md
index a998a76c..da0ab1c5 100644
--- a/packages/ledger/README.md
+++ b/packages/ledger/README.md
@@ -36,11 +36,15 @@ No install, build, service restart, or configuration change is needed.
duplicated entries in copied logs are not deduplicated. No transcript content
leaves the parser. Assistant messages and logs outside repo seats do not count.
Symlink source directories are refused and symlink files are not followed.
-- The first text line alone classifies a message. A bracketed addressing
- preamble whose source session is `control-board` is board; any other valid
- addressing preamble is agent; otherwise human. This is a format count, not
- proof of who typed the message. Text blocks are joined with newlines.
- The entry timestamp is used, falling back to the message timestamp.
+- The first text line alone classifies a message. Two addressing forms count:
+ the tmux preamble `[host:session -> host:session]` that `agent-send.sh`
+ writes, and the T3 header `[from: role (thread-id) -> to: role (thread-id)]`
+ from `docs/guides/T3-AGENT-COMMS.md`. Either may carry ` class=<class>` before
+ the closing bracket. A preamble whose sender is `control-board` (tmux session
+ or T3 role) is board; any other valid preamble is agent; otherwise human.
+ This is a format count, not proof of who typed the message. Text blocks are
+ joined with newlines. The entry timestamp is used, falling back to the
+ message timestamp.
- Seats with no in-range user messages are omitted. Issue seats come from `#N`
mentions anywhere in in-range user text, including quoted text.
- Human messages per closed issue divides Table 2's human sum by issues closed
diff --git a/packages/ledger/src/ledger.mjs b/packages/ledger/src/ledger.mjs
index dc8a3a69..b09e95f5 100644
--- a/packages/ledger/src/ledger.mjs
+++ b/packages/ledger/src/ledger.mjs
@@ -77,8 +77,12 @@ export function messageText(content) {
}
export function messageKind(text) {
const firstLine = text.split(/\r?\n/, 1)[0];
- const match = firstLine.match(/^\[([^\s:\[\]]+):([^\s\[\]]+) -> ([^\s:\[\]]+):([^\s\[\]]+)(?: class=[a-z-]+)?\](?:\s|$)/);
- return !match ? 'human' : match[2] === 'control-board' ? 'board' : 'agent';
+ // tmux preamble from agent-send.sh: [host:session -> host:session class=x]
+ const tmux = firstLine.match(/^\[([^\s:\[\]]+):([^\s\[\]]+) -> ([^\s:\[\]]+):([^\s\[\]]+)(?: class=[a-z-]+)?\](?:\s|$)/);
+ // T3 header (docs/guides/T3-AGENT-COMMS.md): [from: role (id) -> to: role (id) class=x]
+ const t3 = firstLine.match(/^\[from: ([^\s()\[\]]+) \(([^()\[\]]+)\) -> to: ([^\s()\[\]]+) \(([^()\[\]]+)\)(?: class=[a-z-]+)?\](?:\s|$)/);
+ const sender = tmux ? tmux[2] : t3 ? t3[1] : null;
+ return sender === null ? 'human' : sender === 'control-board' ? 'board' : 'agent';
}
async function directories(dir, optional = false) {
try {
diff --git a/packages/ledger/tests/ledger.test.mjs b/packages/ledger/tests/ledger.test.mjs
index 7b4e4d33..175f5d5e 100644
--- a/packages/ledger/tests/ledger.test.mjs
+++ b/packages/ledger/tests/ledger.test.mjs
@@ -110,6 +110,19 @@ test('invalid dates, reverse dates and duplicate options refuse', t => {
assert.throws(() => dateRange('2026-02-30')); assert.throws(() => dateRange('2026-09-12', '2026-09-06'));
const f = fixture(t); assert.equal(f.run(['--since', '2026-09-01']).status, 1);
});
+test('T3 agent assignments do not count as human in Table 2', t => {
+ const f = fixture(t);
+ f.put('.pi/state/bob/sessions/t3.jsonl', [
+ f.entry('[from: sage (1ef1e4f8) -> to: bob (9cb9731e) class=actionable]\nassign #1'),
+ f.entry('[from: sage (1ef1e4f8) -> to: bob (9cb9731e)]\nfollow-up #1'),
+ f.entry('Jason: go ahead'),
+ ].map(x => JSON.stringify(x)).join('\n') + '\n');
+ const result = f.run(['--json']);
+ assert.equal(result.status, 0, result.stderr);
+ const r = JSON.parse(result.stdout);
+ assert.deepEqual(r.seats, [{ seat: 'alice', board: 1, agent: 1, human: 1 }, { seat: 'bob', board: 0, agent: 2, human: 1 }]);
+ assert.equal(r.totals.humanMessagesPerClosedIssue, 2);
+});
test('preamble parsing and issue number boundaries', () => {
assert.equal(messageKind('[h:control-board -> h:seat] hi'), 'board');
assert.equal(messageKind('[h:seat -> h:seat class=actionable] hi'), 'agent');
@@ -117,6 +130,20 @@ test('preamble parsing and issue number boundaries', () => {
assert.equal(messageKind(' [h:seat -> h:seat] quoted'), 'human');
assert.deepEqual(issueNumbers('fix #1 #2 #2 abc#3 #0 #4x'), [1, 2]);
});
+test('T3 header: agent, or board from control-board; anything short of the full header is human', () => {
+ const sage = 'sage (1ef1e4f8-3ead-4208-beca-38f9f1add079)', filbert = 'filbert (9cb9731e-a10f-4c8f-a212-c4fa1f5f4731)';
+ assert.equal(messageKind(`[from: ${sage} -> to: ${filbert}]\nbuild #1506`), 'agent');
+ assert.equal(messageKind(`[from: ${sage} -> to: ${filbert} class=actionable]\nbuild`), 'agent');
+ assert.equal(messageKind(`[from: ${sage} -> to: ${filbert}] same line`), 'agent');
+ assert.equal(messageKind(`[from: darkwing (thread-id: unknown) -> to: reviewer (new-thread)]\nreview`), 'agent');
+ assert.equal(messageKind(`[from: control-board (b) -> to: ${filbert}]\nhi`), 'board');
+ assert.equal(messageKind(`Jason here\n[from: ${sage} -> to: ${filbert}]\nquoted`), 'human');
+ assert.equal(messageKind(` [from: ${sage} -> to: ${filbert}]`), 'human');
+ assert.equal(messageKind(`[from: sage -> to: filbert]\nno thread ids`), 'human');
+ assert.equal(messageKind(`[from: ${sage} -> to: ${filbert} class=Actionable]`), 'human');
+ assert.equal(messageKind(`[from: ${sage} -> to: ${filbert}]trailing`), 'human');
+ assert.equal(messageKind(`[From: ${sage} -> to: ${filbert}]`), 'human');
+});
test('no closed issues with human messages means undefined ratio, not invented zero', () => {
const r = summarize(range, [], [], { rows: [{ seat: 'a', human: 1, board: 0, agent: 0 }], mentions: new Map() });
assert.equal(r.totals.humanMessagesPerClosedIssue, 'unknown');
@@ -0,0 +1,5 @@
0afb0320e9a1833f133426169c389ffb71be3d490e0e402070162aa22e820296 packages/ledger/src/ledger.mjs
d6092a538a059e8869544df903e89a8b9a4c3b6b492645db762b2f03d5630a44 packages/ledger/src/cli.mjs
dfb092aabee5cf5197029c6e8978df570ee50f08d84babde931c6a763befafe9 packages/ledger/src/t3.mjs
7444abd1dbd8e637705def0fd98105e6e69970f1bd397a11f8277996521579d6 packages/ledger/tests/ledger.test.mjs
27f7366dd5edc30a93a8c54bfb46b3fed87e1a44f11e1b22bd159a0718625273 packages/ledger/README.md
@@ -0,0 +1,161 @@
# Gate F build: the ledger's T3 source (#1506), candidate for review
Darkwing built this on 2026-09-26 from the approved brief R3,
`docs/plans/2026-09-26_ledger-t3-source.md` (sha256 f3c05c1b, committed in
ffc22c04). Sage gave the go once Filbert confirmed R3. Filbert reviews the
code; Sage commits after the suites. Base is HEAD 1c5f6bc3. Nothing is
committed or pushed.
## Files
`build-manifest.sha256` pins the five files, and `build.patch` is the diff
against 1c5f6bc3 with `t3.mjs` included as a new file.
- `packages/ledger/src/t3.mjs` (new). `readT3(root, range, {dbPath, isDefault})`:
path checks, one read transaction, schema check, project, title mapping,
header cross-check, counts, mentions, diagnostic.
- `packages/ledger/src/ledger.mjs`. The class fix, `t3Header()`,
`readSeats()`, `mergeSources()`, the `pi` and `t3` keys in the report, the
text line for `--no-t3` or a non-default path, and the U+2028 fix below.
- `packages/ledger/src/cli.mjs`. `--no-t3` and `--t3-db PATH`, which refuse
each other; the usage line.
- `packages/ledger/tests/ledger.test.mjs`. HOME at both spawn sites, the
empty default database, 25 new tests.
- `packages/ledger/README.md`. A new "T3 source" section.
The commit should also carry Filbert's updated review,
`agents/filbert/work/ledger-t3-source-review-2026-09-26.md` (be1aa414), and
this directory's new files.
## Beyond the brief: the Pi reader split valid lines
The brief's live read has to exit 0. It didn't, and T3 wasn't the cause. HEAD
refuses the live checkout the same way:
`Malformed session JSON: filbert/2026-09-12T16-38-58-597Z_01a0967c-….jsonl:611`.
That line parses. It holds a raw U+2028 inside a JSON string, which JSON
allows and `JSON.stringify` writes unescaped. Node 26.8.1's `readline` ends a
line at U+2028 too, so it cut the record in two (733 lines by `readline`, 732
by `\n`). The reader parses every line before it checks the range, so on
Node 26.8.1 every live run refuses, whatever the dates. The file was last
written 2026-09-14. I haven't checked which Node version first split there.
The fix replaces `readline` with a small splitter that ends lines at `\n`
only. It sits in `ledger.mjs`, which this build already changes, and it
blocked acceptance, so I made it here instead of filing it. A new test writes a
Pi log with a raw U+2028 and CRLF endings; it fails with `readline` and passes
with the splitter. Please review it as its own item.
## Choices the brief left open
- Imported threads are excluded by the `import:` prefix alone. Live, all
1678 `historyImport` events sit in `import:` streams, so the two rules agree
today. The events table stays optional, so the exclusion doesn't depend on it.
- The header cross-check runs over every user message in a counted thread, in
range or not. The title mapping is current state, so a conflict in old
history still misassigns counts for any range that includes it.
- Validation (role, text, `created_at`) also covers every message in a
counted thread, assistant rows included, and not only rows in range.
- The diagnostic is in range: `humanSentThroughApi` and `humanWithoutEvent`.
One unparseable event, or an event with no string `messageId`, makes both
`unknown`, the same as a missing table (F5).
- Project and thread matching compare `workspace_root` in JavaScript, so a
declared collation on the column can't loosen byte-for-byte equality.
- JSON adds top-level `pi` (Pi rows) and `t3` (read flag, database, seats with
threads, unmapped, excluded, diagnostic). `seats` and `totals` keep their
shape, so existing consumers and tests are unchanged. `t3.seats` lists a
seat whenever it has a mapped thread, even with zero counts in range.
- The unmapped row comes last in `seats`, and only when it has counts.
## Evidence
- Ledger tests: `node --test packages/ledger/tests/`, 47/47 (ledger 44,
of which 25 are new, and gitea helper 3). The busy-timeout test takes about 5.4 s.
- Class fix against HEAD. HEAD's `messageKind` (from `git show
1c5f6bc3:packages/ledger/src/ledger.mjs`) calls a T3 header with
`class=REVIEW-REQUEST`, a tmux preamble with `class=DECISION` and a T3 header
with `class=Actionable` all human. The build calls them agent. The existing
test asserting `class=Actionable` is human now asserts agent.
- Mutations, each on a scratch copy of the package. Three `gitea-helper`
tests fail in every scratch copy because they need the repository's
`scripts/`, so the counts below leave them out.
- Classes back to `[a-z-]+`: 6 fail.
- No header cross-check: 2 fail.
- No symlink refusal: 4 fail.
- Busy timeout 0: 1 fails.
- Two projects allowed: 1 fails.
- Deleted threads kept, imported threads kept, or range filter removed:
4 fail each.
- No role check: 1 fails.
- No diagnostic table check: 1 fails.
- `readline` restored: 1 fails.
- Two mutations pass, and I'm naming them rather than hiding them:
- Removing `mode=ro` changes nothing, because `readOnly: true` already
opens read-only. Both stay, as the brief says.
- Removing `BEGIN` fails 19 tests, but only because `COMMIT` then has no
transaction. No test proves that the queries share one snapshot.
- Eight suites on a local clone of 1c5f6bc3 with the five files: config 24,
task 90, foundation 43, conductor 17, release 14, auth 15, discord 63,
extension-package 18. The first task run showed 89/1, and I didn't capture
the failing line. Three more task runs passed 90/90. I count it as a flake
I can't name, not as green on the first try.
- Union (control-board, webui, seat, mosaic, ledger, discord) on the same
clone: 434/434 three times, 23 to 24 s each. No fake pi left running.
- No test opens the real `~/.t3`. Every CLI spawn sets `HOME` to a temp
directory, and no test calls `readT3` in process. After the runs, no
`ledger-*` temp directories remained.
## Live read
`node packages/ledger/src/cli.mjs --since 2026-09-01 --until 2026-09-26
--no-issues`, exit 0 three times, no header conflict. The table is the run at
2026-09-26T21:31:02Z.
| Seat | T3 threads | T3 board / agent / human | Pi board / agent / human |
|---|---|---|---|
| darkwing | Darkwing; Darkwing in Claude (archived) | 0 / 19 / 28 | 14 / 109 / 141 |
| dewey | Dewey; Dewey in Claude | 0 / 17 / 7 | 7 / 57 / 16 |
| filbert | Filbert | 0 / 25 / 1 | 5 / 92 / 9 |
| rocko | Rocko | 0 / 20 / 1 | none |
| sage | Sage | 0 / 52 / 10 | 0 / 193 / 72 |
| researcher | none | none | 3 / 1 / 1 |
| t3:unmapped | Discord Bot | 0 / 0 / 68 | none |
This matches the brief, allowing for messages sent since 20:54Z. It maps the
same seven threads. Discord Bot has 68 human: 54 without a header and the 14
free-text headers. The diagnostic reads exactly those 14
(`humanSentThroughApi: 14`, `humanWithoutEvent: 0`). T3 agent messages total
133, against the brief's 96 API headers (80 plus the 16 uppercase ones) at
20:54Z. Two imported threads are excluded, and this project has no deleted
threads.
## Not covered
- Snapshot isolation across the queries (see the `BEGIN` mutation above).
- A seat directory named `t3:unmapped` would share the unmapped row. Directory
names that contain a colon aren't used in `agents/`.
- The live read's effect on the main database file can't be checked while T3
writes to it. The stopped and writer-attached WAL tests check it on
fixtures.
## Review and correction
Filbert approved manifest ba73a163 and the U+2028 fix as its own item:
`agents/filbert/work/ledger-t3-build-review-2026-09-26.md`, sha256 e47ec6da.
Correction to "Beyond the brief" above. Line 611 holds a raw U+2028 and a raw
U+2029, and `readline` ends a line at each. The file has 731 lines by `\n`
(`wc -l` agrees), and `readline` makes 733. I wrote 732 because I counted the
empty string after the final newline. The splitter already ends lines at `\n`
only, so the fix covers both characters. The test and the README name only
U+2028.
Filbert's nonblocking notes, for a follow-up after the Gate F commit, since
changing the pinned files now would void the approval:
1. Add a U+2029 to the splitter test and the README line.
2. Two diagnostic mutations survive: `humanWithoutEvent` hardcoded to 0, and
an unparseable event skipped instead of making the diagnostic `unknown`.
Each needs one fixture message.
3. `readT3`'s catch reports any error that isn't a `SourceError` as a SQLite
read failure. It still exits 1, but a bug would read as a database
problem. Rethrow errors that carry no `errcode`.
4. Snapshot isolation stays untested, as recorded above.
@@ -0,0 +1,828 @@
diff --git a/packages/ledger/README.md b/packages/ledger/README.md
index da0ab1c5..393e9c37 100644
--- a/packages/ledger/README.md
+++ b/packages/ledger/README.md
@@ -1,13 +1,16 @@
# Ledger
Read-only counts from local `refactor` commit subjects, one Gitea issue-list
-request through `scripts/gitea-api.sh`, and repo seats' Pi session logs.
+request through `scripts/gitea-api.sh`, repo seats' Pi session logs, and T3's
+thread messages in `~/.t3/userdata/state.sqlite`.
No board changes, data-root writes, fleet reads, transcript output, or scheduler.
```sh
node packages/ledger/src/cli.mjs --since 2026-09-06 --until 2026-09-12
node packages/ledger/src/cli.mjs --since 2026-09-06 --until 2026-09-12 --json
node packages/ledger/src/cli.mjs --since 2026-09-06 --no-issues
+node packages/ledger/src/cli.mjs --since 2026-09-06 --no-t3
+node packages/ledger/src/cli.mjs --since 2026-09-06 --t3-db /tmp/fixture.sqlite
node --test packages/ledger/tests/
```
@@ -36,12 +39,17 @@ No install, build, service restart, or configuration change is needed.
duplicated entries in copied logs are not deduplicated. No transcript content
leaves the parser. Assistant messages and logs outside repo seats do not count.
Symlink source directories are refused and symlink files are not followed.
+ A line ends at `\n` only. A U+2028 inside a JSON string does not split a record.
+- Table 2 also counts T3 thread messages with role `user`. The T3 source
+ follows. A seat's row sums its Pi and T3 counts; the JSON keeps the split in
+ `pi` (Pi rows) and `t3.seats` (T3 rows).
- The first text line alone classifies a message. Two addressing forms count:
the tmux preamble `[host:session -> host:session]` that `agent-send.sh`
writes, and the T3 header `[from: role (thread-id) -> to: role (thread-id)]`
from `docs/guides/T3-AGENT-COMMS.md`. Either may carry ` class=<class>` before
the closing bracket. A preamble whose sender is `control-board` (tmux session
or T3 role) is board; any other valid preamble is agent; otherwise human.
+ The class may be in either case: seats send `class=DECISION`.
This is a format count, not proof of who typed the message. Text blocks are
joined with newlines. The entry timestamp is used, falling back to the
message timestamp.
@@ -53,6 +61,69 @@ No install, build, service restart, or configuration change is needed.
ratios and durations with one decimal. Titles truncate to 48 characters in
text only. Missing evidence is the literal string `unknown`.
+## T3 source
+
+The rules come from `docs/plans/2026-09-26_ledger-t3-source.md` (Gate F).
+The source is on by default. `--no-t3` skips it, and the report then says
+`T3: not read (--no-t3)`. `--t3-db <path>` reads another database file with
+the same checks. The JSON records the database path and whether it was the
+default. When it wasn't, the text report prints the path, so a fixture result
+can't pass for a live one. The two flags can't be combined.
+
+The reader opens `state.sqlite` read-only through `node:sqlite` and reads no
+other file in `~/.t3`. It runs every query in one read transaction with a 5 s
+busy timeout. It never writes the main database file. Like any SQLite
+connection it may create `-wal` and `-shm` beside it, so a directory that
+isn't writable refuses when SQLite needs them.
+
+- **Project.** Only threads in the one non-deleted T3 project whose
+ `workspace_root` equals this checkout's root byte for byte. The root is the
+ realpath of the package, so a project opened through the compatibility
+ symlink `~/src/mosaic-stack-dev-test` does not match, and the report refuses
+ with no project.
+- **Thread to seat.** A thread belongs to seat `<s>` when `<s>` is a real
+ directory in `agents/` and the lower-cased title equals `<s>` or starts
+ with `<s>` and a space. "Dewey in Claude" maps to `dewey`; "Sagebrush" maps
+ to nothing. Several threads can map to one seat. Threads that map to no
+ seat share one row, `t3:unmapped`, so their human messages still reach the
+ totals. `t3.seats` and `t3.unmapped` in the JSON list the thread ids and
+ titles behind each row.
+- **Titles are current state.** T3 titles an unnamed thread from its first
+ prompt, and a rename moves a thread's whole history to another row. This
+ moves counts between rows, never out of the totals.
+- **Header check.** A user message whose T3 header is addressed to its own
+ thread id must name that thread's seat as the `to:` role (compared lower
+ case). In an unmapped thread the `to:` role must not be a seat. A conflict
+ exits 1 and names the thread, its title and both roles. A header addressed
+ to another thread isn't checked. The check misses a renamed thread that no
+ agent writes to. Such a thread can only add human counts to a row.
+- **Excluded.** Imported threads (id prefix `import:`) are partial copies of
+ Claude Code sessions, not T3 traffic; every T3 event marked `historyImport`
+ sits in one today. Deleted threads don't count; archived threads do.
+ `t3.excluded` gives both thread counts.
+- **Blind spot.** Threads in other T3 projects are not counted, even if they
+ worked on this repository. Live, there is a project at `/home/jwoltje` and
+ a deleted one at `/mnt/storage/src`.
+- **Diagnostic.** `t3.diagnostic.humanSentThroughApi` counts in-range user
+ messages the header rule calls human that T3 recorded as sent through its
+ API (no `appVersion` in the event's origin). Those are seat messages whose
+ header the rule doesn't accept, such as the older free-text Discord Bot
+ headers, and would show the next format drift. `humanWithoutEvent` counts
+ human messages with no `thread.message-sent` event. This is T3's internal
+ metadata, so it feeds no table or total. If `orchestration_events` or a
+ column it needs is missing, or an event doesn't parse, both read `unknown`.
+
+These refuse the report with exit 1, and the ones about the database name
+`--no-t3`: a missing, unreadable or unopenable database (including a busy
+lock past the timeout); a symlink at `~/.t3`, `~/.t3/userdata` or
+`state.sqlite` (with `--t3-db`, the file or its directory); a missing table or
+column the counts need; no project or more than one for this root; a message
+in a counted thread with a role other than `user` or `assistant`, non-text
+content, or a `created_at` that doesn't parse; a header conflict. A missing Pi
+directory means no Pi seats ran here; a missing T3 database means the path or
+T3 changed, so it refuses instead of counting zero. Error messages name ids
+and paths, never message text.
+
## One Gitea call and missing evidence
The client requests issues updated since the start date, all states, first page,
@@ -62,8 +133,8 @@ Use a narrower range or `--no-issues`, not hidden pagination. A commit-linked
issue not returned by the updated-since query still has a row, with unknown
metadata. This is the cost of the brief's one-call boundary.
-Exit 0 means a report was computed. Exit 1 means bad arguments or unreadable git
-or session evidence. Malformed JSONL, including a partially written last line,
+Exit 0 means a report was computed. Exit 1 means bad arguments or unreadable git,
+session or T3 evidence. Malformed JSONL, including a partially written last line,
refuses the report; rerun after the seat finishes writing. Exit 2 means issue
credentials, API, payload, or completeness failure. The CLI never prints API
error bodies or reads authentication files itself. `--no-issues` makes no API
@@ -72,8 +143,9 @@ median duration, and human-per-closed ratio. It cannot invent close-only rows.
For fixtures, a fake `gitea-api.sh` can be placed first on PATH. Otherwise the
repository scripts directory is appended to PATH for the issue request.
-Tests use only temporary repositories, logs, and fake API tools, with no real
-credentials or network. The helper regression stubs Node before any credential
+Tests use only temporary repositories, logs, T3 databases and fake API tools,
+with no real credentials or network. Every CLI run in the tests sets `HOME` to
+a temporary directory, so no test opens the real `~/.t3`. The helper regression stubs Node before any credential
read and checks successful GET, successful POST, and failed HTTP status.
## Acceptance
diff --git a/packages/ledger/src/cli.mjs b/packages/ledger/src/cli.mjs
index 66cd1417..5cf6705c 100644
--- a/packages/ledger/src/cli.mjs
+++ b/packages/ledger/src/cli.mjs
@@ -1,11 +1,12 @@
#!/usr/bin/env node
import path from 'node:path';
import { fileURLToPath } from 'node:url';
-import { dateRange, readCommits, readIssues, readSessions, summarize, formatTable, SourceError } from './ledger.mjs';
+import { dateRange, readCommits, readIssues, readSessions, mergeSources, summarize, formatTable, SourceError } from './ledger.mjs';
+import { readT3, defaultT3Path } from './t3.mjs';
-const usage = 'Usage: node packages/ledger/src/cli.mjs --since YYYY-MM-DD [--until YYYY-MM-DD] [--json] [--no-issues]';
+const usage = 'Usage: node packages/ledger/src/cli.mjs --since YYYY-MM-DD [--until YYYY-MM-DD] [--json] [--no-issues] [--no-t3 | --t3-db PATH]';
export async function main(args, root = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '../../..')) {
- let since, until, json = false, noIssues = false;
+ let since, until, t3Db, json = false, noIssues = false, noT3 = false;
const seen = new Set();
for (let i = 0; i < args.length; i++) {
const flag = args[i];
@@ -14,13 +15,18 @@ export async function main(args, root = path.resolve(path.dirname(fileURLToPath(
if (flag === '--help') { console.log(usage); return; }
if (flag === '--json') json = true;
else if (flag === '--no-issues') noIssues = true;
- else if (flag === '--since' || flag === '--until') {
+ else if (flag === '--no-t3') noT3 = true;
+ else if (flag === '--t3-db') {
+ t3Db = args[++i];
+ if (!t3Db || t3Db.startsWith('--')) throw new SourceError('--t3-db requires a path');
+ } else if (flag === '--since' || flag === '--until') {
const value = args[++i];
if (!value || value.startsWith('--')) throw new SourceError(`${flag} requires a date`);
if (flag === '--since') since = value; else until = value;
} else throw new SourceError('Unknown option; ' + usage);
}
if (!since) throw new SourceError(usage);
+ if (noT3 && t3Db !== undefined) throw new SourceError('--no-t3 and --t3-db cannot be combined');
const range = dateRange(since, until);
const commits = readCommits(root, range);
// Fixture tools may be placed first on PATH. The repository client is the
@@ -30,7 +36,9 @@ export async function main(args, root = path.resolve(path.dirname(fileURLToPath(
let issues;
try { issues = noIssues ? null : readIssues(root, range); }
finally { if (priorPath === undefined) delete process.env.PATH; else process.env.PATH = priorPath; }
- const sessions = await readSessions(root, range);
+ // T3 is on by default. A missing or unreadable database refuses the report.
+ const t3 = noT3 ? null : await readT3(root, range, t3Db === undefined ? { dbPath: defaultT3Path(), isDefault: true } : { dbPath: t3Db, isDefault: false });
+ const sessions = mergeSources(await readSessions(root, range), t3);
const report = summarize(range, commits, issues, sessions);
console.log(json ? JSON.stringify(report, null, 2) : formatTable(report));
return report;
diff --git a/packages/ledger/src/ledger.mjs b/packages/ledger/src/ledger.mjs
index b09e95f5..3669d1cc 100644
--- a/packages/ledger/src/ledger.mjs
+++ b/packages/ledger/src/ledger.mjs
@@ -1,7 +1,6 @@
import { execFileSync } from 'node:child_process';
import { createReadStream } from 'node:fs';
import { readdir, lstat } from 'node:fs/promises';
-import { createInterface } from 'node:readline';
import path from 'node:path';
const DAY = 86400000;
@@ -23,7 +22,7 @@ export function dateRange(since, until = new Date().toISOString().slice(0, 10))
if (end <= start) throw new SourceError('--until must not precede --since');
return { since, until, start, end };
}
-const inRange = (value, range) => {
+export const inRange = (value, range) => {
const ms = typeof value === 'number' ? value : Date.parse(value);
return Number.isFinite(ms) && ms >= range.start && ms < range.end;
};
@@ -75,15 +74,32 @@ export function messageText(content) {
if (Array.isArray(content)) return content.filter(c => c?.type === 'text' && typeof c.text === 'string').map(c => c.text).join('\n');
return '';
}
+// Classes are matched in either case: seats send DECISION and REVIEW-REQUEST.
+// tmux preamble from agent-send.sh: [host:session -> host:session class=x]
+const TMUX = /^\[([^\s:\[\]]+):([^\s\[\]]+) -> ([^\s:\[\]]+):([^\s\[\]]+)(?: class=[A-Za-z-]+)?\](?:\s|$)/;
+// T3 header (docs/guides/T3-AGENT-COMMS.md): [from: role (id) -> to: role (id) class=x]
+const T3 = /^\[from: ([^\s()\[\]]+) \(([^()\[\]]+)\) -> to: ([^\s()\[\]]+) \(([^()\[\]]+)\)(?: class=[A-Za-z-]+)?\](?:\s|$)/;
+const firstLine = text => text.split(/\r?\n/, 1)[0];
+export function t3Header(text) {
+ const m = firstLine(text).match(T3);
+ return m ? { from: m[1], fromId: m[2], to: m[3], toId: m[4] } : null;
+}
export function messageKind(text) {
- const firstLine = text.split(/\r?\n/, 1)[0];
- // tmux preamble from agent-send.sh: [host:session -> host:session class=x]
- const tmux = firstLine.match(/^\[([^\s:\[\]]+):([^\s\[\]]+) -> ([^\s:\[\]]+):([^\s\[\]]+)(?: class=[a-z-]+)?\](?:\s|$)/);
- // T3 header (docs/guides/T3-AGENT-COMMS.md): [from: role (id) -> to: role (id) class=x]
- const t3 = firstLine.match(/^\[from: ([^\s()\[\]]+) \(([^()\[\]]+)\) -> to: ([^\s()\[\]]+) \(([^()\[\]]+)\)(?: class=[a-z-]+)?\](?:\s|$)/);
- const sender = tmux ? tmux[2] : t3 ? t3[1] : null;
+ const tmux = firstLine(text).match(TMUX), t3 = t3Header(text);
+ const sender = tmux ? tmux[2] : t3 ? t3.from : null;
return sender === null ? 'human' : sender === 'control-board' ? 'board' : 'agent';
}
+// JSONL lines end at \n only. readline also ends a line at U+2028, which JSON
+// allows raw inside a string, so it split valid records (Node 26.8.1).
+async function* jsonLines(input) {
+ let rest = '';
+ for await (const chunk of input) {
+ const parts = (rest + chunk).split('\n');
+ rest = parts.pop();
+ yield* parts;
+ }
+ if (rest) yield rest;
+}
async function directories(dir, optional = false) {
try {
if (!(await lstat(dir)).isDirectory()) throw new SourceError('Session source must be a real directory');
@@ -93,11 +109,15 @@ async function directories(dir, optional = false) {
throw new SourceError(`Cannot read ledger directory: ${dir}`);
}
}
+// Seats are the real directories in agents/, sorted.
+export async function readSeats(root) {
+ return (await directories(path.join(root, 'agents'))).filter(e => e.isDirectory()).map(e => e.name).sort((a, b) => a.localeCompare(b));
+}
export async function readSessions(root, range) {
const rows = [];
const mentions = new Map();
// No symlink traversal, no fleet paths, no transcript content in the report.
- const agents = (await directories(path.join(root, 'agents'))).filter(e => e.isDirectory()).sort((a, b) => a.name.localeCompare(b.name));
+ const agents = (await readSeats(root)).map(name => ({ name }));
const state = path.join(root, '.pi', 'state');
// Check every source ancestor, not only the leaf directory.
if (!(await directories(path.join(root, '.pi'), true)).length) return { rows, mentions };
@@ -109,11 +129,10 @@ export async function readSessions(root, range) {
const files = (await directories(dir, true)).filter(e => e.isFile() && e.name.endsWith('.jsonl'));
const row = { seat: agent.name, board: 0, agent: 0, human: 0 };
for (const file of files) {
- const input = createReadStream(path.join(dir, file.name));
- const lines = createInterface({ input, crlfDelay: Infinity });
+ const input = createReadStream(path.join(dir, file.name), { encoding: 'utf8' });
let lineNumber = 0;
try {
- for await (const line of lines) {
+ for await (const line of jsonLines(input)) {
lineNumber++;
if (!line.trim()) continue;
let entry;
@@ -132,12 +151,33 @@ export async function readSessions(root, range) {
mentions.get(number).add(agent.name);
}
}
- } finally { lines.close(); input.destroy(); }
+ } finally { input.destroy(); }
}
if (row.board + row.agent + row.human) rows.push(row);
}
return { rows, mentions };
}
+// Adds T3 counts to the Pi rows per seat. Unmapped T3 threads get one row,
+// last. The report keeps the Pi rows and the T3 section so the split shows.
+export function mergeSources(pi, t3, unmapped = 't3:unmapped') {
+ if (!t3) return { rows: pi.rows, mentions: pi.mentions, pi: pi.rows, t3: { read: false } };
+ const bySeat = new Map(pi.rows.map(r => [r.seat, { ...r }]));
+ for (const [seat, counts] of t3.rows) {
+ if (seat === unmapped || !(counts.board + counts.agent + counts.human)) continue;
+ const row = bySeat.get(seat) ?? { seat, board: 0, agent: 0, human: 0 };
+ for (const kind of ['board', 'agent', 'human']) row[kind] += counts[kind];
+ bySeat.set(seat, row);
+ }
+ const rows = [...bySeat.values()].sort((a, b) => a.seat.localeCompare(b.seat));
+ const extra = t3.rows.get(unmapped);
+ if (extra.board + extra.agent + extra.human) rows.push({ seat: unmapped, ...extra });
+ const mentions = new Map([...pi.mentions].map(([n, seats]) => [n, new Set(seats)]));
+ for (const [n, seats] of t3.mentions) {
+ if (!mentions.has(n)) mentions.set(n, new Set());
+ for (const seat of seats) mentions.get(n).add(seat);
+ }
+ return { rows, mentions, pi: pi.rows, t3: t3.section };
+}
const round = value => Math.round(value * 10) / 10;
function duration(issue) {
if (!issue) return UNKNOWN;
@@ -163,13 +203,14 @@ export function summarize(range, commits, issues, sessions) {
const median = hours.includes(UNKNOWN) ? UNKNOWN : hours.length ?
round(hours.length % 2 ? hours[middle] : (hours[middle - 1] + hours[middle]) / 2) : 0;
const human = sessions.rows.reduce((sum, r) => sum + r.human, 0);
- return { since: range.since, until: range.until, timezone: 'UTC', issues: rows, seats: sessions.rows,
+ const sources = sessions.t3 ? { pi: sessions.pi, t3: sessions.t3 } : {};
+ return { since: range.since, until: range.until, timezone: 'UTC', issues: rows, seats: sessions.rows, ...sources,
totals: { issuesClosed: issues === null ? UNKNOWN : closed.length,
medianHoursOpen: issues === null ? UNKNOWN : median, commits: commits.length,
followUpsPerIssue: rows.length ? round(rows.reduce((sum, r) => sum + r.followUps, 0) / rows.length) : 0,
humanMessagesPerClosedIssue: issues === null ? UNKNOWN : closed.length ? round(human / closed.length) : human ? UNKNOWN : 0 } };
}
-const clean = value => String(value).replace(/[\x00-\x1f\x7f-\x9f]/g, ' ');
+export const clean = value => String(value).replace(/[\x00-\x1f\x7f-\x9f]/g, ' ');
const decimal = value => typeof value === 'number' ? value.toFixed(1) : value;
export function totalsLine(t) {
return `Totals: issues closed ${t.issuesClosed} | median hours open ${decimal(t.medianHoursOpen)} | commits ${t.commits} | follow-ups per issue ${decimal(t.followUpsPerIssue)} | human messages per closed issue ${decimal(t.humanMessagesPerClosedIssue)}`;
@@ -181,5 +222,12 @@ export function formatTable(report) {
decimal(r.hoursOpen), r.commits, r.followUps, r.seats.join(', ')].join(' | ')),
'', 'Seat | Board | Agent | Human',
...report.seats.map(r => [clean(r.seat), r.board, r.agent, r.human].join(' | ')),
+ ...t3Line(report.t3),
'', totalsLine(report.totals)].join('\n');
}
+// One line when T3 was skipped or read from somewhere other than the default.
+function t3Line(t3) {
+ if (!t3) return [];
+ if (!t3.read) return ['T3: not read (--no-t3)'];
+ return t3.database.default ? [] : [`T3: read from ${clean(t3.database.path)}, not the default`];
+}
diff --git a/packages/ledger/tests/ledger.test.mjs b/packages/ledger/tests/ledger.test.mjs
index 175f5d5e..8e60b96b 100644
--- a/packages/ledger/tests/ledger.test.mjs
+++ b/packages/ledger/tests/ledger.test.mjs
@@ -1,6 +1,7 @@
import test from 'node:test';
import assert from 'node:assert/strict';
-import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, cpSync, symlinkSync } from 'node:fs';
+import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, cpSync, symlinkSync, realpathSync } from 'node:fs';
+import { DatabaseSync } from 'node:sqlite';
import os from 'node:os';
import path from 'node:path';
import { fileURLToPath } from 'node:url';
@@ -9,13 +10,50 @@ import { dateRange, messageKind, issueNumbers, totalsLine, summarize } from '../
const source = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '../src');
const range = dateRange('2026-09-06', '2026-09-12');
+// T3 fixture schema: the live tables, cut to the columns the reader uses plus
+// one it doesn't. `text` allows NULL so a non-text row can be tested.
+const T3_SCHEMA = `
+ create table projection_projects (project_id text primary key, title text not null, workspace_root text not null, deleted_at text);
+ create table projection_threads (thread_id text primary key, project_id text not null, title text not null, archived_at text, deleted_at text);
+ create table projection_thread_messages (message_id text primary key, thread_id text not null, role text not null, text, created_at text not null);
+ create table orchestration_events (sequence integer primary key autoincrement, stream_id text not null, event_type text not null, payload_json text not null, metadata_json text not null);`;
+// Writes a T3 database in WAL mode. Threads default to project p1, which is
+// the fixture root. Returns the open writer when keepOpen is set.
+function t3db(file, { root, projects, threads = [], messages = [], after = [], keepOpen = false }) {
+ mkdirSync(path.dirname(file), { recursive: true });
+ for (const old of [file, `${file}-wal`, `${file}-shm`]) rmSync(old, { force: true });
+ const db = new DatabaseSync(file);
+ db.exec('pragma journal_mode=wal'); db.exec(T3_SCHEMA);
+ for (const [id, workspace, deleted = null] of projects ?? [['p1', root]]) {
+ db.prepare('insert into projection_projects values (?, ?, ?, ?)').run(id, 'project', workspace, deleted);
+ }
+ for (const t of threads) {
+ db.prepare('insert into projection_threads values (?, ?, ?, ?, ?)').run(t.id, t.project ?? 'p1', t.title, t.archived ?? null, t.deleted ?? null);
+ }
+ for (const m of messages) addMessage(db, m);
+ for (const sql of after) db.exec(sql);
+ if (keepOpen) return db;
+ db.close();
+}
+let messageId = 0;
+function addMessage(db, { thread, text, role = 'user', at = '2026-09-08T12:00:00Z', origin = 'app' }) {
+ const id = `m${++messageId}`;
+ db.prepare('insert into projection_thread_messages values (?, ?, ?, ?, ?)').run(id, thread, role, text, at);
+ if (origin !== 'none') db.prepare('insert into orchestration_events (stream_id, event_type, payload_json, metadata_json) values (?, ?, ?, ?)')
+ .run(thread, 'thread.message-sent', JSON.stringify({ messageId: id, threadId: thread, role, text }), JSON.stringify({ origin: origin === 'app' ? { appVersion: '0.0.0' } : {} }));
+}
const fixtureIssues = [
{ number: 1, title: 'First issue', created_at: '2026-09-06T00:00:00Z', closed_at: '2026-09-07T12:00:00Z' },
{ number: 2, title: 'Second issue', created_at: '2026-09-06T00:00:00Z', closed_at: null },
];
function fixture(t) {
const root = mkdtempSync(path.join(os.tmpdir(), 'ledger-test-'));
- t.after(() => rmSync(root, { recursive: true, force: true }));
+ // No test opens the real ~/.t3: every CLI run gets this HOME, with an empty
+ // T3 database at the default path. The CLI's root is a realpath.
+ const home = mkdtempSync(path.join(os.tmpdir(), 'ledger-home-'));
+ t.after(() => { rmSync(root, { recursive: true, force: true }); rmSync(home, { recursive: true, force: true }); });
+ const defaultDb = path.join(home, '.t3/userdata/state.sqlite');
+ t3db(defaultDb, { root: realpathSync(root) });
const put = (name, data) => { const p = path.join(root, name); mkdirSync(path.dirname(p), { recursive: true }); writeFileSync(p, data); return p; };
const git = (args, date = '2026-09-07T00:00:00Z') => execFileSync('git', args, { cwd: root, env: { ...process.env, GIT_AUTHOR_DATE: date, GIT_COMMITTER_DATE: date, GIT_CONFIG_NOSYSTEM: '1', GIT_CONFIG_GLOBAL: '/dev/null' }, stdio: 'pipe' });
git(['init', '-b', 'refactor']); git(['config', 'user.email', '[email protected]']); git(['config', 'user.name', 'Fixture']);
@@ -30,8 +68,8 @@ function fixture(t) {
const entry = (text, timestamp = '2026-09-08T12:00:00Z') => ({ type: 'message', timestamp, message: { role: 'user', content: [{ type: 'text', text }] } });
const logs = [entry('[host:control-board -> host:alice] do #1'), entry('[host:bob -> host:alice] review #2'), entry('build #2'), entry('old #1', '2026-09-05T23:59:59Z'), { type: 'message', timestamp: '2026-09-08T00:00:00Z', message: { role: 'assistant', content: 'not a user #1' } }];
put('.pi/state/alice/sessions/one.jsonl', logs.map(x => JSON.stringify(x)).join('\n') + '\n');
- const run = (args = [], env = {}) => spawnSync(process.execPath, [path.join(root, 'packages/ledger/src/cli.mjs'), '--since', '2026-09-06', '--until', '2026-09-12', ...args], { cwd: root, encoding: 'utf8', env: { ...process.env, PATH: `${path.join(root, 'bin')}:${process.env.PATH}`, ISSUES: path.join(root, 'issues.json'), CALLS: path.join(root, 'calls.jsonl'), ...env } });
- return { root, put, commit, run, entry, logs };
+ const run = (args = [], env = {}) => spawnSync(process.execPath, [path.join(root, 'packages/ledger/src/cli.mjs'), '--since', '2026-09-06', '--until', '2026-09-12', ...args], { cwd: root, encoding: 'utf8', env: { ...process.env, HOME: home, PATH: `${path.join(root, 'bin')}:${process.env.PATH}`, ISSUES: path.join(root, 'issues.json'), CALLS: path.join(root, 'calls.jsonl'), ...env } });
+ return { root, real: realpathSync(root), home, defaultDb, put, commit, run, entry, logs };
}
test('fixture git subjects only, follow-ups and three session kinds', t => {
const f = fixture(t), result = f.run(['--json']);
@@ -63,7 +101,7 @@ test('missing credentials exit 2, no-issues never calls API and shows unknown',
test('empty range gives no rows and zero totals', t => {
const f = fixture(t);
f.put('issues.json', '[]');
- const result = spawnSync(process.execPath, [path.join(f.root, 'packages/ledger/src/cli.mjs'), '--since', '2027-01-01', '--until', '2027-01-02', '--json'], { encoding: 'utf8', env: { ...process.env, PATH: `${f.root}/bin:${process.env.PATH}`, ISSUES: `${f.root}/issues.json`, CALLS: `${f.root}/calls.jsonl` } });
+ const result = spawnSync(process.execPath, [path.join(f.root, 'packages/ledger/src/cli.mjs'), '--since', '2027-01-01', '--until', '2027-01-02', '--json'], { encoding: 'utf8', env: { ...process.env, HOME: f.home, PATH: `${f.root}/bin:${process.env.PATH}`, ISSUES: `${f.root}/issues.json`, CALLS: `${f.root}/calls.jsonl` } });
assert.equal(result.status, 0, result.stderr); const r = JSON.parse(result.stdout);
assert.deepEqual(r.issues, []); assert.deepEqual(r.seats, []); assert.ok(Object.values(r.totals).every(n => n === 0));
});
@@ -96,6 +134,13 @@ test('partial or malformed session log refuses with location, not content', t =>
const f = fixture(t); f.put('.pi/state/alice/sessions/bad.jsonl', '{sensitive'); const r = f.run();
assert.equal(r.status, 1); assert.match(r.stderr, /Malformed session JSON: alice\/bad.jsonl:1/); assert.doesNotMatch(r.stderr, /sensitive/);
});
+test('a U+2028 inside a session string is one line, not a malformed record', t => {
+ const f = fixture(t);
+ f.put('.pi/state/bob/sessions/sep.jsonl', [f.entry('Jason: one\u2028two #2'), f.entry('[h:alice -> h:bob] ok')].map(x => JSON.stringify(x)).join('\r\n') + '\r\n');
+ assert.ok(readFileSync(path.join(f.root, '.pi/state/bob/sessions/sep.jsonl'), 'utf8').includes('\u2028'));
+ const r = f.run(['--json']); assert.equal(r.status, 0, r.stderr);
+ assert.deepEqual(JSON.parse(r.stdout).seats[1], { seat: 'bob', board: 0, agent: 1, human: 1 });
+});
test('no sessions is an empty table; symlink source refuses', t => {
const f = fixture(t); rmSync(path.join(f.root, '.pi'), { recursive: true });
assert.deepEqual(JSON.parse(f.run(['--json']).stdout).seats, []);
@@ -140,7 +185,11 @@ test('T3 header: agent, or board from control-board; anything short of the full
assert.equal(messageKind(`Jason here\n[from: ${sage} -> to: ${filbert}]\nquoted`), 'human');
assert.equal(messageKind(` [from: ${sage} -> to: ${filbert}]`), 'human');
assert.equal(messageKind(`[from: sage -> to: filbert]\nno thread ids`), 'human');
- assert.equal(messageKind(`[from: ${sage} -> to: ${filbert} class=Actionable]`), 'human');
+ // Classes match in either case (Gate F). HEAD before the fix called these human.
+ assert.equal(messageKind(`[from: ${sage} -> to: ${filbert} class=Actionable]`), 'agent');
+ assert.equal(messageKind(`[from: ${sage} -> to: ${filbert} class=REVIEW-REQUEST]\nreview`), 'agent');
+ assert.equal(messageKind('[h:sage -> h:bob class=DECISION] go'), 'agent');
+ assert.equal(messageKind(`[from: ${sage} -> to: ${filbert} class=review_request]`), 'human');
assert.equal(messageKind(`[from: ${sage} -> to: ${filbert}]trailing`), 'human');
assert.equal(messageKind(`[From: ${sage} -> to: ${filbert}]`), 'human');
});
@@ -148,3 +197,218 @@ test('no closed issues with human messages means undefined ratio, not invented z
const r = summarize(range, [], [], { rows: [{ seat: 'a', human: 1, board: 0, agent: 0 }], mentions: new Map() });
assert.equal(r.totals.humanMessagesPerClosedIssue, 'unknown');
});
+
+// T3 thread source (Gate F, docs/plans/2026-09-26_ledger-t3-source.md).
+const T1 = 't-alice', T2 = 't-bob', T3 = 't-sagebrush', T4 = 't-researcher', T5 = 't-discord';
+const header = (from, to, toId, cls = '') => `[from: ${from} (x1) -> to: ${to} (${toId})${cls}]`;
+function t3Fixture(t) {
+ const f = fixture(t);
+ for (const seat of ['sage', 'researcher']) mkdirSync(path.join(f.root, 'agents', seat), { recursive: true });
+ const db = path.join(f.home, 'fixture/t3.sqlite');
+ const threads = [
+ { id: T1, title: 'Alice' }, { id: T2, title: 'Bob in Claude', archived: '2026-09-09T00:00:00Z' },
+ { id: T3, title: 'Sagebrush' }, { id: T4, title: 'Researcher' }, { id: T5, title: 'Discord Bot' },
+ { id: 'import:claudeAgent:1', title: 'alice' }, { id: 't-deleted', title: 'Alice', deleted: '2026-09-09T00:00:00Z' },
+ { id: 't-other', project: 'p2', title: 'Alice' },
+ ];
+ const messages = [
+ { thread: T1, text: 'Jason: go #1' },
+ { thread: T1, text: `${header('sage', 'alice', T1, ' class=REVIEW-REQUEST')}\nreview #2`, origin: 'api' },
+ { thread: T1, text: '[h:sage -> h:alice class=DECISION] go', origin: 'api' },
+ { thread: T1, text: `${header('control-board', 'alice', T1)}\nbuzz`, origin: 'api' },
+ { thread: T1, text: 'outside', at: '2026-09-13T00:00:00Z' },
+ { thread: T1, text: 'an answer #9', role: 'assistant', origin: 'none' },
+ { thread: T2, text: 'archived still counts #2' },
+ { thread: T3, text: 'Sagebrush is not sage' },
+ { thread: T3, text: `${header('sage', 'discord', T3)}\nnot a seat role`, origin: 'api' },
+ { thread: T4, text: 'research this' },
+ { thread: T5, text: '[from: SetSpark coordinator (x1) -> to: Discord Bot (x2)]\nfree text', origin: 'api' },
+ { thread: 'import:claudeAgent:1', text: 'imported' },
+ { thread: 't-deleted', text: 'deleted' },
+ { thread: 't-other', text: 'other project' },
+ ];
+ const write = (overrides = {}) => t3db(db, { root: f.real, projects: [['p1', f.real], ['p2', '/elsewhere']], threads, messages, ...overrides });
+ return { ...f, db, threads, messages, write };
+}
+test('T3: seat, archived, unmapped and Researcher threads count; imported, deleted and other-project threads do not', t => {
+ const f = t3Fixture(t); f.write();
+ const result = f.run(['--json', '--t3-db', f.db]);
+ assert.equal(result.status, 0, result.stderr);
+ const r = JSON.parse(result.stdout);
+ assert.deepEqual(r.seats, [
+ { seat: 'alice', board: 2, agent: 3, human: 2 }, { seat: 'bob', board: 0, agent: 0, human: 1 },
+ { seat: 'researcher', board: 0, agent: 0, human: 1 }, { seat: 't3:unmapped', board: 0, agent: 1, human: 2 },
+ ]);
+ assert.deepEqual(r.pi, [{ seat: 'alice', board: 1, agent: 1, human: 1 }]);
+ assert.deepEqual(r.t3.database, { path: f.db, default: false });
+ assert.deepEqual(r.t3.seats, [
+ { seat: 'alice', board: 1, agent: 2, human: 1, threads: [{ id: T1, title: 'Alice', archived: false }] },
+ { seat: 'bob', board: 0, agent: 0, human: 1, threads: [{ id: T2, title: 'Bob in Claude', archived: true }] },
+ { seat: 'researcher', board: 0, agent: 0, human: 1, threads: [{ id: T4, title: 'Researcher', archived: false }] },
+ ]);
+ assert.deepEqual(r.t3.unmapped, { board: 0, agent: 1, human: 2, threads: [
+ { id: T5, title: 'Discord Bot', archived: false }, { id: T3, title: 'Sagebrush', archived: false }] });
+ assert.deepEqual(r.t3.excluded, { importedThreads: 1, deletedThreads: 1 });
+ // The free-text header counts as human; only the diagnostic shows it was sent through the API.
+ assert.deepEqual(r.t3.diagnostic, { humanSentThroughApi: 1, humanWithoutEvent: 0 });
+ assert.equal(r.totals.humanMessagesPerClosedIssue, 6);
+ assert.deepEqual(r.issues.map(x => [x.issue, x.seats]), [[1, ['alice']], [2, ['alice', 'bob']]]);
+ const text = f.run(['--t3-db', f.db]);
+ assert.equal(text.status, 0, text.stderr);
+ assert.ok(text.stdout.includes(`T3: read from ${f.db}, not the default`));
+ assert.match(text.stdout, /t3:unmapped \| 0 \| 1 \| 2/);
+});
+test('T3: the default path is read from HOME and prints no path line; --no-t3 says so', t => {
+ const f = t3Fixture(t); f.write();
+ rmSync(f.defaultDb); cpSync(f.db, f.defaultDb);
+ const json = JSON.parse(f.run(['--json']).stdout);
+ assert.deepEqual(json.t3.database, { path: f.defaultDb, default: true });
+ assert.equal(json.seats.at(-1).seat, 't3:unmapped');
+ const text = f.run(); assert.equal(text.status, 0, text.stderr); assert.doesNotMatch(text.stdout, /^T3:/m);
+ rmSync(path.join(f.home, '.t3'), { recursive: true });
+ const off = f.run(['--no-t3']); assert.equal(off.status, 0, off.stderr);
+ assert.match(off.stdout, /^T3: not read \(--no-t3\)$/m);
+ const offJson = JSON.parse(f.run(['--no-t3', '--json']).stdout);
+ assert.deepEqual(offJson.t3, { read: false }); assert.deepEqual(offJson.seats, [{ seat: 'alice', board: 1, agent: 1, human: 1 }]);
+ const both = f.run(['--no-t3', '--t3-db', f.db]); assert.equal(both.status, 1); assert.match(both.stderr, /cannot be combined/);
+ assert.equal(f.run(['--t3-db']).status, 1);
+});
+test('T3: a HOME with no database exits 1 and names --no-t3', t => {
+ const f = fixture(t); rmSync(path.join(f.home, '.t3'), { recursive: true });
+ const r = f.run(); assert.equal(r.status, 1); assert.equal(r.stdout, '');
+ assert.match(r.stderr, /T3 database unavailable: .*\.t3 is missing or unreadable; use --no-t3/);
+});
+test('T3: a file that is not a database exits 1 and names --no-t3', t => {
+ const f = fixture(t); writeFileSync(f.defaultDb, 'not sqlite'.repeat(100));
+ const r = f.run(); assert.equal(r.status, 1); assert.match(r.stderr, /T3 database cannot be read: .*\(SQLite \d+\); use --no-t3/);
+});
+test('T3: a seat thread renamed to another seat exits 1 naming thread, title and roles', t => {
+ const f = t3Fixture(t);
+ f.write({ messages: [...f.messages, { thread: T2, text: `${header('sage', 'alice', T2, ' class=INFO')}\nfor alice`, origin: 'api' }] });
+ const r = f.run(['--t3-db', f.db]); assert.equal(r.status, 1);
+ assert.equal(r.stderr.trim(), `T3 header conflict: thread ${T2} "Bob in Claude" maps to bob, but a header addresses alice`);
+});
+test('T3: an unmapped thread addressed as a seat exits 1', t => {
+ const f = t3Fixture(t);
+ f.write({ messages: [...f.messages, { thread: T3, text: `${header('bob', 'Sage', T3)}\nhi`, origin: 'api' }] });
+ const r = f.run(['--t3-db', f.db]); assert.equal(r.status, 1);
+ assert.match(r.stderr, /thread t-sagebrush "Sagebrush" maps to no seat, but a header addresses Sage/);
+});
+test('T3: a header to another thread id is not cross-checked', t => {
+ const f = t3Fixture(t);
+ f.write({ messages: [...f.messages, { thread: T2, text: `${header('sage', 'alice', T1)}\ncopied`, origin: 'api' }] });
+ const r = f.run(['--json', '--t3-db', f.db]); assert.equal(r.status, 0, r.stderr);
+ assert.equal(JSON.parse(r.stdout).t3.seats[1].agent, 1);
+});
+test('T3: no project, or two, for this root exits 1', t => {
+ const f = t3Fixture(t);
+ f.write({ projects: [['p1', `${f.real}-link`], ['p2', '/elsewhere']] });
+ let r = f.run(['--t3-db', f.db]); assert.equal(r.status, 1); assert.match(r.stderr, /T3 has no project for .*symlink does not match/);
+ f.write({ projects: [['p1', f.real], ['p2', f.real]] });
+ r = f.run(['--t3-db', f.db]); assert.equal(r.status, 1); assert.match(r.stderr, /T3 has more than one project for/);
+ f.write({ projects: [['p1', f.real], ['p2', f.real, '2026-09-01T00:00:00Z']] });
+ assert.equal(f.run(['--t3-db', f.db]).status, 0);
+});
+for (const [name, after, pattern] of [
+ ['a removed column', ['alter table projection_threads drop column title'], /T3 schema changed: missing projection_threads.title/],
+ ['a missing table', ['drop table projection_thread_messages'], /T3 schema changed: missing projection_thread_messages$/m],
+]) test(`T3: ${name} exits 1 and names it`, t => {
+ const f = t3Fixture(t); f.write({ after, messages: [] });
+ const r = f.run(['--t3-db', f.db]); assert.equal(r.status, 1); assert.match(r.stderr, pattern);
+});
+for (const [name, message, pattern] of [
+ ['an unknown role', { role: 'system' }, /T3 message m\d+ in thread t-alice has an unknown role/],
+ ['non-text content', { text: null }, /has non-text content/],
+ ['an unparseable created_at', { at: 'yesterday' }, /has an invalid created_at/],
+]) test(`T3: a counted row with ${name} exits 1 without its text`, t => {
+ const f = t3Fixture(t);
+ f.write({ messages: [...f.messages, { thread: T1, text: 'secret words', ...message }] });
+ const r = f.run(['--t3-db', f.db]); assert.equal(r.status, 1); assert.match(r.stderr, pattern); assert.doesNotMatch(r.stderr, /secret/);
+});
+test('T3: a missing orchestration_events makes the diagnostic unknown and keeps the counts', t => {
+ const f = t3Fixture(t); f.write();
+ const before = JSON.parse(f.run(['--json', '--t3-db', f.db]).stdout);
+ f.write({ after: ['drop table orchestration_events'] });
+ const r = f.run(['--json', '--t3-db', f.db]); assert.equal(r.status, 0, r.stderr);
+ const after = JSON.parse(r.stdout);
+ assert.deepEqual(after.t3.diagnostic, { humanSentThroughApi: 'unknown', humanWithoutEvent: 'unknown' });
+ assert.deepEqual(after.seats, before.seats); assert.deepEqual(after.totals, before.totals);
+});
+for (const link of ['.t3', '.t3/userdata', '.t3/userdata/state.sqlite']) test(`T3: a symlink at ~/${link} exits 1`, t => {
+ const f = fixture(t), target = path.join(f.home, 'real', link);
+ mkdirSync(path.dirname(target), { recursive: true });
+ cpSync(path.join(f.home, link), target, { recursive: true });
+ rmSync(path.join(f.home, link), { recursive: true }); symlinkSync(target, path.join(f.home, link));
+ const r = f.run(); assert.equal(r.status, 1); assert.match(r.stderr, new RegExp(`${link.replaceAll('.', '\\.')} is a symlink; use --no-t3`));
+});
+test('T3: with --t3-db, a symlinked file or directory exits 1', t => {
+ const f = t3Fixture(t); f.write();
+ const file = path.join(f.home, 'file-link.sqlite'); symlinkSync(f.db, file);
+ let r = f.run(['--t3-db', file]); assert.equal(r.status, 1); assert.match(r.stderr, /file-link.sqlite is a symlink/);
+ const dir = path.join(f.home, 'dir-link'); symlinkSync(path.dirname(f.db), dir);
+ r = f.run(['--t3-db', path.join(dir, 't3.sqlite')]); assert.equal(r.status, 1); assert.match(r.stderr, /dir-link is a symlink/);
+});
+
+// WAL states. The CLI reads with mode=ro; it may create -wal and -shm but must
+// never change the main file.
+const sha = file => execFileSync('sha256sum', [file], { encoding: 'utf8' }).split(' ')[0];
+const humans = r => JSON.parse(r.stdout).t3.seats.find(s => s.seat === 'alice').human;
+const asRoot = process.getuid?.() === 0;
+function killedWriter(db) {
+ // A writer that commits into the WAL and dies without a checkpoint.
+ const code = `const { DatabaseSync } = require('node:sqlite'); const db = new DatabaseSync(${JSON.stringify(db)});
+ db.exec('pragma wal_autocheckpoint=0');
+ db.prepare("insert into projection_thread_messages values ('late', 't-alice', 'user', 'late human', '2026-09-08T13:00:00Z')").run();
+ process.kill(process.pid, 'SIGKILL');`;
+ const r = spawnSync(process.execPath, ['-e', code]);
+ assert.equal(r.signal, 'SIGKILL');
+ rmSync(`${db}-shm`);
+}
+function inReadOnlyDir(dir, check) {
+ execFileSync('chmod', ['0555', dir]);
+ try { check(); } finally { execFileSync('chmod', ['0755', dir]); }
+}
+test('T3 WAL: the newest message only in -wal, writer attached, is counted', t => {
+ const f = t3Fixture(t), writer = f.write({ keepOpen: true });
+ t.after(() => writer.close());
+ writer.exec('pragma wal_autocheckpoint=0');
+ addMessage(writer, { thread: T1, text: 'newest', at: '2026-09-08T13:00:00Z' });
+ const main = sha(f.db);
+ const r = f.run(['--json', '--t3-db', f.db]); assert.equal(r.status, 0, r.stderr);
+ assert.equal(humans(r), 2); assert.equal(sha(f.db), main);
+});
+test('T3 WAL: stopped cleanly, counts are correct and the main file is unchanged', t => {
+ const f = t3Fixture(t); f.write();
+ assert.throws(() => readFileSync(`${f.db}-wal`));
+ const main = sha(f.db);
+ const r = f.run(['--json', '--t3-db', f.db]); assert.equal(r.status, 0, r.stderr);
+ assert.equal(humans(r), 1); assert.equal(sha(f.db), main);
+});
+test('T3 WAL: -wal without -shm in a writable directory is read', t => {
+ const f = t3Fixture(t); f.write(); killedWriter(f.db);
+ const main = sha(f.db);
+ const r = f.run(['--json', '--t3-db', f.db]); assert.equal(r.status, 0, r.stderr);
+ assert.equal(humans(r), 2); assert.equal(sha(f.db), main);
+});
+test('T3 WAL: -wal without -shm in a read-only directory exits 1', { skip: asRoot && 'mode bits do not bind root' }, t => {
+ const f = t3Fixture(t); f.write(); killedWriter(f.db);
+ inReadOnlyDir(path.dirname(f.db), () => {
+ const r = f.run(['--t3-db', f.db]); assert.equal(r.status, 1); assert.match(r.stderr, /cannot be read: .*\(SQLite 14\); use --no-t3/);
+ });
+});
+test('T3 WAL: stopped cleanly in a read-only directory exits 1', { skip: asRoot && 'mode bits do not bind root' }, t => {
+ const f = t3Fixture(t); f.write();
+ inReadOnlyDir(path.dirname(f.db), () => {
+ const r = f.run(['--t3-db', f.db]); assert.equal(r.status, 1); assert.match(r.stderr, /cannot be read: .*\(SQLite 1544\); use --no-t3/);
+ });
+});
+test('T3: a lock held past the 5 s busy timeout exits 1 and names --no-t3', t => {
+ const f = t3Fixture(t), writer = f.write({ keepOpen: true });
+ t.after(() => writer.close());
+ writer.exec('pragma locking_mode=exclusive'); writer.exec('begin exclusive');
+ addMessage(writer, { thread: T1, text: 'held' });
+ const started = Date.now(), r = f.run(['--t3-db', f.db]);
+ writer.exec('commit');
+ assert.equal(r.status, 1); assert.match(r.stderr, /cannot be read: .*\(SQLite 5\); use --no-t3/);
+ assert.ok(Date.now() - started >= 4500, 'the reader waited for the busy timeout');
+});
diff --git a/packages/ledger/src/t3.mjs b/packages/ledger/src/t3.mjs
new file mode 100644
index 00000000..9672fcc9
--- /dev/null
+++ b/packages/ledger/src/t3.mjs
@@ -0,0 +1,151 @@
+import { lstat } from 'node:fs/promises';
+import { DatabaseSync } from 'node:sqlite';
+import { pathToFileURL } from 'node:url';
+import os from 'node:os';
+import path from 'node:path';
+import { SourceError, UNKNOWN, clean, inRange, issueNumbers, messageKind, readSeats, t3Header } from './ledger.mjs';
+
+// T3 keeps every thread message in one SQLite database. This reader opens that
+// file read-only and nothing else in ~/.t3. See
+// docs/plans/2026-09-26_ledger-t3-source.md for the rules below.
+export const UNMAPPED = 't3:unmapped';
+const SKIP = 'use --no-t3 to skip T3';
+const REQUIRED = {
+ projection_projects: ['project_id', 'workspace_root', 'deleted_at'],
+ projection_threads: ['thread_id', 'project_id', 'title', 'archived_at', 'deleted_at'],
+ projection_thread_messages: ['message_id', 'thread_id', 'role', 'text', 'created_at'],
+};
+const DIAGNOSTIC = { orchestration_events: ['stream_id', 'event_type', 'payload_json', 'metadata_json'] };
+
+export const defaultT3Path = () => path.join(os.homedir(), '.t3', 'userdata', 'state.sqlite');
+
+// Every named path must exist and must not be a symlink. Skipping one would be
+// a silent zero, so each problem refuses the report.
+async function checkPaths(dbPath, isDefault) {
+ const dirs = isDefault ? [path.dirname(path.dirname(dbPath)), path.dirname(dbPath)] : [path.dirname(dbPath)];
+ for (const [target, wantDir] of [...dirs.map(d => [d, true]), [dbPath, false]]) {
+ let stat;
+ try { stat = await lstat(target); }
+ catch { throw new SourceError(`T3 database unavailable: ${target} is missing or unreadable; ${SKIP}`); }
+ if (stat.isSymbolicLink()) throw new SourceError(`T3 database refused: ${target} is a symlink; ${SKIP}`);
+ if (wantDir ? !stat.isDirectory() : !stat.isFile()) {
+ throw new SourceError(`T3 database refused: ${target} is not a ${wantDir ? 'directory' : 'regular file'}; ${SKIP}`);
+ }
+ }
+}
+
+function missingColumns(db, tables) {
+ const missing = [];
+ for (const [table, columns] of Object.entries(tables)) {
+ const have = new Set(db.prepare('select name from pragma_table_info(?)').all(table).map(r => r.name));
+ if (!have.size) missing.push(table);
+ else for (const column of columns) if (!have.has(column)) missing.push(`${table}.${column}`);
+ }
+ return missing;
+}
+
+// Seat for a thread title: the lower-cased title equals the seat or starts
+// with the seat and a space. Longest seat first, so the most specific wins.
+export function seatForTitle(title, seats) {
+ const lower = title.toLowerCase();
+ return [...seats].sort((a, b) => b.length - a.length).find(s => lower === s || lower.startsWith(`${s} `)) ?? null;
+}
+
+// Origin per message id from thread.message-sent events. Any missing table,
+// column or unparseable event makes the diagnostic unknown; it decides nothing.
+function origins(db, projectId) {
+ if (missingColumns(db, DIAGNOSTIC).length) return null;
+ const byMessage = new Map();
+ const events = db.prepare(`select e.payload_json, e.metadata_json from orchestration_events e
+ join projection_threads t on t.thread_id = e.stream_id
+ where e.event_type = 'thread.message-sent' and t.project_id = ?`).all(projectId);
+ for (const event of events) {
+ let payload, metadata;
+ try { payload = JSON.parse(event.payload_json); metadata = JSON.parse(event.metadata_json); }
+ catch { return null; }
+ if (typeof payload?.messageId !== 'string') return null;
+ byMessage.set(payload.messageId, typeof metadata?.origin?.appVersion === 'string');
+ }
+ return byMessage;
+}
+
+function query(db, root, range, seats) {
+ const missing = missingColumns(db, REQUIRED);
+ if (missing.length) throw new SourceError(`T3 schema changed: missing ${missing.join(', ')}`);
+ // Compared in JavaScript so a declared collation can't loosen the match.
+ const projects = db.prepare('select project_id, workspace_root from projection_projects where deleted_at is null').all()
+ .filter(p => p.workspace_root === root);
+ if (projects.length !== 1) {
+ throw new SourceError(`T3 has ${projects.length ? 'more than one project' : 'no project'} for ${root}; a project opened through a symlink does not match; ${SKIP}`);
+ }
+ const projectId = projects[0].project_id;
+ const threads = new Map(), excluded = { importedThreads: 0, deletedThreads: 0 };
+ for (const t of db.prepare('select thread_id, title, archived_at, deleted_at from projection_threads where project_id = ?').all(projectId)) {
+ if (typeof t.thread_id !== 'string' || typeof t.title !== 'string') throw new SourceError('T3 thread with a non-text id or title');
+ if (t.thread_id.startsWith('import:')) { excluded.importedThreads++; continue; }
+ if (t.deleted_at !== null) { excluded.deletedThreads++; continue; }
+ threads.set(t.thread_id, { id: t.thread_id, title: t.title, archived: t.archived_at !== null, seat: seatForTitle(t.title, seats) });
+ }
+ const rows = new Map([...seats, UNMAPPED].map(s => [s, { board: 0, agent: 0, human: 0 }]));
+ const mentions = new Map(), human = [];
+ const messages = db.prepare(`select m.message_id, m.thread_id, m.role, m.text, m.created_at from projection_thread_messages m
+ join projection_threads t on t.thread_id = m.thread_id where t.project_id = ?`).all(projectId);
+ for (const m of messages) {
+ const thread = threads.get(m.thread_id);
+ if (!thread) continue;
+ const where = `T3 message ${clean(m.message_id)} in thread ${clean(m.thread_id)}`;
+ if (m.role !== 'user' && m.role !== 'assistant') throw new SourceError(`${where} has an unknown role`);
+ if (typeof m.text !== 'string') throw new SourceError(`${where} has non-text content`);
+ if (typeof m.created_at !== 'string' || !Number.isFinite(Date.parse(m.created_at))) throw new SourceError(`${where} has an invalid created_at`);
+ if (m.role !== 'user') continue;
+ // A header addressed to its own thread must agree with the title mapping.
+ const header = t3Header(m.text);
+ if (header && header.toId === thread.id) {
+ const to = header.to.toLowerCase();
+ if (thread.seat ? to !== thread.seat : seats.includes(to)) {
+ throw new SourceError(`T3 header conflict: thread ${clean(thread.id)} "${clean(thread.title)}" maps to ${thread.seat ?? 'no seat'}, but a header addresses ${clean(header.to)}`);
+ }
+ }
+ if (!inRange(m.created_at, range)) continue;
+ const kind = messageKind(m.text), seat = thread.seat ?? UNMAPPED;
+ rows.get(seat)[kind]++;
+ if (kind === 'human') human.push(m.message_id);
+ for (const number of issueNumbers(m.text)) {
+ if (!mentions.has(number)) mentions.set(number, new Set());
+ mentions.get(number).add(seat);
+ }
+ }
+ const byMessage = origins(db, projectId);
+ const sentThroughApi = byMessage === null ? UNKNOWN : human.filter(id => byMessage.get(id) === false).length;
+ const noEvent = byMessage === null ? UNKNOWN : human.filter(id => !byMessage.has(id)).length;
+ const listed = seat => [...threads.values()].filter(t => (t.seat ?? UNMAPPED) === seat)
+ .sort((a, b) => a.id.localeCompare(b.id)).map(({ id, title, archived }) => ({ id, title, archived }));
+ const seatRows = seats.map(seat => ({ seat, ...rows.get(seat), threads: listed(seat) })).filter(r => r.threads.length);
+ return { rows, mentions, excluded, seats: seatRows, unmapped: { ...rows.get(UNMAPPED), threads: listed(UNMAPPED) },
+ diagnostic: { humanSentThroughApi: sentThroughApi, humanWithoutEvent: noEvent } };
+}
+
+// Reads one snapshot of T3's database. Returns the per-seat rows and issue
+// mentions the ledger merges with Pi, and the report's `t3` section.
+export async function readT3(root, range, { dbPath = defaultT3Path(), isDefault = true } = {}) {
+ dbPath = path.resolve(dbPath);
+ await checkPaths(dbPath, isDefault);
+ const seats = await readSeats(root);
+ const url = pathToFileURL(dbPath);
+ url.searchParams.set('mode', 'ro');
+ let db, result;
+ try {
+ db = new DatabaseSync(url, { readOnly: true, timeout: 5000 });
+ db.exec('BEGIN');
+ result = query(db, root, range, seats);
+ db.exec('COMMIT');
+ } catch (error) {
+ if (error instanceof SourceError) throw error;
+ throw new SourceError(`T3 database cannot be read: ${dbPath} (SQLite ${error.errcode ?? 'error'}); ${SKIP}`);
+ } finally {
+ try { if (db?.isTransaction) db.exec('ROLLBACK'); } catch { /* the close below still runs */ }
+ try { db?.close(); } catch { /* nothing was written */ }
+ }
+ const { rows, mentions, ...section } = result;
+ return { rows, mentions, section: { read: true, database: { path: dbPath, default: isDefault }, ...section } };
+}
@@ -0,0 +1,3 @@
5acbc1075a5d0ad709faf14698235c8c2332c408cb4fccc580bd4a75e2c314fb packages/ledger/src/t3.mjs
6546dbaf59c046d599a9d378c1a1f2d9afc9487189db06f50fa3201c9cadc63b packages/ledger/tests/ledger.test.mjs
101013def3b168ae1b7e291ff86200b49ed6b27927787585d5ca32a283bf38bd packages/ledger/README.md
@@ -0,0 +1,67 @@
# Gate F follow-up: Filbert's notes 1 to 3 (#1506), candidate for review
Darkwing, 2026-09-26. Filbert's build review
(`agents/filbert/work/ledger-t3-build-review-2026-09-26.md`, e47ec6da) left
four nonblocking notes on Gate F (136958c9). Sage asked for 1 to 3 as one small
change that Filbert reviews and Sage commits. Note 4, snapshot isolation, went
to DEFERRED (a68dc174). Base is HEAD a4d38a3d, which changes nothing under
`packages/ledger` since 136958c9. Nothing is committed or pushed.
`followup-manifest.sha256` pins the three files. `followup.patch` is the diff
against a4d38a3d.
## Changes
1. **U+2029.** The splitter test now writes a Pi entry holding a raw U+2028
and a raw U+2029, with CRLF endings, and asserts that the file contains
both. The README line names both characters. `ledger.mjs` is unchanged,
because the splitter already ends lines at `\n` only.
2. **Diagnostic.** Three new tests:
- A human message with no `thread.message-sent` event gives
`humanWithoutEvent: 1`.
- A `thread.message-sent` event whose payload doesn't parse makes both
diagnostic fields `unknown` and leaves `seats` unchanged.
- The same for an event whose `messageId` isn't a string. Filbert didn't
list this one, but it's the third `return null` in `origins()` and had
no test either.
3. **Rethrow.** `readT3`'s catch now rethrows anything that is not a
`SourceError` and carries no numeric `errcode`. The CLI prints such an
error as `Ledger failed: cannot read source evidence`, exit 1. That's the
CLI's existing message for a non-source error, and it no longer points
at SQLite or `--no-t3`. With only numeric errcodes left, the message's
`?? 'error'` fallback could no longer fire, so I removed it. The new test
calls `readT3` in process with an explicit fixture path and a `null`
range, so `inRange` throws a `TypeError` inside the read transaction. It
asserts the `TypeError` comes out. It never touches the real `~/.t3`, and
the fixture comment says so.
## Evidence
- Ledger tests: 51/51, the Gate F 47 plus 4 new.
- Mutations on a scratch copy of the package. The three `gitea-helper` tests
fail in every scratch copy, as before, so the counts leave them out:
| Mutation | Result |
|---|---|
| `humanWithoutEvent` hardcoded to 0 | 1 fails (no-event test) |
| unparseable event skipped (`continue`) | 1 fails (unparseable test) |
| non-string `messageId` skipped | 1 fails (messageId test) |
| rethrow removed (Gate F catch) | 1 fails (rethrow test) |
| splitter also splits at U+2028 | 1 fails (splitter test) |
| splitter also splits at U+2029 | 1 fails (splitter test) |
My first try at the last two put a raw U+2028 or U+2029 in the regex
source. That ends a JS regex literal, so the whole test file failed to
load, which doesn't count as a kill. I reran with the escape written out
literally, and the rows above come from that rerun.
- Eight suites on a local clone of a4d38a3d with the three files: config 24,
task 90, foundation 43, conductor 17, release 14, auth 15, discord 63,
extension-package 18. I ran them twice, and the second run was on the final
files after the errcode edit.
- Union on the same clone. Control-board, webui, seat, mosaic, ledger and
discord, plus conversation, which CHAT-02 committed: 474/474 twice before
the errcode edit and once after. No `ledger-*` temp directories remained.
- Live read, `--since 2026-09-01 --until 2026-09-26 --no-issues --json`, at
2026-09-26T21:47Z: exit 0, no header conflict, diagnostic
`{humanSentThroughApi: 15, humanWithoutEvent: 0}`, two imported threads
excluded. The Gate F build read 14; messages have been sent since then.
@@ -0,0 +1,99 @@
diff --git a/packages/ledger/README.md b/packages/ledger/README.md
index 393e9c37..3a2ce27c 100644
--- a/packages/ledger/README.md
+++ b/packages/ledger/README.md
@@ -39,7 +39,8 @@ No install, build, service restart, or configuration change is needed.
duplicated entries in copied logs are not deduplicated. No transcript content
leaves the parser. Assistant messages and logs outside repo seats do not count.
Symlink source directories are refused and symlink files are not followed.
- A line ends at `\n` only. A U+2028 inside a JSON string does not split a record.
+ A line ends at `\n` only. A U+2028 or U+2029 inside a JSON string does not
+ split a record.
- Table 2 also counts T3 thread messages with role `user`. The T3 source
follows. A seat's row sums its Pi and T3 counts; the JSON keeps the split in
`pi` (Pi rows) and `t3.seats` (T3 rows).
diff --git a/packages/ledger/src/t3.mjs b/packages/ledger/src/t3.mjs
index 9672fcc9..fc9da14e 100644
--- a/packages/ledger/src/t3.mjs
+++ b/packages/ledger/src/t3.mjs
@@ -140,8 +140,10 @@ export async function readT3(root, range, { dbPath = defaultT3Path(), isDefault
result = query(db, root, range, seats);
db.exec('COMMIT');
} catch (error) {
- if (error instanceof SourceError) throw error;
- throw new SourceError(`T3 database cannot be read: ${dbPath} (SQLite ${error.errcode ?? 'error'}); ${SKIP}`);
+ // Only a SQLite failure carries an errcode. Anything else is a bug and
+ // surfaces as itself, not as a database problem.
+ if (error instanceof SourceError || typeof error?.errcode !== 'number') throw error;
+ throw new SourceError(`T3 database cannot be read: ${dbPath} (SQLite ${error.errcode}); ${SKIP}`);
} finally {
try { if (db?.isTransaction) db.exec('ROLLBACK'); } catch { /* the close below still runs */ }
try { db?.close(); } catch { /* nothing was written */ }
diff --git a/packages/ledger/tests/ledger.test.mjs b/packages/ledger/tests/ledger.test.mjs
index 8e60b96b..834dbd55 100644
--- a/packages/ledger/tests/ledger.test.mjs
+++ b/packages/ledger/tests/ledger.test.mjs
@@ -7,6 +7,7 @@ import path from 'node:path';
import { fileURLToPath } from 'node:url';
import { execFileSync, spawnSync } from 'node:child_process';
import { dateRange, messageKind, issueNumbers, totalsLine, summarize } from '../src/ledger.mjs';
+import { readT3 } from '../src/t3.mjs';
const source = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '../src');
const range = dateRange('2026-09-06', '2026-09-12');
@@ -49,7 +50,8 @@ const fixtureIssues = [
function fixture(t) {
const root = mkdtempSync(path.join(os.tmpdir(), 'ledger-test-'));
// No test opens the real ~/.t3: every CLI run gets this HOME, with an empty
- // T3 database at the default path. The CLI's root is a realpath.
+ // T3 database at the default path. The one in-process readT3 call passes an
+ // explicit fixture path. The CLI's root is a realpath.
const home = mkdtempSync(path.join(os.tmpdir(), 'ledger-home-'));
t.after(() => { rmSync(root, { recursive: true, force: true }); rmSync(home, { recursive: true, force: true }); });
const defaultDb = path.join(home, '.t3/userdata/state.sqlite');
@@ -134,10 +136,11 @@ test('partial or malformed session log refuses with location, not content', t =>
const f = fixture(t); f.put('.pi/state/alice/sessions/bad.jsonl', '{sensitive'); const r = f.run();
assert.equal(r.status, 1); assert.match(r.stderr, /Malformed session JSON: alice\/bad.jsonl:1/); assert.doesNotMatch(r.stderr, /sensitive/);
});
-test('a U+2028 inside a session string is one line, not a malformed record', t => {
+test('a U+2028 or U+2029 inside a session string is one line, not a malformed record', t => {
const f = fixture(t);
- f.put('.pi/state/bob/sessions/sep.jsonl', [f.entry('Jason: one\u2028two #2'), f.entry('[h:alice -> h:bob] ok')].map(x => JSON.stringify(x)).join('\r\n') + '\r\n');
- assert.ok(readFileSync(path.join(f.root, '.pi/state/bob/sessions/sep.jsonl'), 'utf8').includes('\u2028'));
+ f.put('.pi/state/bob/sessions/sep.jsonl', [f.entry('Jason: one\u2028two\u2029three #2'), f.entry('[h:alice -> h:bob] ok')].map(x => JSON.stringify(x)).join('\r\n') + '\r\n');
+ const written = readFileSync(path.join(f.root, '.pi/state/bob/sessions/sep.jsonl'), 'utf8');
+ assert.ok(written.includes('\u2028') && written.includes('\u2029'));
const r = f.run(['--json']); assert.equal(r.status, 0, r.stderr);
assert.deepEqual(JSON.parse(r.stdout).seats[1], { seat: 'bob', board: 0, agent: 1, human: 1 });
});
@@ -334,6 +337,30 @@ test('T3: a missing orchestration_events makes the diagnostic unknown and keeps
assert.deepEqual(after.t3.diagnostic, { humanSentThroughApi: 'unknown', humanWithoutEvent: 'unknown' });
assert.deepEqual(after.seats, before.seats); assert.deepEqual(after.totals, before.totals);
});
+test('T3: a human message with no event counts in humanWithoutEvent', t => {
+ const f = t3Fixture(t);
+ f.write({ messages: [...f.messages, { thread: T1, text: 'typed, no event', origin: 'none' }] });
+ const r = f.run(['--json', '--t3-db', f.db]); assert.equal(r.status, 0, r.stderr);
+ assert.deepEqual(JSON.parse(r.stdout).t3.diagnostic, { humanSentThroughApi: 1, humanWithoutEvent: 1 });
+});
+const badEvent = payload => `insert into orchestration_events (stream_id, event_type, payload_json, metadata_json) values ('${T1}', 'thread.message-sent', '${payload}', '{}')`;
+for (const [name, payload] of [['an unparseable event', '{bad'], ['an event with no string messageId', '{"messageId":7}']]) {
+ test(`T3: ${name} makes the diagnostic unknown and keeps the counts`, t => {
+ const f = t3Fixture(t); f.write();
+ const before = JSON.parse(f.run(['--json', '--t3-db', f.db]).stdout);
+ f.write({ after: [badEvent(payload)] });
+ const r = f.run(['--json', '--t3-db', f.db]); assert.equal(r.status, 0, r.stderr);
+ const after = JSON.parse(r.stdout);
+ assert.deepEqual(after.t3.diagnostic, { humanSentThroughApi: 'unknown', humanWithoutEvent: 'unknown' });
+ assert.deepEqual(after.seats, before.seats);
+ });
+}
+test('T3: an error that is not from SQLite is rethrown, not reported as a database failure', async t => {
+ const f = t3Fixture(t); f.write();
+ // In process with an explicit path, so the real ~/.t3 stays closed. A null
+ // range makes inRange throw a TypeError inside the read transaction.
+ await assert.rejects(readT3(f.real, null, { dbPath: f.db, isDefault: false }), TypeError);
+});
for (const link of ['.t3', '.t3/userdata', '.t3/userdata/state.sqlite']) test(`T3: a symlink at ~/${link} exits 1`, t => {
const f = fixture(t), target = path.join(f.home, 'real', link);
mkdirSync(path.dirname(target), { recursive: true });
@@ -0,0 +1,5 @@
08959a05574264e4f8243a90af94746e73a2fde3706f22e38f4ff1105b7a45a8 agents/darkwing/work/ledger-t3-source/r1.md
e8300cb6abea70819aba7cf10040d19b4d6019b5663c37203209537a5f10ee62 agents/darkwing/work/ledger-t3-source/r2.md
f3c05c1b4d28a419ab817621e15708b147f71dff588664980e22696eaafbc342 docs/plans/2026-09-26_ledger-t3-source.md
aa4740ae5d045aa12971af5de36a5107805bd31da839aae4e4d398a818d471fc agents/darkwing/work/ledger-t3-source/r1-to-r2.diff
4128375121673e0a49ef6789390503ff7395ba59f87bfc450764ea56b536c5b2 agents/darkwing/work/ledger-t3-source/r2-to-r3.diff
@@ -0,0 +1,419 @@
--- r1.md
+++ docs/plans/2026-09-26_ledger-t3-source.md
@@ -1,8 +1,11 @@
# Ledger: a read-only T3 thread source for Table 2 (Gate F brief)
-Brief only, no code. Darkwing wrote it on 2026-09-26 at Sage's request. Filbert
-reviews it, and Jason sees it on the decision sheet before anyone builds it.
-Issue #1506.
+Brief only, no code. Darkwing wrote it on 2026-09-26 at Sage's request, issue
+#1506. R1 (sha256 08959a05) went to Filbert, whose review asked for
+revisions: `agents/filbert/work/ledger-t3-source-review-2026-09-26.md`, sha256
+19dda29a. This is R2. It takes every finding, and it records Sage's rulings
+on the three open questions. Section 1 has one measurement that differs from
+the review.
## Why
@@ -18,8 +21,8 @@
## Where T3 keeps messages
T3 keeps its state in one SQLite database, `~/.t3/userdata/state.sqlite`, in
-WAL mode (`state.sqlite-wal` and `state.sqlite-shm` sit beside it). Three
-projection tables are enough:
+WAL mode (`state.sqlite-wal` and `state.sqlite-shm` sit beside it). The counts
+need three projection tables:
- `projection_projects`: `project_id`, `workspace_root`, `deleted_at`.
- `projection_threads`: `thread_id`, `project_id`, `title`, `archived_at`,
@@ -27,51 +30,72 @@
- `projection_thread_messages`: `message_id` (primary key), `thread_id`,
`role` (`user` or `assistant`), `text`, `created_at` (ISO UTC).
-One more table is optional. In `orchestration_events`, each
+The JSON diagnostic reads one more. In `orchestration_events`, each
`thread.message-sent` event carries `metadata_json.origin`. Messages typed in
the T3 app carry an `appVersion` there. Messages sent through T3's API or MCP
-tools, which is how seats talk to each other, don't. See the cross-check below.
+tools, which is how seats talk to each other, don't.
The same directory also holds `secrets/`, `clerk-tokens.json` and other
settings files. The reader opens `state.sqlite` and nothing else, and it
selects named columns only, never `*`.
-## Reading it, with T3 running or not
+## 1. Reading it, with T3 running or not
-The file stays on disk whether T3 runs or not. The reader opens it with Node's
-built-in `node:sqlite` (`DatabaseSync`, `file:<path>?mode=ro`, `readOnly:
-true`). That needs no dependency, and Node 26.8.1 prints no warning for it. I
-read the live database this way today, while T3 was running, with no errors
-and no locks. A WAL reader sees every committed message, including those still
-in the `-wal` file.
-
-Two rules:
-- Never open with `immutable=1` and never copy the file. Both skip the WAL
- and silently lose the newest messages. A copy of the three files is also
- not atomic.
-- If T3 stopped uncleanly and left a `-wal` without its `-shm`, a read-only
- connection may be unable to rebuild the index. If the open fails, the
- ledger reports it and refuses. I have not tested this case or the fully
- stopped case. Both are acceptance checks below.
+The reader uses Node's built-in `node:sqlite` (`DatabaseSync`). That needs no
+dependency, and Node 26.8.1 (SQLite 3.53.4) prints no warning for it.
-## Jason or agent
+- **URI.** Build it with `pathToFileURL(dbPath)` and set `mode=ro` through
+ `searchParams`, then pass `readOnly: true`. A `?`, `#` or `%` in the home
+ path would break a string-built URI.
+- **One snapshot.** Run every query, from the schema checks through the
+ diagnostic, inside one `BEGIN` … `COMMIT`. In autocommit mode each
+ statement sees its own snapshot while T3 writes between them.
+- **Busy timeout.** Set `DatabaseSync`'s `timeout` to 5 s. A transient
+ `SQLITE_BUSY` during a T3 checkpoint then waits instead of failing. A busy
+ error after the timeout exits 1 like any open failure.
+- **No `immutable=1` and no copy.** Both lose the WAL. Filbert found worse
+ than lost messages: with a table created inside the WAL, `immutable=1`
+ fails with `no such table`.
+
+What happens on disk. Filbert and I both tested these in scratch
+directories:
+
+| State | Directory writable | Result |
+|---|---|---|
+| T3 running, writer attached, newest rows only in `-wal` | yes | reads them |
+| `-wal` without `-shm` (writer killed, `-shm` removed) | yes | reads the WAL rows and creates `-shm` |
+| `-wal` without `-shm` | no | open fails, SQLite 14 |
+| T3 stopped cleanly, no `-wal` or `-shm` | yes | reads, then leaves an empty `-wal` and a 32 KiB `-shm` |
+| T3 stopped cleanly | no | fails, SQLite 1544 "attempt to write a readonly database" |
+
+In every case the main file's bytes stayed the same. The last row is where
+Filbert and I differ. His review says the stopped-case read works with the
+directory read-only. In my run it failed with and without the read
+transaction. The build's test settles it. Either way a failed open is exit 1.
+
+So the accurate claim: the reader never writes the main database file. Like
+any SQLite connection, it may create or update `-wal` and `-shm` beside it
+and takes read locks in `-shm`. T3 opens normally afterwards.
+
+## 2. Jason or agent
Reuse the 6a rule. The first line of `text` decides: a T3 header or the tmux
preamble counts as agent, `control-board` as the sender counts as board, and
anything else counts as human. Messages with role `user` count; assistant
messages don't.
-6a has a defect this source would expose. Its regex allows only a lowercase
-class (`class=[a-z-]+`). Seats send uppercase classes: Sage's DECISION, INFO,
-REVIEW-REQUEST and REVIEW-NOTE, and my own REVIEW-REQUEST. In this project's
-threads, 16 real agent headers fail on that alone and would count as human.
-The fix is to make the class match case-insensitive. It belongs in this build
-or just before it, reviewed with it. The ms-communications table lists
-lowercase names, so the fix follows what seats send, not the table.
+The class fix rides in this build (Sage's ruling). HEAD's
+`packages/ledger/src/ledger.mjs:81` (tmux) and `:83` (T3) both allow only
+`class=[a-z-]+`. Both become case-insensitive. Seats send uppercase classes:
+Sage's DECISION, INFO, REVIEW-REQUEST and REVIEW-NOTE, and my own
+REVIEW-REQUEST. In this project's threads 16 real agent headers failed on
+that alone at 20:54Z. The ms-communications table lists lowercase names, so
+the fix follows what seats send, not the table.
Cross-check, read at 2026-09-26T20:54Z for the mosaic-stack project (209
user messages outside imported and deleted threads, every one with its
-`thread.message-sent` event):
+`thread.message-sent` event). Filbert's later read agreed, plus messages sent
+since.
| T3 origin | Header matches 6a | Count |
|---|---|---|
@@ -83,110 +107,193 @@
No message typed in the app carries a header, and every API message in this
project carries one of the three forms. The 14 free-text ones are older
Discord Bot thread headers such as `[from: SetSpark coordinator (…) -> to:
-Discord Bot (…)]`, written before the guide fixed the format. With the class
-fix they still count as human. That's 14 wrong human counts, all dated
-2026-09-17 to 2026-09-22.
-
-Recommendation: the header rule decides, as Sage asked. The reader also
-reports one diagnostic number, not used in any table: user messages the rule
-calls human that T3 recorded as sent through the API. That count is how the
-uppercase-class bug showed up, and it would catch the next format drift. The
-origin field is T3's internal metadata, not a documented contract, so it
-shouldn't decide anything. I'd make it JSON only, so Table 2's layout stays
-the same.
-
-## Thread to seat
-
-A thread counts for this checkout only if its project's `workspace_root` is
-the ledger's repository root. That is `/mnt/storage/src/mosaic-stack`, project
-`34050c07`.
-
-Thread IDs change whenever Jason starts a new thread for a seat, so there's no
-fixed map. T3-AGENT-COMMS.md already names threads after the seat ("Darkwing",
-"Sage", "Dewey in Claude"). Proposed rule: a thread belongs to seat `<s>` when
-`<s>` is a real directory under `agents/` and the lower-cased title equals
-`<s>` or starts with `<s>` followed by a space. Several threads can map to one
-seat. Their counts add up, as several Pi session files already do.
+Discord Bot (…)]`, written before the guide fixed the format. Sage ruled they
+stay as recorded: they count as human, dated 2026-09-17 to 2026-09-22.
+
+The header rule decides. The JSON also carries one diagnostic that feeds no
+table or total: user messages the rule calls human that T3 recorded as sent
+through the API. That number exposed the class bug and would catch the next
+format drift. `origin` is T3's internal metadata, not a documented contract,
+so it decides nothing. If `orchestration_events` or a column it needs is
+missing, the diagnostic reads `unknown` and the report goes on (Sage's
+ruling on F5). Missing tables the counts depend on still exit 1.
+
+## 3. Thread to seat
+
+**Project.** A thread counts for this checkout only if its project's
+`workspace_root` equals the ledger's repository root, byte for byte. The CLI
+already takes that root from the realpath of its own URL, today
+`/mnt/storage/src/mosaic-stack`, project `34050c07`. So a T3 project opened
+through the compatibility symlink `~/src/mosaic-stack-dev-test` doesn't
+match, and "no project row" is the right refusal. The README says so.
+
+**Title rule.** Thread IDs change whenever Jason starts a new thread for a
+seat, so there's no fixed map. T3-AGENT-COMMS.md already names threads after
+the seat ("Darkwing", "Sage", "Dewey in Claude"). A thread belongs to seat
+`<s>` when `<s>` is a real directory under `agents/` and the lower-cased
+title equals `<s>` or starts with `<s>` followed by a space. So "Sagebrush"
+stays unmapped. Several threads can map to one seat, and their counts add
+up, as several Pi session files already do.
Today that maps Sage, Darkwing, Filbert, Dewey and Rocko (one thread each,
-created 2026-09-26), plus "Darkwing in Claude" (archived) and "Dewey in
-Claude". Three threads map to no seat. Two are imported and excluded anyway
-("FINDINGS.md review" and "[dragon-lin:darkwing -> …"). The third is
-"Discord Bot" with 68 user messages: 54 without a header, and the 14
-free-text headers above. The guide's own advice, titles like `review:
-<topic>`, will produce more unmapped threads.
-
-Unmapped threads go in one Table 2 row, `t3:unmapped`, so Jason's messages
-there still count toward the Human column and the human-per-closed ratio. The
-other choice is to drop them, which would hide those 54 headerless prompts.
-That is Jason's decision. I recommend the row.
-
-A seat's row sums its Pi and T3 counts. JSON splits them by source. Nothing is
-counted twice: every T3 session today runs on `claudeAgent` or `codex`, which
-don't write `.pi/state`, and Filbert found no T3 header in any Pi log.
+created 2026-09-26, titles set by hand), plus "Darkwing in Claude" (archived)
+and "Dewey in Claude". Researcher has a directory and no thread. Three
+threads map to no seat. Two are imported and excluded anyway ("FINDINGS.md
+review" and "[dragon-lin:darkwing -> …"). The third is "Discord Bot" with 68
+user messages: 54 without a header, and the 14 free-text headers.
+
+Titles are current state, and T3 can write them itself. They go wrong three
+ways. T3 auto-titles an unnamed thread from Jason's first prompt, so "Rocko
+review of the plan" maps to rocko. A rename moves the whole history to
+another row. A seat thread titled for a topic drops into `t3:unmapped`.
+None of this changes the Human total or the human-per-closed ratio. It only
+moves counts between rows, but Gate F reads one seat's row.
+
+**Header cross-check.** The headers already say which seat a thread belongs
+to. For every user message whose header matches the fixed 6a rule and whose
+`to:` id equals the message's own `thread_id`:
+- in a mapped thread, the `to:` role, lower-cased, must equal that thread's
+ seat;
+- in an unmapped thread, the `to:` role must not be a seat name.
+
+A conflict exits 1 and names the thread id, its title and both roles. A
+header whose `to:` id is some other thread is not checked. The check reads
+message text only, not T3 metadata. In a live read at 21:02Z every header
+agreed: all 104 addressed to their own thread carried the full thread id and
+named that thread's seat (Sage 40, Darkwing 15, Filbert 18, Dewey 15, Rocko
+16).
+
+It catches a seat thread renamed to another seat or to a topic, once any
+agent writes to it. It also catches an auto-titled thread that agents
+address by a different seat. It misses a thread no agent ever writes to.
+Such a thread can only add human counts to a seat's row, never hide them, so
+for Gate F it errs toward a visible failure. The README says so.
+
+**Unmapped row.** Unmapped threads go in one Table 2 row, `t3:unmapped`
+(Sage's ruling), so their human messages still reach the Human column and
+the human-per-closed ratio.
+
+**Mapping in the JSON.** For each seat, the T3 thread ids and titles that
+made its row, and the unmapped thread ids and titles. Anyone checking a Gate
+F result can then see which threads the row came from.
+
+A seat's row sums its Pi and T3 counts, and the JSON splits them by source.
+Nothing is counted twice. Every T3 session today runs on `claudeAgent` or
+`codex`, which don't write `.pi/state`, and Filbert found no T3 header in any
+Pi log (6a record).
-Excluded, with the reason stated in the README:
+**Excluded,** with the reason stated in the README:
- Imported threads (`thread_id` starting `import:`, events marked
`historyImport`). They are partial copies of Claude Code sessions, not T3
traffic: 55 user messages in two threads here.
- Deleted threads (`deleted_at` set). Across all projects there are 3, with
3 messages. Archived threads count.
-
-## What fails closed
-
-With the T3 source on, each of these refuses the report with exit 1, the
-code the ledger already uses for unreadable session evidence. The report
-never falls back to Pi logs alone. As with `--no-issues`, `--no-t3` turns the
-source off, and the report then says T3 was not read.
-- The database is missing, unreadable, or won't open read-only (including
- the `-wal` without `-shm` case). This differs from the Pi reader, which
- treats a missing `.pi` as no messages. A missing Pi directory means no Pi
- seats ran here. A missing T3 database on this host means the path or T3
- changed, and a silent zero is the failure Gate F exists to prevent.
-- A required table or column is missing. The reader checks `PRAGMA
+- Threads in other T3 projects. Live, there is a project at `/home/jwoltje`
+ and a deleted one at `/mnt/storage/src`. A thread in either could work on
+ this repository and would not be counted. The workspace-root rule is still
+ the right one, but the README names this blind spot.
+
+## 4. What fails closed
+
+The source is on by default (Sage's ruling). `--no-t3` turns it off, and the
+report then says T3 was not read. `--t3-db <path>` reads another database
+file instead of `~/.t3/userdata/state.sqlite`. It exists for fixtures and
+gets the same checks.
+
+Each of these refuses the report with exit 1, the code the ledger already
+uses for unreadable session evidence. The report never falls back to Pi logs
+alone. Where the database is missing or won't open, the message names
+`--no-t3`.
+- The database is missing or unreadable, or won't open read-only. That
+ includes a directory that isn't writable when SQLite needs to create
+ `-shm`, and a busy error after the timeout. The Pi reader treats a missing
+ `.pi` as no messages, and this departs from it on purpose. A missing Pi
+ directory means no Pi seats ran here. A missing T3 database on this host
+ means the path or T3 changed, and a silent zero is the failure Gate F
+ exists to prevent.
+- `~/.t3`, `~/.t3/userdata` or `state.sqlite` is a symlink. With `--t3-db`,
+ the file and its directory are checked. The Pi reader checks every
+ ancestor too, but it skips symlinked entries. Skipping one named file
+ would be another silent zero, so this reader refuses.
+- A table or column the counts need is missing. The reader checks `PRAGMA
table_info` and names what's missing. This catches a T3 upgrade that
changes the schema.
- No project row, or more than one non-deleted row, for this repository root.
- A counted row has a bad `role`, non-string `text`, or a `created_at` that
doesn't parse. The Pi reader already refuses malformed JSONL and bad
timestamps the same way.
-- `state.sqlite` or `~/.t3/userdata` is a symlink. The Pi reader skips
- symlinked entries instead. For one named file, skipping would be another
- silent zero, so this reader refuses.
-
-The source never writes to the database. It never reads other files in
-`~/.t3`, and it passes no message text beyond `messageKind` and
-`issueNumbers`, the same rule as for Pi logs. The one outside effect is
-SQLite's own: a WAL reader takes read locks in the `-shm` file, as T3's own
-connections do.
-
-## Decisions for Jason
-
-1. The source is on by default, with `--no-t3` to turn it off. The other
- choice is off by default with `--t3` to turn it on. I recommend on by
- default, because Gate F exists to count these messages.
-2. Unmapped threads get a `t3:unmapped` row. The other choice is to drop
- them. I recommend the row.
-3. The 14 free-text headers from 09-17 to 09-22 stay counted as human. Fixing
- them would mean loosening the header grammar for history only, and I don't
- recommend it.
-
-## Acceptance for the build
-
-- Fixture databases built with `node:sqlite` in a temp dir, in WAL mode:
- seat threads and an unmapped thread; imported, deleted and archived
- threads; all three header forms, uppercase classes included; a message
- outside the date range; another project with the same seat titles.
-- Each fail-closed case above has its own test, including a schema column
- removed and `-wal` without `-shm`. One test opens a database whose newest
- message is still in the WAL and counts it. Another reads a database closed
- cleanly with no writer attached, which is the T3-stopped case.
-- The class fix is proven against HEAD's `messageKind`: an uppercase class
- counts as agent after the fix and as human before it.
+- A header conflicts with the title mapping (section 3).
+
+The reader never reads other files in `~/.t3`. It passes no message text
+beyond `messageKind`, `issueNumbers` and the header's `to:` role and id, the
+same rule as for Pi logs.
+
+## 5. Rulings
+
+Sage ruled on the three questions R1 put to Jason, as lead calls:
+1. The source is on by default. A missing or unreadable database exits 1,
+ and the message names `--no-t3`.
+2. Unmapped threads get the `t3:unmapped` row.
+3. The 14 free-text headers stay as recorded. They show only in the JSON
+ diagnostic.
+
+Sage also ruled that the class fix rides in this build, and that a missing
+diagnostic table reads `unknown` (F5).
+
+## 6. Acceptance for the build
+
+**No test opens the real `~/.t3`.** Both places in
+`packages/ledger/tests/ledger.test.mjs` that spawn `cli.mjs` (the shared
+`run()` helper and the direct `spawnSync` at line 66) set `HOME` to the
+fixture's temp directory. A test that forgets `--t3-db` or `--no-t3` then
+finds no database and fails closed. The existing tests aren't about T3. Each
+gets an empty fixture database at the fixture `HOME`'s default path, with
+one project row for the fixture root. So they run with the source on, and
+their expected rows don't change. One test asserts that a `HOME` with no
+database exits 1 and names `--no-t3`.
+
+Fixture databases are built with `node:sqlite` in a temp directory, in WAL
+mode:
+- seat threads and an unmapped thread; imported, deleted and archived
+ threads; a message outside the date range;
+- all three header forms, with uppercase classes in both the tmux preamble
+ and the T3 header;
+- a thread with the same seat title in another project;
+- a seat thread renamed to another seat, with an agent header to its own
+ id, which exits 1;
+- a "Sagebrush" title, which stays unmapped;
+- a thread titled "Researcher", which maps to the seat that has no thread
+ live.
+
+WAL states, each with its own test:
+- The newest message is only in `-wal`, with the writer still attached (the
+ live-T3 case). It is counted.
+- T3 stopped: the database closed cleanly with no writer. Counts are
+ correct, and the main file's bytes are unchanged afterwards.
+- `-wal` without `-shm` in a writable directory: made by a child writer with
+ `wal_autocheckpoint=0` that is SIGKILLed, then `-shm` deleted. The WAL
+ rows are counted.
+- `-wal` without `-shm` in a directory that isn't writable: exit 1, naming
+ `--no-t3`. Skipped when the tests run as root, where the mode bits don't
+ bind.
+- The stopped case in a directory that isn't writable: exit 1, naming
+ `--no-t3`, as I measured it. If the build reads there instead, the builder
+ changes this test to assert correct counts and records the correction in
+ the BUILD-LOG entry. Skipped as root too.
+
+Also:
+- Each other fail-closed case in section 4 has its own test, including a
+ removed schema column and each symlink.
+- A missing `orchestration_events` gives `unknown` for the diagnostic and
+ the same counts.
+- The class fix is proven against HEAD's `messageKind`. An uppercase class
+ in either preamble counts as agent after the fix and as human before it.
+- The JSON lists each seat's threads and the unmapped threads.
- A read against the live database gives the counts in this brief, allowing
- for messages sent since.
-- The ledger README's counting rules name the new source, the mapping rule
- and the exclusions.
+ for messages sent since. It exits 0 with no header conflict.
+- The ledger README's counting rules name the new source, both flags, the
+ mapping rule and the header check, and the exclusions. That includes the
+ symlinked-checkout case and the other-project blind spot.
- No suite runs the ledger tests, so the BUILD-LOG entry names the test file.
## Not in scope
@@ -194,5 +301,6 @@
- Claude Code transcripts (`~/.claude/projects`) and Codex sessions
(`~/.codex/sessions`). T3's database already holds every message T3
delivered, so those files would only duplicate it.
-- Any write to T3, any T3 API call, or anything that needs T3 running.
+- Any T3 API call, anything that needs T3 running, and any write beyond
+ SQLite's own `-wal` and `-shm` handling.
- Fixing the two tmux misclassifications Filbert found in 6a.
+198
View File
@@ -0,0 +1,198 @@
# Ledger: a read-only T3 thread source for Table 2 (Gate F brief)
Brief only, no code. Darkwing wrote it on 2026-09-26 at Sage's request. Filbert
reviews it, and Jason sees it on the decision sheet before anyone builds it.
Issue #1506.
## Why
Table 2 counts user messages per seat from `.pi/state/<seat>/sessions/*.jsonl`
only. Development seats now run in T3 on the Claude and Codex harnesses, so
their prompts, Jason's included, never reach a Pi log. Today the Human column
can't see T3 at all, and the zero it shows for T3 seats means "no source", not
"no human prompts". 6a (ef0020ad) taught `messageKind` the T3 header, but no
source the ledger reads contains one. Gate F (QUEUE row 6) passes when
Filbert's item closes with zero human messages from Jason. While the ledger
can't see T3, a zero there proves nothing.
## Where T3 keeps messages
T3 keeps its state in one SQLite database, `~/.t3/userdata/state.sqlite`, in
WAL mode (`state.sqlite-wal` and `state.sqlite-shm` sit beside it). Three
projection tables are enough:
- `projection_projects`: `project_id`, `workspace_root`, `deleted_at`.
- `projection_threads`: `thread_id`, `project_id`, `title`, `archived_at`,
`deleted_at`.
- `projection_thread_messages`: `message_id` (primary key), `thread_id`,
`role` (`user` or `assistant`), `text`, `created_at` (ISO UTC).
One more table is optional. In `orchestration_events`, each
`thread.message-sent` event carries `metadata_json.origin`. Messages typed in
the T3 app carry an `appVersion` there. Messages sent through T3's API or MCP
tools, which is how seats talk to each other, don't. See the cross-check below.
The same directory also holds `secrets/`, `clerk-tokens.json` and other
settings files. The reader opens `state.sqlite` and nothing else, and it
selects named columns only, never `*`.
## Reading it, with T3 running or not
The file stays on disk whether T3 runs or not. The reader opens it with Node's
built-in `node:sqlite` (`DatabaseSync`, `file:<path>?mode=ro`, `readOnly:
true`). That needs no dependency, and Node 26.8.1 prints no warning for it. I
read the live database this way today, while T3 was running, with no errors
and no locks. A WAL reader sees every committed message, including those still
in the `-wal` file.
Two rules:
- Never open with `immutable=1` and never copy the file. Both skip the WAL
and silently lose the newest messages. A copy of the three files is also
not atomic.
- If T3 stopped uncleanly and left a `-wal` without its `-shm`, a read-only
connection may be unable to rebuild the index. If the open fails, the
ledger reports it and refuses. I have not tested this case or the fully
stopped case. Both are acceptance checks below.
## Jason or agent
Reuse the 6a rule. The first line of `text` decides: a T3 header or the tmux
preamble counts as agent, `control-board` as the sender counts as board, and
anything else counts as human. Messages with role `user` count; assistant
messages don't.
6a has a defect this source would expose. Its regex allows only a lowercase
class (`class=[a-z-]+`). Seats send uppercase classes: Sage's DECISION, INFO,
REVIEW-REQUEST and REVIEW-NOTE, and my own REVIEW-REQUEST. In this project's
threads, 16 real agent headers fail on that alone and would count as human.
The fix is to make the class match case-insensitive. It belongs in this build
or just before it, reviewed with it. The ms-communications table lists
lowercase names, so the fix follows what seats send, not the table.
Cross-check, read at 2026-09-26T20:54Z for the mosaic-stack project (209
user messages outside imported and deleted threads, every one with its
`thread.message-sent` event):
| T3 origin | Header matches 6a | Count |
|---|---|---|
| typed in the app (has `appVersion`) | no | 99 |
| sent through the API (no `appVersion`) | yes | 80 |
| sent through the API | no, uppercase class | 16 |
| sent through the API | no, free-text roles | 14 |
No message typed in the app carries a header, and every API message in this
project carries one of the three forms. The 14 free-text ones are older
Discord Bot thread headers such as `[from: SetSpark coordinator (…) -> to:
Discord Bot (…)]`, written before the guide fixed the format. With the class
fix they still count as human. That's 14 wrong human counts, all dated
2026-09-17 to 2026-09-22.
Recommendation: the header rule decides, as Sage asked. The reader also
reports one diagnostic number, not used in any table: user messages the rule
calls human that T3 recorded as sent through the API. That count is how the
uppercase-class bug showed up, and it would catch the next format drift. The
origin field is T3's internal metadata, not a documented contract, so it
shouldn't decide anything. I'd make it JSON only, so Table 2's layout stays
the same.
## Thread to seat
A thread counts for this checkout only if its project's `workspace_root` is
the ledger's repository root. That is `/mnt/storage/src/mosaic-stack`, project
`34050c07`.
Thread IDs change whenever Jason starts a new thread for a seat, so there's no
fixed map. T3-AGENT-COMMS.md already names threads after the seat ("Darkwing",
"Sage", "Dewey in Claude"). Proposed rule: a thread belongs to seat `<s>` when
`<s>` is a real directory under `agents/` and the lower-cased title equals
`<s>` or starts with `<s>` followed by a space. Several threads can map to one
seat. Their counts add up, as several Pi session files already do.
Today that maps Sage, Darkwing, Filbert, Dewey and Rocko (one thread each,
created 2026-09-26), plus "Darkwing in Claude" (archived) and "Dewey in
Claude". Three threads map to no seat. Two are imported and excluded anyway
("FINDINGS.md review" and "[dragon-lin:darkwing -> …"). The third is
"Discord Bot" with 68 user messages: 54 without a header, and the 14
free-text headers above. The guide's own advice, titles like `review:
<topic>`, will produce more unmapped threads.
Unmapped threads go in one Table 2 row, `t3:unmapped`, so Jason's messages
there still count toward the Human column and the human-per-closed ratio. The
other choice is to drop them, which would hide those 54 headerless prompts.
That is Jason's decision. I recommend the row.
A seat's row sums its Pi and T3 counts. JSON splits them by source. Nothing is
counted twice: every T3 session today runs on `claudeAgent` or `codex`, which
don't write `.pi/state`, and Filbert found no T3 header in any Pi log.
Excluded, with the reason stated in the README:
- Imported threads (`thread_id` starting `import:`, events marked
`historyImport`). They are partial copies of Claude Code sessions, not T3
traffic: 55 user messages in two threads here.
- Deleted threads (`deleted_at` set). Across all projects there are 3, with
3 messages. Archived threads count.
## What fails closed
With the T3 source on, each of these refuses the report with exit 1, the
code the ledger already uses for unreadable session evidence. The report
never falls back to Pi logs alone. As with `--no-issues`, `--no-t3` turns the
source off, and the report then says T3 was not read.
- The database is missing, unreadable, or won't open read-only (including
the `-wal` without `-shm` case). This differs from the Pi reader, which
treats a missing `.pi` as no messages. A missing Pi directory means no Pi
seats ran here. A missing T3 database on this host means the path or T3
changed, and a silent zero is the failure Gate F exists to prevent.
- A required table or column is missing. The reader checks `PRAGMA
table_info` and names what's missing. This catches a T3 upgrade that
changes the schema.
- No project row, or more than one non-deleted row, for this repository root.
- A counted row has a bad `role`, non-string `text`, or a `created_at` that
doesn't parse. The Pi reader already refuses malformed JSONL and bad
timestamps the same way.
- `state.sqlite` or `~/.t3/userdata` is a symlink. The Pi reader skips
symlinked entries instead. For one named file, skipping would be another
silent zero, so this reader refuses.
The source never writes to the database. It never reads other files in
`~/.t3`, and it passes no message text beyond `messageKind` and
`issueNumbers`, the same rule as for Pi logs. The one outside effect is
SQLite's own: a WAL reader takes read locks in the `-shm` file, as T3's own
connections do.
## Decisions for Jason
1. The source is on by default, with `--no-t3` to turn it off. The other
choice is off by default with `--t3` to turn it on. I recommend on by
default, because Gate F exists to count these messages.
2. Unmapped threads get a `t3:unmapped` row. The other choice is to drop
them. I recommend the row.
3. The 14 free-text headers from 09-17 to 09-22 stay counted as human. Fixing
them would mean loosening the header grammar for history only, and I don't
recommend it.
## Acceptance for the build
- Fixture databases built with `node:sqlite` in a temp dir, in WAL mode:
seat threads and an unmapped thread; imported, deleted and archived
threads; all three header forms, uppercase classes included; a message
outside the date range; another project with the same seat titles.
- Each fail-closed case above has its own test, including a schema column
removed and `-wal` without `-shm`. One test opens a database whose newest
message is still in the WAL and counts it. Another reads a database closed
cleanly with no writer attached, which is the T3-stopped case.
- The class fix is proven against HEAD's `messageKind`: an uppercase class
counts as agent after the fix and as human before it.
- A read against the live database gives the counts in this brief, allowing
for messages sent since.
- The ledger README's counting rules name the new source, the mapping rule
and the exclusions.
- No suite runs the ledger tests, so the BUILD-LOG entry names the test file.
## Not in scope
- Claude Code transcripts (`~/.claude/projects`) and Codex sessions
(`~/.codex/sessions`). T3's database already holds every message T3
delivered, so those files would only duplicate it.
- Any write to T3, any T3 API call, or anything that needs T3 running.
- Fixing the two tmux misclassifications Filbert found in 6a.
@@ -0,0 +1,74 @@
--- r2.md
+++ docs/plans/2026-09-26_ledger-t3-source.md
@@ -3,9 +3,9 @@
Brief only, no code. Darkwing wrote it on 2026-09-26 at Sage's request, issue
#1506. R1 (sha256 08959a05) went to Filbert, whose review asked for
revisions: `agents/filbert/work/ledger-t3-source-review-2026-09-26.md`, sha256
-19dda29a. This is R2. It takes every finding, and it records Sage's rulings
-on the three open questions. Section 1 has one measurement that differs from
-the review.
+19dda29a. R2 (sha256 e8300cb6) took every finding and recorded Sage's
+rulings on the three open questions. Filbert approved R2 with three nits,
+review sha256 bb02d8d3. This is R3, which takes the nits.
## Why
@@ -68,10 +68,10 @@
| T3 stopped cleanly, no `-wal` or `-shm` | yes | reads, then leaves an empty `-wal` and a 32 KiB `-shm` |
| T3 stopped cleanly | no | fails, SQLite 1544 "attempt to write a readonly database" |
-In every case the main file's bytes stayed the same. The last row is where
-Filbert and I differ. His review says the stopped-case read works with the
-directory read-only. In my run it failed with and without the read
-transaction. The build's test settles it. Either way a failed open is exit 1.
+In every case the main file's bytes stayed the same. Filbert's first
+review said the last case reads. His test had reused a database whose empty
+`-wal` and `-shm` were still present. On a true clean stop he also got 1544,
+and his review records the correction. A failed open is exit 1.
So the accurate claim: the reader never writes the main database file. Like
any SQLite connection, it may create or update `-wal` and `-shm` beside it
@@ -198,7 +198,9 @@
The source is on by default (Sage's ruling). `--no-t3` turns it off, and the
report then says T3 was not read. `--t3-db <path>` reads another database
file instead of `~/.t3/userdata/state.sqlite`. It exists for fixtures and
-gets the same checks.
+gets the same checks. The JSON records the database path read and whether
+it was the default. When it wasn't, the text report adds one line naming the
+path, so a Gate F result can't come from a fixture unnoticed.
Each of these refuses the report with exit 1, the code the ledger already
uses for unreadable session evidence. The report never falls back to Pi logs
@@ -248,7 +250,9 @@
fixture's temp directory. A test that forgets `--t3-db` or `--no-t3` then
finds no database and fails closed. The existing tests aren't about T3. Each
gets an empty fixture database at the fixture `HOME`'s default path, with
-one project row for the fixture root. So they run with the source on, and
+one project row for the fixture root. That row stores
+`fs.realpathSync(root)`, because the CLI resolves its root through realpath
+and a symlinked temp directory would otherwise not match. So they run with the source on, and
their expected rows don't change. One test asserts that a `HOME` with no
database exits 1 and names `--no-t3`.
@@ -277,9 +281,7 @@
`--no-t3`. Skipped when the tests run as root, where the mode bits don't
bind.
- The stopped case in a directory that isn't writable: exit 1, naming
- `--no-t3`, as I measured it. If the build reads there instead, the builder
- changes this test to assert correct counts and records the correction in
- the BUILD-LOG entry. Skipped as root too.
+ `--no-t3`. Skipped as root too.
Also:
- Each other fail-closed case in section 4 has its own test, including a
@@ -288,7 +290,9 @@
the same counts.
- The class fix is proven against HEAD's `messageKind`. An uppercase class
in either preamble counts as agent after the fix and as human before it.
-- The JSON lists each seat's threads and the unmapped threads.
+- The JSON lists each seat's threads and the unmapped threads, and the
+ database path with whether it was the default. A `--t3-db` run prints the
+ path line in the text report, and a default run doesn't.
- A read against the live database gives the counts in this brief, allowing
for messages sent since. It exits 0 with no header conflict.
- The ledger README's counting rules name the new source, both flags, the
+306
View File
@@ -0,0 +1,306 @@
# Ledger: a read-only T3 thread source for Table 2 (Gate F brief)
Brief only, no code. Darkwing wrote it on 2026-09-26 at Sage's request, issue
#1506. R1 (sha256 08959a05) went to Filbert, whose review asked for
revisions: `agents/filbert/work/ledger-t3-source-review-2026-09-26.md`, sha256
19dda29a. This is R2. It takes every finding, and it records Sage's rulings
on the three open questions. Section 1 has one measurement that differs from
the review.
## Why
Table 2 counts user messages per seat from `.pi/state/<seat>/sessions/*.jsonl`
only. Development seats now run in T3 on the Claude and Codex harnesses, so
their prompts, Jason's included, never reach a Pi log. Today the Human column
can't see T3 at all, and the zero it shows for T3 seats means "no source", not
"no human prompts". 6a (ef0020ad) taught `messageKind` the T3 header, but no
source the ledger reads contains one. Gate F (QUEUE row 6) passes when
Filbert's item closes with zero human messages from Jason. While the ledger
can't see T3, a zero there proves nothing.
## Where T3 keeps messages
T3 keeps its state in one SQLite database, `~/.t3/userdata/state.sqlite`, in
WAL mode (`state.sqlite-wal` and `state.sqlite-shm` sit beside it). The counts
need three projection tables:
- `projection_projects`: `project_id`, `workspace_root`, `deleted_at`.
- `projection_threads`: `thread_id`, `project_id`, `title`, `archived_at`,
`deleted_at`.
- `projection_thread_messages`: `message_id` (primary key), `thread_id`,
`role` (`user` or `assistant`), `text`, `created_at` (ISO UTC).
The JSON diagnostic reads one more. In `orchestration_events`, each
`thread.message-sent` event carries `metadata_json.origin`. Messages typed in
the T3 app carry an `appVersion` there. Messages sent through T3's API or MCP
tools, which is how seats talk to each other, don't.
The same directory also holds `secrets/`, `clerk-tokens.json` and other
settings files. The reader opens `state.sqlite` and nothing else, and it
selects named columns only, never `*`.
## 1. Reading it, with T3 running or not
The reader uses Node's built-in `node:sqlite` (`DatabaseSync`). That needs no
dependency, and Node 26.8.1 (SQLite 3.53.4) prints no warning for it.
- **URI.** Build it with `pathToFileURL(dbPath)` and set `mode=ro` through
`searchParams`, then pass `readOnly: true`. A `?`, `#` or `%` in the home
path would break a string-built URI.
- **One snapshot.** Run every query, from the schema checks through the
diagnostic, inside one `BEGIN` … `COMMIT`. In autocommit mode each
statement sees its own snapshot while T3 writes between them.
- **Busy timeout.** Set `DatabaseSync`'s `timeout` to 5 s. A transient
`SQLITE_BUSY` during a T3 checkpoint then waits instead of failing. A busy
error after the timeout exits 1 like any open failure.
- **No `immutable=1` and no copy.** Both lose the WAL. Filbert found worse
than lost messages: with a table created inside the WAL, `immutable=1`
fails with `no such table`.
What happens on disk. Filbert and I both tested these in scratch
directories:
| State | Directory writable | Result |
|---|---|---|
| T3 running, writer attached, newest rows only in `-wal` | yes | reads them |
| `-wal` without `-shm` (writer killed, `-shm` removed) | yes | reads the WAL rows and creates `-shm` |
| `-wal` without `-shm` | no | open fails, SQLite 14 |
| T3 stopped cleanly, no `-wal` or `-shm` | yes | reads, then leaves an empty `-wal` and a 32 KiB `-shm` |
| T3 stopped cleanly | no | fails, SQLite 1544 "attempt to write a readonly database" |
In every case the main file's bytes stayed the same. The last row is where
Filbert and I differ. His review says the stopped-case read works with the
directory read-only. In my run it failed with and without the read
transaction. The build's test settles it. Either way a failed open is exit 1.
So the accurate claim: the reader never writes the main database file. Like
any SQLite connection, it may create or update `-wal` and `-shm` beside it
and takes read locks in `-shm`. T3 opens normally afterwards.
## 2. Jason or agent
Reuse the 6a rule. The first line of `text` decides: a T3 header or the tmux
preamble counts as agent, `control-board` as the sender counts as board, and
anything else counts as human. Messages with role `user` count; assistant
messages don't.
The class fix rides in this build (Sage's ruling). HEAD's
`packages/ledger/src/ledger.mjs:81` (tmux) and `:83` (T3) both allow only
`class=[a-z-]+`. Both become case-insensitive. Seats send uppercase classes:
Sage's DECISION, INFO, REVIEW-REQUEST and REVIEW-NOTE, and my own
REVIEW-REQUEST. In this project's threads 16 real agent headers failed on
that alone at 20:54Z. The ms-communications table lists lowercase names, so
the fix follows what seats send, not the table.
Cross-check, read at 2026-09-26T20:54Z for the mosaic-stack project (209
user messages outside imported and deleted threads, every one with its
`thread.message-sent` event). Filbert's later read agreed, plus messages sent
since.
| T3 origin | Header matches 6a | Count |
|---|---|---|
| typed in the app (has `appVersion`) | no | 99 |
| sent through the API (no `appVersion`) | yes | 80 |
| sent through the API | no, uppercase class | 16 |
| sent through the API | no, free-text roles | 14 |
No message typed in the app carries a header, and every API message in this
project carries one of the three forms. The 14 free-text ones are older
Discord Bot thread headers such as `[from: SetSpark coordinator (…) -> to:
Discord Bot (…)]`, written before the guide fixed the format. Sage ruled they
stay as recorded: they count as human, dated 2026-09-17 to 2026-09-22.
The header rule decides. The JSON also carries one diagnostic that feeds no
table or total: user messages the rule calls human that T3 recorded as sent
through the API. That number exposed the class bug and would catch the next
format drift. `origin` is T3's internal metadata, not a documented contract,
so it decides nothing. If `orchestration_events` or a column it needs is
missing, the diagnostic reads `unknown` and the report goes on (Sage's
ruling on F5). Missing tables the counts depend on still exit 1.
## 3. Thread to seat
**Project.** A thread counts for this checkout only if its project's
`workspace_root` equals the ledger's repository root, byte for byte. The CLI
already takes that root from the realpath of its own URL, today
`/mnt/storage/src/mosaic-stack`, project `34050c07`. So a T3 project opened
through the compatibility symlink `~/src/mosaic-stack-dev-test` doesn't
match, and "no project row" is the right refusal. The README says so.
**Title rule.** Thread IDs change whenever Jason starts a new thread for a
seat, so there's no fixed map. T3-AGENT-COMMS.md already names threads after
the seat ("Darkwing", "Sage", "Dewey in Claude"). A thread belongs to seat
`<s>` when `<s>` is a real directory under `agents/` and the lower-cased
title equals `<s>` or starts with `<s>` followed by a space. So "Sagebrush"
stays unmapped. Several threads can map to one seat, and their counts add
up, as several Pi session files already do.
Today that maps Sage, Darkwing, Filbert, Dewey and Rocko (one thread each,
created 2026-09-26, titles set by hand), plus "Darkwing in Claude" (archived)
and "Dewey in Claude". Researcher has a directory and no thread. Three
threads map to no seat. Two are imported and excluded anyway ("FINDINGS.md
review" and "[dragon-lin:darkwing -> …"). The third is "Discord Bot" with 68
user messages: 54 without a header, and the 14 free-text headers.
Titles are current state, and T3 can write them itself. They go wrong three
ways. T3 auto-titles an unnamed thread from Jason's first prompt, so "Rocko
review of the plan" maps to rocko. A rename moves the whole history to
another row. A seat thread titled for a topic drops into `t3:unmapped`.
None of this changes the Human total or the human-per-closed ratio. It only
moves counts between rows, but Gate F reads one seat's row.
**Header cross-check.** The headers already say which seat a thread belongs
to. For every user message whose header matches the fixed 6a rule and whose
`to:` id equals the message's own `thread_id`:
- in a mapped thread, the `to:` role, lower-cased, must equal that thread's
seat;
- in an unmapped thread, the `to:` role must not be a seat name.
A conflict exits 1 and names the thread id, its title and both roles. A
header whose `to:` id is some other thread is not checked. The check reads
message text only, not T3 metadata. In a live read at 21:02Z every header
agreed: all 104 addressed to their own thread carried the full thread id and
named that thread's seat (Sage 40, Darkwing 15, Filbert 18, Dewey 15, Rocko
16).
It catches a seat thread renamed to another seat or to a topic, once any
agent writes to it. It also catches an auto-titled thread that agents
address by a different seat. It misses a thread no agent ever writes to.
Such a thread can only add human counts to a seat's row, never hide them, so
for Gate F it errs toward a visible failure. The README says so.
**Unmapped row.** Unmapped threads go in one Table 2 row, `t3:unmapped`
(Sage's ruling), so their human messages still reach the Human column and
the human-per-closed ratio.
**Mapping in the JSON.** For each seat, the T3 thread ids and titles that
made its row, and the unmapped thread ids and titles. Anyone checking a Gate
F result can then see which threads the row came from.
A seat's row sums its Pi and T3 counts, and the JSON splits them by source.
Nothing is counted twice. Every T3 session today runs on `claudeAgent` or
`codex`, which don't write `.pi/state`, and Filbert found no T3 header in any
Pi log (6a record).
**Excluded,** with the reason stated in the README:
- Imported threads (`thread_id` starting `import:`, events marked
`historyImport`). They are partial copies of Claude Code sessions, not T3
traffic: 55 user messages in two threads here.
- Deleted threads (`deleted_at` set). Across all projects there are 3, with
3 messages. Archived threads count.
- Threads in other T3 projects. Live, there is a project at `/home/jwoltje`
and a deleted one at `/mnt/storage/src`. A thread in either could work on
this repository and would not be counted. The workspace-root rule is still
the right one, but the README names this blind spot.
## 4. What fails closed
The source is on by default (Sage's ruling). `--no-t3` turns it off, and the
report then says T3 was not read. `--t3-db <path>` reads another database
file instead of `~/.t3/userdata/state.sqlite`. It exists for fixtures and
gets the same checks.
Each of these refuses the report with exit 1, the code the ledger already
uses for unreadable session evidence. The report never falls back to Pi logs
alone. Where the database is missing or won't open, the message names
`--no-t3`.
- The database is missing or unreadable, or won't open read-only. That
includes a directory that isn't writable when SQLite needs to create
`-shm`, and a busy error after the timeout. The Pi reader treats a missing
`.pi` as no messages, and this departs from it on purpose. A missing Pi
directory means no Pi seats ran here. A missing T3 database on this host
means the path or T3 changed, and a silent zero is the failure Gate F
exists to prevent.
- `~/.t3`, `~/.t3/userdata` or `state.sqlite` is a symlink. With `--t3-db`,
the file and its directory are checked. The Pi reader checks every
ancestor too, but it skips symlinked entries. Skipping one named file
would be another silent zero, so this reader refuses.
- A table or column the counts need is missing. The reader checks `PRAGMA
table_info` and names what's missing. This catches a T3 upgrade that
changes the schema.
- No project row, or more than one non-deleted row, for this repository root.
- A counted row has a bad `role`, non-string `text`, or a `created_at` that
doesn't parse. The Pi reader already refuses malformed JSONL and bad
timestamps the same way.
- A header conflicts with the title mapping (section 3).
The reader never reads other files in `~/.t3`. It passes no message text
beyond `messageKind`, `issueNumbers` and the header's `to:` role and id, the
same rule as for Pi logs.
## 5. Rulings
Sage ruled on the three questions R1 put to Jason, as lead calls:
1. The source is on by default. A missing or unreadable database exits 1,
and the message names `--no-t3`.
2. Unmapped threads get the `t3:unmapped` row.
3. The 14 free-text headers stay as recorded. They show only in the JSON
diagnostic.
Sage also ruled that the class fix rides in this build, and that a missing
diagnostic table reads `unknown` (F5).
## 6. Acceptance for the build
**No test opens the real `~/.t3`.** Both places in
`packages/ledger/tests/ledger.test.mjs` that spawn `cli.mjs` (the shared
`run()` helper and the direct `spawnSync` at line 66) set `HOME` to the
fixture's temp directory. A test that forgets `--t3-db` or `--no-t3` then
finds no database and fails closed. The existing tests aren't about T3. Each
gets an empty fixture database at the fixture `HOME`'s default path, with
one project row for the fixture root. So they run with the source on, and
their expected rows don't change. One test asserts that a `HOME` with no
database exits 1 and names `--no-t3`.
Fixture databases are built with `node:sqlite` in a temp directory, in WAL
mode:
- seat threads and an unmapped thread; imported, deleted and archived
threads; a message outside the date range;
- all three header forms, with uppercase classes in both the tmux preamble
and the T3 header;
- a thread with the same seat title in another project;
- a seat thread renamed to another seat, with an agent header to its own
id, which exits 1;
- a "Sagebrush" title, which stays unmapped;
- a thread titled "Researcher", which maps to the seat that has no thread
live.
WAL states, each with its own test:
- The newest message is only in `-wal`, with the writer still attached (the
live-T3 case). It is counted.
- T3 stopped: the database closed cleanly with no writer. Counts are
correct, and the main file's bytes are unchanged afterwards.
- `-wal` without `-shm` in a writable directory: made by a child writer with
`wal_autocheckpoint=0` that is SIGKILLed, then `-shm` deleted. The WAL
rows are counted.
- `-wal` without `-shm` in a directory that isn't writable: exit 1, naming
`--no-t3`. Skipped when the tests run as root, where the mode bits don't
bind.
- The stopped case in a directory that isn't writable: exit 1, naming
`--no-t3`, as I measured it. If the build reads there instead, the builder
changes this test to assert correct counts and records the correction in
the BUILD-LOG entry. Skipped as root too.
Also:
- Each other fail-closed case in section 4 has its own test, including a
removed schema column and each symlink.
- A missing `orchestration_events` gives `unknown` for the diagnostic and
the same counts.
- The class fix is proven against HEAD's `messageKind`. An uppercase class
in either preamble counts as agent after the fix and as human before it.
- The JSON lists each seat's threads and the unmapped threads.
- A read against the live database gives the counts in this brief, allowing
for messages sent since. It exits 0 with no header conflict.
- The ledger README's counting rules name the new source, both flags, the
mapping rule and the header check, and the exclusions. That includes the
symlinked-checkout case and the other-project blind spot.
- No suite runs the ledger tests, so the BUILD-LOG entry names the test file.
## Not in scope
- Claude Code transcripts (`~/.claude/projects`) and Codex sessions
(`~/.codex/sessions`). T3's database already holds every message T3
delivered, so those files would only duplicate it.
- Any T3 API call, anything that needs T3 running, and any write beyond
SQLite's own `-wal` and `-shm` handling.
- Fixing the two tmux misclassifications Filbert found in 6a.
@@ -0,0 +1,78 @@
# Queue row 33, round 1 review (#1508)
Darkwing, 2026-10-04. Request: #1508 comment 26648. Brief:
`docs/plans/2026-10-04_queue-follow-ups.md`. Candidate:
`agents/filbert/work/queue-33/candidate-manifest.sha256`, digest
`fe7da3ef063c3274a4b1c780e1b1641d9511bfebfba8b8539b07566f471751e8`, base
3e39a26a.
Verdict: approve.
## Checks
- The manifest hashes to fe7da3ef. All 10 listed files match it in the
canonical working tree. Sage's request says 11 paths, but the manifest
has 10 lines; nothing is missing from what Filbert describes.
- Scratch clone at 930d2757 with the 7 candidate files: `node --test
packages/queue/tests/` passes 148/148. `bash scripts/test-queue.sh`
gives 27/0; the live checks skip there, as in an export.
- Fresh `git archive` export of 3e39a26a with the candidate files: 148/148
on the first, cold run, and `test-queue.sh` 27/0.
- `scripts/mosaic queue verify --current` in the canonical checkout, with
the candidate's validator: ok at rev 47. The live log replays under the
calendar check.
## The five items
1. Genesis. `parseMigrationMap` has one caller, `genesis` in store.mjs,
and it runs after the retry check. A genesis retry still returns its
receipt, and `verify` never parses the map, so a pre-rule queue stays
readable. C1 is right.
2. Assign wording matches the brief.
3. HEAD moved. `head_moved` runs before each canary and again when the
clean run fails. The step-7 check now catches a move that update-ref
used to catch. Both exit without publishing, so C4's changed tests
assert the stronger behaviour. C2 ("another commit landed") is right:
the check fires for any commit. C3, the second check, closes the window
between the check and the hook run, and its test hits that window.
4. Calendar. I parsed every hour 00 to 99 with minutes and seconds 0, 59,
60 and 99 in V8. Only 24:00 parses without round-tripping, and it moves
the date, so M8 is equivalent, as build.md says. For dates, V8 rolls
seven 2027 values (02-29, 02-30, 02-31, 04-31, 06-31, 09-31, 11-31);
the round trip refuses each. Year 0000 and 9999-12-31T23:59:59.999Z
round-trip and pass. Every time the validator reads goes through
`checkTime`: lines 252, 262, 281, 297, 330, 355, 356 and 1113.
5. Cold runs: 20 recorded, all 148/148. The brief's rule closes the item.
## My mutants
Ten, on the candidate. Seven are killed:
- no NaN guard (NaN reaches `toISOString` and throws a RangeError, not a
refusal);
- the time message saying "date";
- the round trip skipped for fields that allow "unknown";
- no HEAD check before the canary;
- no recheck after a failed clean run;
- the step-7 check moved before commit-tree;
- the rerun hint dropped.
Three survive:
- `Date.parse(v)` on the bare date in place of the `T00:00:00.000Z`
suffix. Equivalent: a bare ISO date parses as UTC midnight.
- The genesis rule checking only `reviewers[0]`. Every new test puts the
owner first. The code uses `includes`, so this isn't a defect.
- `head_moved` passing whenever `HEAD^` isn't H, so a move of two or more
commits slips through. Every test moves HEAD by one commit. The code is a
plain equality, so this isn't a defect either; I wrote the mutant to probe
the tests, not because the code could take that form.
## Non-blocking
- n1. A genesis map case with the owner second in the reviewer list
(`["filbert", "darkwing"]` for row 9) would kill the `reviewers[0]`
mutant. Optional.
- n2. The 20 cold runs ran one at a time on an idle machine. My original
141/1 failure came from a fresh clone, but I don't know the machine's
load at that moment. If it was load-dependent, idle runs wouldn't show
it. The brief's rule is met, and I agree with closing the item. If the
failure comes back, capture the test name and the load.
@@ -0,0 +1,20 @@
76833a3bf536bb9592a0040cdb10d9a8351cf0b3bd2828feee928d3415136a79 docs/plans/BRIEF-TEMPLATE.md
5d4b4c7a4624ef267d76d4c032dbcddf3a3d7e7b73a5cfed06953307992bbde7 packages/queue/package.json
576d8ed44a19e1cca96fd7128ca34fcc62d840580b2c0936fa3b296df2e72ddf packages/queue/README.md
188ade96cabb73e06b6b8fbe3d30e8d4d174843877f3a151dcf08082c92d3068 packages/queue/src/cli.mjs
7a851814dfff6f392de814fc31f8dc8cbe9c79cd40ef71313dca1939115ee879 packages/queue/src/errors.mjs
b18e120cb9ddba4c5576d7bc7f7f378ed86e6ec1e084d436474b80765261d2e5 packages/queue/src/io.mjs
52d9f68f01f29e84943fc359fdb1d1ddfaf58d1650c6b15b253b83f1daba9927 packages/queue/src/lock.mjs
c11235a6b99acf6baf1257c63eced410060065ef4f21f79a18860adfa19571cf packages/queue/src/queue.mjs
756cbc9ab13de85757cc24f903d02e8a1d20bb45bb8020c5e1d91e26fbaacaa4 packages/queue/src/store.mjs
6005da4c809cb9045f9480e1e29077b8e7ab185c3ed91567717e8cc13d67b5b0 packages/queue/tests/commit.test.mjs
d29d58427c712a69e8a818d8c74ce724d519ddea8780cc0b5f0b0035cec0f498 packages/queue/tests/data.test.mjs
5cccea50d5a40e07891a090dea001c095a26f4a6c2e0af40c144b1d00f239b5e packages/queue/tests/fixtures/kill-at.mjs
59cc8092fbcddbe9854da7d5014f706f4ffb86573f62ad909aa0b0e09c6d99b9 packages/queue/tests/fixtures/lock-child.mjs
5769b3618918fe36398a75e449d644932332ad5a60c3127098a1d011eda2c182 packages/queue/tests/helpers.mjs
9f98a388ce91438c3238be36049bcd5171b365c68d4b7bf1ae5ff910f4b7b1e8 packages/queue/tests/lock.test.mjs
2a3d2be8cb25b6e7cd18ba56393a284415c66e7ee26a3148fa39f32885efbd35 packages/queue/tests/store.test.mjs
74378acbd41ef21a0b171b08aa85677e4c471966d2ffd8cfb310adb0e044b246 packages/queue/tests/write.test.mjs
3cbd40575dc728dc5407c5029f4f5fff747ce93a5508807233f4362ad37d7f9c scripts/git-hooks/pre-commit
2632078bea45e0249b3fdd9a335100bade7c9106222931603ede9d414c404f54 scripts/queue-commit.sh
92cea23b9ada2edb1b0482ca2daf5864666cc1f546e1f2e6558c3826f1ddf9e7 scripts/test-queue.sh
@@ -0,0 +1,20 @@
20363f5dafbb1be8b7380d7603fd04cf38f5284457b634c9a98ce5a6d8e4832a docs/plans/BRIEF-TEMPLATE.md
5d4b4c7a4624ef267d76d4c032dbcddf3a3d7e7b73a5cfed06953307992bbde7 packages/queue/package.json
9ebdb6a3f239051a39e63fcf8f59c8fba540bb3cc63b7880d15e6c6f1e549a09 packages/queue/README.md
711db25594d78e0ba603a9e91221f6c32241ce1d9217b01bd4da3a99967ae871 packages/queue/src/cli.mjs
7a851814dfff6f392de814fc31f8dc8cbe9c79cd40ef71313dca1939115ee879 packages/queue/src/errors.mjs
b18e120cb9ddba4c5576d7bc7f7f378ed86e6ec1e084d436474b80765261d2e5 packages/queue/src/io.mjs
1095cb6611f2d8d38f535930bbe03208e854b4a57176cea813f7ea0487b3c4f1 packages/queue/src/lock.mjs
8e9230901f550b829ef55e754819c9cd5703e98e447fcebdba6ce5dbef11b0e9 packages/queue/src/queue.mjs
172cf529b2c0e915fbd4a130faa5dbec9023bb19160012246801b16482b4ffaa packages/queue/src/store.mjs
74d04d0de9f4e068fe66bb465ed1575fb38b6ea6592fca6cd6fbacdd553fe9e1 packages/queue/tests/commit.test.mjs
e3f774f823bfb1bcb6025cc3688d72ec110144d5df064557d1645a62472fd663 packages/queue/tests/data.test.mjs
5cccea50d5a40e07891a090dea001c095a26f4a6c2e0af40c144b1d00f239b5e packages/queue/tests/fixtures/kill-at.mjs
59cc8092fbcddbe9854da7d5014f706f4ffb86573f62ad909aa0b0e09c6d99b9 packages/queue/tests/fixtures/lock-child.mjs
5769b3618918fe36398a75e449d644932332ad5a60c3127098a1d011eda2c182 packages/queue/tests/helpers.mjs
a2dbe3dc69b9d53c42246e41e61b9b7d2395697a53ca12ef3481965b43321ff6 packages/queue/tests/lock.test.mjs
34f4b0e3eeadc882ffcc6ba7f0b439651957f2f0c41c9387cf8fb51901f46202 packages/queue/tests/store.test.mjs
70e8a068efda8da8fe1e1628cc1cd5b7fa7796a475941bf7349747244c002f1b packages/queue/tests/write.test.mjs
3cbd40575dc728dc5407c5029f4f5fff747ce93a5508807233f4362ad37d7f9c scripts/git-hooks/pre-commit
2632078bea45e0249b3fdd9a335100bade7c9106222931603ede9d414c404f54 scripts/queue-commit.sh
92cea23b9ada2edb1b0482ca2daf5864666cc1f546e1f2e6558c3826f1ddf9e7 scripts/test-queue.sh
+186
View File
@@ -0,0 +1,186 @@
# Queue A1 build (#1508), candidate for review
Darkwing built this on 2026-09-26 from section 8 of
`agents/filbert/work/queue-as-data-plan-2026-09-26.md` (sha256 282fabbb,
the only spec), split as Sage approved: render is in A1, `move in-review`
needs `--candidate` until piece D, and a round's issue is the row's first
issue. Filbert reviews the code; Sage commits after the suites. Base is HEAD
3a209eea. Nothing is committed, staged or pushed.
A stray pkill at 22:02:32Z stopped my first turn with only
`src/errors.mjs` and `src/io.mjs` on disk. I reread both against what I had
meant them to be. They match: the exit-code class, and the file layer with
`realIo` as the only layer the CLI uses. Everything else was written after
the restart.
## Files
`build-manifest.sha256` pins the 20 files. `build.patch` (sha256 419804f2)
adds all 20 as new files with their modes. It applies cleanly to 3a209eea,
and the applied tree matches the manifest and passes `scripts/test-queue.sh`.
- `packages/queue/src/`: `errors.mjs`, `io.mjs` (the fault-injectable file
layer), `lock.mjs` (lock and unlock gate, 8.4), `queue.mjs` (serialization,
replay, the transition matrix, `next`, render), `store.mjs` (canonical
checks, the write path, witness, views, snapshot, verify), `cli.mjs`.
- `packages/queue/tests/`: data 19, lock 17, store 18, write 20, commit 21
tests, plus `helpers.mjs` and two child fixtures.
- `packages/queue/package.json` and `README.md`. The package has no
dependencies.
- `scripts/queue-commit.sh` (0755): the 8.12 procedure and
`--install-hook`.
- `scripts/git-hooks/pre-commit` (0755, POSIX sh): the queue guard.
- `scripts/test-queue.sh` (0755): the suite, in the style of
`test-discord.sh`.
- `docs/plans/BRIEF-TEMPLATE.md`: the 8.13 template.
## Which path runs `verify` once genesis is in
`scripts/test-queue.sh` runs `node packages/queue/src/cli.mjs verify` when
`git cat-file -e HEAD:docs/plans/queue.json` succeeds. At HEAD today there is
no `queue.json`, so it prints `skip queue verify: HEAD has no
docs/plans/queue.json (before the genesis commit)` and stays green. A2 moves
that call to `scripts/mosaic queue verify` when it adds the dispatch.
`queue-commit.sh` also calls `node packages/queue/src/cli.mjs` directly
(`snapshot`, then `verify --snapshot` from HEAD's archive) until A2.
## Sage's five conditions
1. Nothing ran against the canonical `.git`. After all runs, `.git/hooks`
holds only the samples, `.git` has no `mosaic-queue*` file, and
`git config --show-scope --get-all core.hooksPath` returns nothing in any
scope (rc 1). Every hook install, genesis, lock and gate test runs in a
scratch repo under the system temp directory. Suite runs used a
`--shared` clone at `/tmp/qa1-verify`.
2. `docs/plans/QUEUE.md`, `AGENTS.md` and `docs/TOOLS.md` are unedited.
`docs/SESSIONS.md` shows as modified in the working tree, but that was
someone else's edit before my session began; I didn't touch it.
3. `scripts/test-queue.sh` is green at HEAD with no `queue.json`: 19 checks
passed, `verify` skipped as above.
4. H is recorded before the canary. `commit.test.mjs` has "H recorded before
the canary": a shim commits on the first `git hook run`, and
`queue-commit.sh` exits 1 with `refs/heads/<branch> moved since <H>;
nothing published`. A second test moves HEAD after `commit-tree` with the
same result. Mutation M1 (read H after the canary) fails the first test.
5. The fault file layer is reachable only from tests. Faults enter through
the options the API takes (`io`, `proc`, `hook`, `now`, `readOrder`,
`lockWaitMs`); `cli.mjs` passes none. The only `process.env` read in
`src/` is the default `env` in `store.mjs`'s context.
## Bugs found while building
- A nested `node --test` inherits `NODE_TEST_CONTEXT` and exits 0 whatever
its tests do. Step 4's run of HEAD's archived tests therefore passed with a
failing test in the archive. `queue-commit.sh` and `test-queue.sh` now run
it under `env -u NODE_TEST_CONTEXT`, and a test commits an archive with a
failing test and expects a refusal (mutation M4). Other suites in this repo
that nest `node --test` may have the same blind spot. I haven't checked
them.
- git 2.55 does not hold `index.lock` while the commit editor is open. The
plan expected a paused `git commit -e` to block step 8. It doesn't: the
paused commit loses later at its own HEAD update with `cannot lock ref
'HEAD': is at C but expected H`. The test now asserts that outcome.
Nothing is lost, but the reason differs from the plan's.
- ext4 hands a freed inode number straight back. My first test for release's
inode check wrote a byte-identical lock after unlinking the original and
got the same inode back, so it proved nothing. It now writes a copy and
renames it over the lock, which guarantees a new inode.
## Choices the spec left open
- Verb names `release` and `set`. `add` requires `--gate`. A null brief is
allowed only on rows that genesis creates as done. `sync --op` is
optional.
- Reads need no actor. `next` with no seat and no `$MOSAIC_AGENT_NAME`
refuses.
- `accept-history` accepts a stale table but not an unknown one. When a
write succeeds but the table write is skipped, the CLI warns and exits 0.
- An invalid witness is treated as absent.
- `GIT_DIR`, `GIT_WORK_TREE` and `GIT_COMMON_DIR` refuse in the CLI.
`queue-commit.sh` also refuses `GIT_INDEX_FILE`, `GIT_OBJECT_DIRECTORY`
and `GIT_ALTERNATE_OBJECT_DIRECTORIES`.
- Leftover temp files are unlinked. Genesis uses `link`, so it cannot
replace an existing file.
- `unlock` works on a missing or invalid lock file.
- `note` on a blocked row edits `blockedReason`.
- Messages already name `scripts/mosaic queue`. The lost-history refusal
names `sync` when the file holds genesis alone.
- `--candidate` auto-detects: an existing file is a manifest, anything else
is a commit reachable from `refs/heads` or `refs/tags`.
- Only `--install-hook` is privileged in `queue-commit.sh` (jason or sage).
The commit itself relies on the protocol that the lead runs it.
- After the snapshot, `queue-commit.sh` checks the genesis entry's branch
and root against the current branch and root. For `--genesis` it also
checks that `H:<map>` is the map blob genesis read.
- Step 8 compares the index's two queue entries with H's before it looks
for `index.lock`, so "someone staged a queue path" is reported ahead of a
lock.
## Deferred
- 8.12's test of `verify-commit` on a prospective tree belongs to piece D,
which adds `verify-commit`. It is not in A1.
- A2 holds the real migration map, the QUEUE.md markers and header, the
row-7 pointer, the golden render, `scripts/mosaic` dispatch and
`docs/TOOLS.md`.
- Sage adds `queue` to the suite list when A1 lands (lead decision 20).
- Bootstrap happens after A2's map and markers land:
`scripts/queue-commit.sh --install-hook --by sage`, then `queue genesis`,
then `scripts/queue-commit.sh --genesis -m MSG`.
## Verification
In `/tmp/qa1-verify` (HEAD 3a209eea plus the 20 files):
| Suite | Result |
|---|---|
| config | 24/24 |
| task | 90/90 |
| foundation | 43/43 |
| conductor | 17/17 |
| release | 14/14 |
| auth | 15/15 |
| discord | 63/63 |
| extension-package | 18/18 |
| queue | 19 checks; `node --test` 95/95 |
The queue tests take about 18 s and were stable over two runs. A combined
`node --test` run over every package came to 474/474.
I also ran the post-genesis path in a scratch repo: install the hook,
genesis, `--genesis` commit, then `test-queue.sh`. `verify` printed `ok
verify rev 1: file valid, witness matches, view current`. After a hand edit
to one table cell it failed with `view unknown`.
### Mutations
Each mutation went into the verify clone, the queue tests ran, and the
original was restored. Every one was caught; the number is how many tests
failed.
| Id | Mutation | Failing tests |
|---|---|---|
| M1 | read H after the canary | 1 |
| M2 | drop the step-7 guard recheck | 2 |
| M3 | drop step 8's entry comparison | 1 |
| M4 | keep `NODE_TEST_CONTEXT` | 1 |
| M5 | drop step 1's staged-path check | 1 |
| M6 | `update-ref` without the old value | 2 |
| M7 | skip the canary | 2 |
| M8 | guard hook always passes | 21 |
| M9 | drop step 8's `index.lock` check | 1 |
| M10 | drop the exec-bit check | 2 |
| M11 | drop the `core.hooksPath` check | 1 |
| M12 | drop the map-blob check | 1 |
| S1 | drop the "unchanged since read" check | 1 |
| S2 | swallow the directory fsync error | 2 |
| S3 | drop the recheck under the lock for unlocked reads | 2 |
| S4 | drop the unlock-gate check | 2 |
| S5a | release ignores the inode | 1 |
| S5b | release ignores the record bytes | 1 |
| S6 | write the witness before the rename | 10 |
| S7 | drop genesis's fsync | 1 |
| S8 | treat a reused pid as dead | 11 |
S5 survived at first: the delayed-release test was caught by the byte
comparison alone. The inode test in `lock.test.mjs` closes that gap.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,819 @@
diff --git a/docs/plans/BRIEF-TEMPLATE.md b/docs/plans/BRIEF-TEMPLATE.md
index c5ece4ba..1b0facef 100644
--- a/docs/plans/BRIEF-TEMPLATE.md
+++ b/docs/plans/BRIEF-TEMPLATE.md
@@ -13,7 +13,8 @@ Rules the queue enforces (queue-as-data plan 8.13):
refuse until the lead re-pins it.
`queued` means the brief exists, not that it is accepted. The row moves to
-`briefed` when its owner accepts it.
+`briefed` when a privileged actor (jason or sage) accepts it; the owner
+can't.
---
diff --git a/packages/queue/README.md b/packages/queue/README.md
index 74956816..1c8774fa 100644
--- a/packages/queue/README.md
+++ b/packages/queue/README.md
@@ -44,7 +44,18 @@ Exit codes: 0 ok; 1 the operation failed; 2 invalid data or refused;
Until piece D, `move ID in-review` needs `--candidate`: an existing file is
read as a manifest (one `<sha256> <path>` line per file), anything else as a
commit reachable from `refs/heads` or `refs/tags`. The candidate is frozen
-for the round. The review's issue is the row's first issue.
+for the round.
+
+The review's issue follows lead decision 23. A row with no issues can't
+request review. A row with one issue uses it. A row with several needs
+`--issue N`, one of its issues. Later rounds keep the previous round's issue
+unless `--issue` names another; if the row no longer lists the kept issue,
+the request refuses until `--issue` names one.
+
+`move ID done` from in-review needs `--evidence
+comment=<id>,round=<n>,candidate=<digest>`. The round must be the current
+one and the digest its candidate's, so a comment from an earlier round
+can't close a later one, even when the candidate is the same.
## Where the files live
diff --git a/packages/queue/src/cli.mjs b/packages/queue/src/cli.mjs
index c5517747..e8df80f7 100644
--- a/packages/queue/src/cli.mjs
+++ b/packages/queue/src/cli.mjs
@@ -4,7 +4,7 @@
// Reads: list | show ID | next [SEAT]
// Changes: add --piece TEXT --gate TEXT --brief PATH#ANCHOR [--issue N]... [--note TEXT]
// [--owner SEAT] [--gate-owner SEAT] [--after ID[:settled]]... [--reviewer SEAT]... [--required]
-// move ID STATE [--reason TEXT] [--candidate COMMIT|MANIFEST] [--evidence TEXT]
+// move ID STATE [--reason TEXT] [--candidate COMMIT|MANIFEST] [--issue N] [--evidence TEXT]
// release ID | assign ID SEAT | note ID TEXT | set ID FIELD VALUE [--reason TEXT]
// genesis --root PATH --branch NAME --map PATH
// accept-history --reason TEXT --yes
@@ -23,7 +23,7 @@ import { list, mutate, next, renderView, show, snapshot, sync, unlock, verify, v
const USAGE = [
"usage: queue list | show ID | next [SEAT]",
" queue add --op ID --piece TEXT --gate TEXT --brief PATH#ANCHOR [--issue N]... [--note TEXT] [--owner SEAT] [--gate-owner SEAT] [--after ID[:settled]]... [--reviewer SEAT]... [--required]",
- " queue move ID STATE --op ID [--reason TEXT] [--candidate COMMIT|MANIFEST] [--evidence TEXT]",
+ " queue move ID STATE --op ID [--reason TEXT] [--candidate COMMIT|MANIFEST] [--issue N] [--evidence TEXT]",
" queue release ID --op ID | assign ID SEAT --op ID | note ID TEXT --op ID",
` queue set ID FIELD VALUE --op ID [--reason TEXT] (fields: ${SET_FIELDS.join(", ")})`,
" queue genesis --op ID --root PATH --branch NAME --map PATH",
@@ -134,8 +134,12 @@ export function run(argv, opts = {}) {
});
}
case "move":
- allow(flags, [...CHANGE, "--reason", "--candidate", "--evidence"]); positional(pos, 2, "move ID STATE");
- return change("move", { id: intArg(pos[0], "ID"), to: pos[1], reason: f("--reason"), candidate: f("--candidate"), evidence: f("--evidence") });
+ allow(flags, [...CHANGE, "--reason", "--candidate", "--issue", "--evidence"]); positional(pos, 2, "move ID STATE");
+ if ((flags.get("--issue") ?? []).length > 1) throw usage("move takes one --issue");
+ return change("move", {
+ id: intArg(pos[0], "ID"), to: pos[1], reason: f("--reason"), candidate: f("--candidate"), evidence: f("--evidence"),
+ issue: flags.has("--issue") ? intArg(flags.get("--issue")[0].replace(/^#/, ""), "--issue") : null,
+ });
case "release":
allow(flags, CHANGE); positional(pos, 1, "release ID");
return change("release", { id: intArg(pos[0], "ID") });
diff --git a/packages/queue/src/lock.mjs b/packages/queue/src/lock.mjs
index 56c43f0f..34e88e3b 100644
--- a/packages/queue/src/lock.mjs
+++ b/packages/queue/src/lock.mjs
@@ -70,6 +70,7 @@ function describe(c) {
function publish(io, target, bytes, hook, waitMs, stepMs) {
const tmp = `${target}.${process.pid}.${randomBytes(6).toString("hex")}.tmp`;
let fd;
+ let st;
try {
fd = io.openExcl(tmp, 0o600);
} catch (err) {
@@ -82,6 +83,9 @@ function publish(io, target, bytes, hook, waitMs, stepMs) {
fd = null;
const back = io.readFile(tmp);
if (!back.equals(bytes)) throw Object.assign(new Error("read-back differs from the record"), { code: "EREADBACK" });
+ // The link gives the target this inode, so read it before linking:
+ // nothing that can fail runs between a successful link and the return.
+ st = io.stat(tmp);
} catch (err) {
if (fd !== null) { try { io.close(fd); } catch { /* already failing */ } }
unlinkQuiet(io, tmp);
@@ -99,7 +103,6 @@ function publish(io, target, bytes, hook, waitMs, stepMs) {
sleepMs(stepMs);
continue;
}
- const st = io.stat(tmp);
return { linked: true, dev: st.dev, ino: st.ino, bytes };
}
} finally {
@@ -122,8 +125,14 @@ export function acquire({ gitDir, io, proc = realProc, op = null, verb, waitMs =
const handle = { path, dev: got.dev, ino: got.ino, bytes: got.bytes };
hook("lock-linked");
const gate = join(gitDir, GATE_NAME);
- if (lstatOrNull(io, gate) !== null) {
- const c = classify(readOrNull(io, gate), proc);
+ let c = null;
+ try {
+ if (lstatOrNull(io, gate) !== null) c = classify(readOrNull(io, gate), proc);
+ } catch (err) {
+ const left = release(handle, io);
+ throw new QueueError(`cannot check the unlock gate ${gate}: ${errno(err)}; ${left ?? "lock released"}`, 1);
+ }
+ if (c !== null) {
release(handle, io);
throw new QueueError(`unlock gate ${gate} is present (${describe(c)}); check it with \`scripts/mosaic queue unlock --check-gate\``, 2);
}
@@ -154,18 +163,29 @@ export function unlock({ gitDir, io, proc = realProc, hook = () => {} }) {
}
const gate = { path: gatePath, dev: got.dev, ino: got.ino, bytes };
hook("gate-held");
+ let result;
+ let failure = null;
try {
const lockBytes = readOrNull(io, lockPath);
- if (lockBytes === null) return "no queue lock present; nothing removed";
- const c = classify(lockBytes, proc);
- if (c.state !== "dead" && c.state !== "mismatch") {
- throw new QueueError(`queue lock owner is ${describe(c)}; unlock refuses`, 2);
+ if (lockBytes === null) {
+ result = "no queue lock present; nothing removed";
+ } else {
+ const c = classify(lockBytes, proc);
+ if (c.state !== "dead" && c.state !== "mismatch") {
+ throw new QueueError(`queue lock owner is ${describe(c)}; unlock refuses`, 2);
+ }
+ io.unlink(lockPath);
+ result = `removed queue lock (${describe(c)}): ${lockBytes.toString("utf8").trim()}`;
}
- io.unlink(lockPath);
- return `removed queue lock (${describe(c)}): ${lockBytes.toString("utf8").trim()}`;
- } finally {
- release(gate, io);
+ } catch (err) {
+ failure = err;
}
+ let msg;
+ try { msg = release(gate, io); } catch (err) { msg = `cannot release the unlock gate (${errno(err)})`; }
+ // Like the lock, a swapped gate is reported on success and on refusal (8.4).
+ if (msg && failure instanceof Error) failure.message += `\nwarning: ${msg}`;
+ if (failure) throw failure;
+ return msg ? `${result}\nwarning: ${msg}` : result;
}
export function checkGate({ gitDir, io, proc = realProc }) {
diff --git a/packages/queue/src/queue.mjs b/packages/queue/src/queue.mjs
index 9999ba18..06e77c17 100644
--- a/packages/queue/src/queue.mjs
+++ b/packages/queue/src/queue.mjs
@@ -208,7 +208,7 @@ export function validateRow(row) {
checkNames(row.reviewers, `${w} reviewers`);
if (row.review !== null) {
keysExactly(row.review, ["issue", "rounds"], `${w} review`);
- if (row.review.issue !== null) checkId(row.review.issue, `${w} review issue`);
+ checkId(row.review.issue, `${w} review issue`);
if (!Array.isArray(row.review.rounds) || row.review.rounds.length === 0) throw refuse(`${w} review needs at least one round`);
row.review.rounds.forEach((r, i) => checkRound(r, i + 1));
}
@@ -338,6 +338,7 @@ export function canonArgs(verb, a) {
reason: nullable(a.reason, (v) => checkText(v, "reason")),
candidate: nullable(a.candidate, (v) => checkText(v, "candidate", { max: 300 })),
evidence: nullable(a.evidence, (v) => checkText(v, "evidence")),
+ issue: nullable(a.issue, (v) => checkId(v, "issue")),
};
case "release":
return { id: checkId(a.id) };
@@ -393,11 +394,30 @@ function afterSatisfied(rows, row) {
return missing;
}
-// `comment=<id>,candidate=<digest>`: the J5 evidence before Piece D.
+// `comment=<id>,round=<n>,candidate=<digest>`: the J5 evidence before Piece D.
export function parseReviewEvidence(text) {
- const m = /^comment=([1-9][0-9]{0,19}),candidate=([0-9a-f]{40}|[0-9a-f]{64})$/.exec(text ?? "");
- if (!m) throw refuse("in-review to done needs --evidence comment=<id>,candidate=<digest> for the current round");
- return { comment: m[1], candidate: m[2] };
+ const m = /^comment=([1-9][0-9]{0,19}),round=([1-9][0-9]{0,5}),candidate=([0-9a-f]{40}|[0-9a-f]{64})$/.exec(text ?? "");
+ if (!m) throw refuse("in-review to done needs --evidence comment=<id>,round=<n>,candidate=<digest> for the current round");
+ return { comment: m[1], round: Number(m[2]), candidate: m[3] };
+}
+
+// The issue a review round posts to (lead decision 23). The row must list
+// one; with several, --issue names it. A later round keeps the previous
+// round's issue unless --issue names another, and the kept issue must still
+// be one of the row's.
+function reviewIssue(row, issue) {
+ const list = row.issues.map((n) => `#${n}`).join(", ");
+ if (row.issues.length === 0) throw refuse(`row ${row.id} lists no issues; a privileged actor sets one before review`);
+ if (issue !== null) {
+ if (!row.issues.includes(issue)) throw refuse(`--issue #${issue} is not one of row ${row.id}'s issues (${list})`);
+ return issue;
+ }
+ if (row.review) {
+ if (!row.issues.includes(row.review.issue)) throw refuse(`row ${row.id}'s review issue #${row.review.issue} is no longer one of its issues (${list}); name one with --issue`);
+ return row.review.issue;
+ }
+ if (row.issues.length > 1) throw refuse(`row ${row.id} lists several issues (${list}); name the review's issue with --issue`);
+ return row.issues[0];
}
function touch(row, entry) {
@@ -424,7 +444,7 @@ function getRow(rows, id) {
}
function applyMove(rows, row, entry, resolved) {
- const { to, reason, candidate, evidence } = entry.args;
+ const { to, reason, candidate, evidence, issue } = entry.args;
const by = entry.by;
const from = row.state;
const illegal = () => refuse(`row ${row.id}: ${from}→${to} is not a transition`);
@@ -432,9 +452,11 @@ function applyMove(rows, row, entry, resolved) {
if (reason !== null && to !== "blocked") throw refuse("--reason applies only to a move to blocked");
if (candidate !== null && !(from === "in-progress" && to === "in-review")) throw refuse("--candidate applies only to in-progress→in-review");
if (evidence !== null && to !== "done") throw refuse("--evidence applies only to a move to done");
+ if (issue !== null && !(from === "in-progress" && to === "in-review")) throw refuse("--issue applies only to in-progress→in-review");
let next = { ...row, state: to };
let round = null;
let cand = null;
+ let revIssue = null;
if (to === "blocked") {
if (from === "blocked") throw refuse(`row ${row.id} is already blocked; update the reason with note`);
if (!NON_TERMINAL.has(from)) throw illegal();
@@ -457,11 +479,12 @@ function applyMove(rows, row, entry, resolved) {
} else if (from === "in-progress" && to === "in-review") {
if (row.claim === null || by !== row.claim.seat) throw refuse(`only the claimant (${row.claim?.seat ?? "nobody"}) may request review of row ${row.id}`);
if (candidate === null) throw refuse("in-progress→in-review needs --candidate <commit|manifest>");
+ revIssue = reviewIssue(row, issue);
cand = checkCandidate(resolved.candidate);
const rounds = row.review ? row.review.rounds : [];
round = rounds.length + 1;
next.review = {
- issue: row.review ? row.review.issue : (row.issues[0] ?? null),
+ issue: revIssue,
rounds: [...rounds, { n: round, op: entry.op, by, at: entry.at, candidate: cand, request: "none" }],
};
} else if (from === "in-review" && (to === "in-progress" || to === "waiting-on-jason")) {
@@ -479,6 +502,7 @@ function applyMove(rows, row, entry, resolved) {
const ev = parseReviewEvidence(evidence);
const cur = row.review?.rounds.at(-1);
if (!cur) throw refuse(`row ${row.id} has no review round to cite`);
+ if (ev.round !== cur.n) throw refuse(`evidence names round ${ev.round}; row ${row.id} is in round ${cur.n}`);
if (ev.candidate !== cur.candidate.digest) throw refuse(`evidence candidate ${ev.candidate} is not round ${cur.n}'s candidate ${cur.candidate.digest}`);
round = cur.n;
next.claim = null;
@@ -491,7 +515,7 @@ function applyMove(rows, row, entry, resolved) {
throw illegal();
}
next = touch(next, entry);
- return { row: next, result: { row: row.id, from, to, round, candidate: cand } };
+ return { row: next, result: { row: row.id, from, to, round, issue: revIssue, candidate: cand } };
}
function applySet(rows, row, entry, resolved) {
@@ -588,7 +612,7 @@ export function applyEntry(state, entry, resolved) {
const out = applyMove(rows, row, entry, resolved);
rows.set(row.id, out.row);
const r = out.result;
- result = { ...r, receipt: receipt(entry, rev, `row ${row.id} ${r.from}→${r.to}${r.round ? ` round ${r.round}` : ""}`) };
+ result = { ...r, receipt: receipt(entry, rev, `row ${row.id} ${r.from}→${r.to}${r.round ? ` round ${r.round}` : ""}${r.issue ? ` on #${r.issue}` : ""}`) };
break;
}
case "release": {
diff --git a/packages/queue/src/store.mjs b/packages/queue/src/store.mjs
index 97f8784c..c3a66b8c 100644
--- a/packages/queue/src/store.mjs
+++ b/packages/queue/src/store.mjs
@@ -242,7 +242,11 @@ function writeWitness(ctx, loc, doc, bytes) {
unlinkQuiet(ctx.io, tmp);
throw err;
}
- ctx.io.fsyncDir(loc.gitDir);
+ try {
+ ctx.io.fsyncDir(loc.gitDir);
+ } catch (err) {
+ throw Object.assign(new Error(errno(err)), { code: err.code, renamed: true });
+ }
}
function confirmTail(ctx, loc, cur) {
@@ -316,6 +320,7 @@ function writeView(ctx, loc, before, parts, body, tag) {
const now = readOrNull(ctx.io, loc.viewPath);
if (now === null || !now.equals(before)) return stale;
const tmp = `${loc.viewPath}.${tag}.tmp`;
+ let renamed = false;
try {
const mode = Number(ctx.io.stat(loc.viewPath).mode & 0o777n);
unlinkQuiet(ctx.io, tmp);
@@ -326,9 +331,11 @@ function writeView(ctx, loc, before, parts, body, tag) {
return stale;
}
ctx.io.rename(tmp, loc.viewPath);
+ renamed = true;
ctx.io.fsyncDir(loc.docsDir);
return null;
} catch (err) {
+ if (renamed) return `the view is written but not confirmed durable (${errno(err)}); the op stands; after a host crash, check the table with \`${FIX} verify\``;
unlinkQuiet(ctx.io, tmp);
return `the view write failed (${errno(err)}); the op stands and the view is stale; run \`${FIX} render\``;
}
@@ -390,7 +397,8 @@ function commitWrite(ctx, loc, cur, doc, bytes, op, view, body, exclusive) {
try {
writeWitness(ctx, loc, doc, bytes);
} catch (err) {
- throw new QueueError(`uncertain ${op} rev ${rev}: durable, witness not updated (${errno(err)})`, 3);
+ const what = err.renamed ? "witness written, its directory fsync failed" : "witness not updated";
+ throw new QueueError(`uncertain ${op} rev ${rev}: durable, ${what} (${errno(err)})`, 3);
}
ctx.hook("witnessed");
const warn = writeView(ctx, loc, view.bytes, view.parts, body, op);
@@ -402,14 +410,19 @@ function withLock(ctx, loc, { op = null, verb }, fn) {
checkPlatform(ctx.io, [loc.docsDir, loc.gitDir]);
const handle = acquire({ gitDir: loc.gitDir, io: ctx.io, proc: ctx.proc, op, verb, waitMs: ctx.lockWaitMs, stepMs: ctx.lockStepMs, hook: ctx.hook });
const res = { out: [], err: [], code: 0 };
+ let failure = null;
try {
ctx.hook("locked");
fn(res);
- } finally {
- let msg;
- try { msg = release(handle, ctx.io); } catch (err) { msg = `cannot release the queue lock (${errno(err)})`; }
- if (msg) res.err.push(`warning: ${msg}`);
+ } catch (err) {
+ failure = err;
}
+ let msg;
+ try { msg = release(handle, ctx.io); } catch (err) { msg = `cannot release the queue lock (${errno(err)})`; }
+ // A refusal still reports what release found (8.4).
+ if (msg && failure instanceof Error) failure.message += `\nwarning: ${msg}`;
+ else if (msg) res.err.push(`warning: ${msg}`);
+ if (failure) throw failure;
return res;
}
@@ -763,5 +776,6 @@ export function unlock(opts, { checkGateOnly = false } = {}) {
const ctx = makeCtx(opts);
const loc = unlockLoc(ctx);
if (checkGateOnly) return { out: [checkGate({ gitDir: loc.gitDir, io: ctx.io, proc: ctx.proc }).line], err: [], code: 0 };
- return { out: [unlockLock({ gitDir: loc.gitDir, io: ctx.io, proc: ctx.proc, hook: ctx.hook })], err: [], code: 0 };
+ const [line, ...warnings] = unlockLock({ gitDir: loc.gitDir, io: ctx.io, proc: ctx.proc, hook: ctx.hook }).split("\n");
+ return { out: [line], err: warnings, code: 0 };
}
diff --git a/packages/queue/tests/commit.test.mjs b/packages/queue/tests/commit.test.mjs
index 55105163..9a1b6ec1 100644
--- a/packages/queue/tests/commit.test.mjs
+++ b/packages/queue/tests/commit.test.mjs
@@ -152,7 +152,12 @@ fi`);
assert.equal(r.blob("HEAD", "src.txt"), "src\n");
});
-test("F1: a commit whose guard ran before update-ref fails at its own HEAD update", async (t) => {
+// A commit paused in its editor after its guard passed against H. Whether
+// git holds index.lock during the editor depends on the form: git 2.55
+// doesn't for plain `commit -e` and does for `commit -e -- path`. Step 8
+// reconciles when the lock is free and exits 3 when it isn't; either way
+// the paused commit loses at its own HEAD update.
+async function pausedCommit(t, form) {
const r = ready(t);
note(r);
stageFile(r, "src.txt");
@@ -161,27 +166,38 @@ test("F1: a commit whose guard ran before update-ref fails at its own HEAD updat
const go = join(r.ctl, "editor-go");
const editor = join(r.ctl, "editor.sh");
writeFileSync(editor, `#!/bin/sh\n: > ${q(started)}\nwhile [ ! -e ${q(go)} ]; do sleep 0.05; done\necho "ordinary" > "$1"\n`, { mode: 0o755 });
- const child = spawn("git", ["-C", r.root, "commit", "-e", "-q"], { env: { ...r.env, GIT_EDITOR: editor }, stdio: ["ignore", "pipe", "pipe"] });
+ const child = spawn("git", ["-C", r.root, "commit", "-e", "-q", ...form], { env: { ...r.env, GIT_EDITOR: editor }, stdio: ["ignore", "pipe", "pipe"] });
let childErr = "";
child.stderr.on("data", (d) => { childErr += d; });
const exited = new Promise((resolve) => child.on("exit", resolve));
for (let i = 0; i < 200 && !existsSync(started); i++) sleepMs(50);
assert.ok(existsSync(started), "the editor never started");
- // The paused commit ran its guard against H and holds index.lock.
+ const locked = existsSync(join(r.gitDir, "index.lock"));
+ t.diagnostic(`git commit -e${form.map((a) => ` ${a}`).join("")}: index.lock ${locked ? "held" : "free"} during the editor`);
const res = r.qc(["-m", "queue rev 1"]);
writeFileSync(go, "");
const code = await exited;
- // git 2.55 does not hold index.lock while the editor runs, so step 8
- // reconciles; the paused commit then loses at its HEAD update.
- assert.equal(res.code, 0, res.err);
+ assert.equal(res.code, locked ? 3 : 0, `index.lock ${locked ? "held" : "free"} during the editor: ${res.err}`);
+ if (locked) assert.match(res.err, /another git process holds \.git\/index\.lock/);
const c = r.head();
assert.equal(r.g("rev-parse", "HEAD^").trim(), h);
+ assert.equal(r.revAt(c), 1);
assert.notEqual(code, 0);
assert.match(childErr, new RegExp(`cannot lock ref 'HEAD': is at ${c} but expected ${h}`));
assert.equal(r.head(), c, "the old queue landed on top of C");
+ if (locked) r.g("reset", "-q", "--", "docs/plans/queue.json", "docs/plans/QUEUE.md");
assert.equal(r.g("diff", "--cached", "--name-only").trim(), "src.txt");
r.g("commit", "-q", "-m", "ordinary");
assert.equal(r.revAt("HEAD"), 1);
+ return locked;
+}
+
+test("F1: a plain `commit -e` whose guard ran before update-ref fails at its own HEAD update", async (t) => {
+ await pausedCommit(t, []);
+});
+
+test("F1: a `commit -e -- path` whose guard ran before update-ref fails at its own HEAD update", async (t) => {
+ await pausedCommit(t, ["--", "src.txt"]);
});
test("F1: step 8 with index.lock held exits 3, and ordinary commits stay refused until the printed command runs", (t) => {
diff --git a/packages/queue/tests/data.test.mjs b/packages/queue/tests/data.test.mjs
index 2e6c019f..99080d95 100644
--- a/packages/queue/tests/data.test.mjs
+++ b/packages/queue/tests/data.test.mjs
@@ -3,8 +3,8 @@
import assert from "node:assert/strict";
import { test } from "node:test";
import {
- CALLER_OP_RE, LOG_OP_RE, applyEntry, buildDoc, canonArgs, classifyView, countHeading, genesisReceipt, genesisRows,
- gitBlobId, loadDoc, nextFor, parseManifest, parseMigrationMap, render, rowsArray, serialize, sha256, splitView,
+ CALLER_OP_RE, LOG_OP_RE, STATES, applyEntry, buildDoc, canonArgs, classifyView, countHeading, genesisReceipt, genesisRows,
+ gitBlobId, loadDoc, nextFor, parseManifest, parseMigrationMap, render, rowsArray, serialize, sha256, splitView, validateRow,
} from "../src/queue.mjs";
import { QueueError } from "../src/errors.mjs";
import { MAP_ROWS, mapText } from "./helpers.mjs";
@@ -49,7 +49,7 @@ function refused(fn, re) {
assert.throws(fn, (err) => err instanceof QueueError && err.code === 2 && re.test(err.message));
}
-const mv = (id, to, extra = {}) => ({ id, to, reason: null, candidate: null, evidence: null, ...extra });
+const mv = (id, to, extra = {}) => ({ id, to, reason: null, candidate: null, evidence: null, issue: null, ...extra });
const row = (doc, id) => doc.rows.find((r) => r.id === id);
const MANIFEST = `${"c".repeat(64)} packages/queue/src/queue.mjs\n`;
const CAND = { kind: "manifest", digest: sha256(MANIFEST), text: MANIFEST };
@@ -156,7 +156,7 @@ test("matrix: release, review round, changes requested and waiting-on-jason", ()
const rv = row(d, 9).review;
assert.equal(rv.issue, 1508);
assert.deepEqual(rv.rounds.map((r) => [r.n, r.request, r.candidate.digest]), [[1, "none", CAND.digest]]);
- assert.match(d.log.at(-1).result.receipt, /in-progress→in-review round 1$/);
+ assert.match(d.log.at(-1).result.receipt, /in-progress→in-review round 1 on #1508$/);
refused(() => step(d, "move", mv(9, "in-progress"), "dewey"), /claimed by darkwing/);
d = step(d, "move", mv(9, "in-progress"), "darkwing");
assert.equal(row(d, 9).claim.seat, "darkwing");
@@ -176,9 +176,11 @@ test("matrix: release, review round, changes requested and waiting-on-jason", ()
test("matrix J5: in-review→done by the gate owner with evidence naming the current round", () => {
let d = row9Started();
d = step(d, "move", mv(9, "in-review", { candidate: "x" }), "darkwing", { candidate: CAND });
- const ev = `comment=4242,candidate=${CAND.digest}`;
- refused(() => step(d, "move", mv(9, "done"), "filbert"), /--evidence comment=<id>,candidate=<digest>/);
- refused(() => step(d, "move", mv(9, "done", { evidence: `comment=1,candidate=${"d".repeat(64)}` }), "filbert"), /is not round 1's candidate/);
+ const ev = `comment=4242,round=1,candidate=${CAND.digest}`;
+ refused(() => step(d, "move", mv(9, "done"), "filbert"), /--evidence comment=<id>,round=<n>,candidate=<digest>/);
+ refused(() => step(d, "move", mv(9, "done", { evidence: `comment=4242,candidate=${CAND.digest}` }), "filbert"), /round=<n>/);
+ refused(() => step(d, "move", mv(9, "done", { evidence: `comment=1,round=1,candidate=${"d".repeat(64)}` }), "filbert"), /is not round 1's candidate/);
+ refused(() => step(d, "move", mv(9, "done", { evidence: `comment=4242,round=2,candidate=${CAND.digest}` }), "filbert"), /evidence names round 2; row 9 is in round 1/);
refused(() => step(d, "move", mv(9, "done", { evidence: ev }), "rocko"), /only the gate owner \(filbert\)/);
const done = step(d, "move", mv(9, "done", { evidence: ev }), "filbert");
assert.equal(row(done, 9).state, "done");
@@ -187,6 +189,155 @@ test("matrix J5: in-review→done by the gate owner with evidence naming the cur
let e = step(genesisDoc(), "move", mv(11, "in-progress"), "dewey");
e = step(e, "move", mv(11, "in-review", { candidate: "x" }), "dewey", { candidate: CAND });
refused(() => step(e, "move", mv(11, "done", { evidence: ev }), "jason"), /gate is Jason's/);
+ // Changes requested, then the same candidate again: round 1's comment
+ // does not close round 2 (8.7, R3).
+ d = step(d, "move", mv(9, "in-progress"), "darkwing");
+ d = step(d, "move", mv(9, "in-review", { candidate: "x" }), "darkwing", { candidate: CAND });
+ assert.deepEqual(row(d, 9).review.rounds.map((r) => r.candidate.digest), [CAND.digest, CAND.digest]);
+ refused(() => step(d, "move", mv(9, "done", { evidence: ev }), "filbert"), /evidence names round 1; row 9 is in round 2/);
+ const done2 = step(d, "move", mv(9, "done", { evidence: `comment=4343,round=2,candidate=${CAND.digest}` }), "filbert");
+ assert.equal(done2.log.at(-1).result.round, 2);
+});
+
+// A row 12 owned by darkwing with the given issues, started, so the next
+// move is the review request.
+function row12Started(issues, { gateOwner = "filbert" } = {}) {
+ let d = step(genesisDoc(), "add", { piece: "N", gate: "g", brief: "docs/plans/brief-b.md#Queue", issues, owner: "darkwing", gateOwner }, "sage", { brief: brief() });
+ d = step(d, "move", mv(12, "briefed"), "sage");
+ return step(d, "move", mv(12, "in-progress"), "darkwing");
+}
+
+const review = (d, extra = {}) => step(d, "move", mv(12, "in-review", { candidate: "x", ...extra }), "darkwing", { candidate: CAND });
+const again = (d) => step(d, "move", mv(12, "in-progress"), "darkwing");
+
+test("review issue, lead decision 23: none refuses, one is used, several need --issue, later rounds keep it", () => {
+ // No issues: refused before the round opens.
+ refused(() => review(row12Started([])), /row 12 lists no issues; a privileged actor sets one before review/);
+ // One issue: used without --issue; --issue may name it; any other refuses.
+ const one = review(row12Started([1508]));
+ assert.equal(row(one, 12).review.issue, 1508);
+ assert.match(one.log.at(-1).result.receipt, /round 1 on #1508$/);
+ assert.equal(one.log.at(-1).result.issue, 1508);
+ refused(() => review(row12Started([1508]), { issue: 1495 }), /--issue #1495 is not one of row 12's issues \(#1508\)/);
+ // Several: --issue is required and must be one of the row's; the lowest
+ // number is not a default.
+ const several = row12Started([1495, 1508]);
+ refused(() => review(several), /row 12 lists several issues \(#1495, #1508\); name the review's issue with --issue/);
+ refused(() => review(several, { issue: 1600 }), /--issue #1600 is not one of row 12's issues \(#1495, #1508\)/);
+ let d = review(several, { issue: 1508 });
+ assert.equal(row(d, 12).review.issue, 1508);
+ // Later rounds keep the previous round's issue unless --issue names another.
+ d = review(again(d));
+ assert.deepEqual([row(d, 12).review.issue, row(d, 12).review.rounds.length], [1508, 2]);
+ d = review(again(d), { issue: 1495 });
+ assert.deepEqual([row(d, 12).review.issue, row(d, 12).review.rounds.length], [1495, 3]);
+ // A kept issue the row no longer lists refuses until --issue names one.
+ d = step(again(d), "set", { id: 12, field: "issues", value: [1508, 1600] }, "sage");
+ refused(() => review(d), /review issue #1495 is no longer one of its issues \(#1508, #1600\); name one with --issue/);
+ assert.equal(row(review(d, { issue: 1600 }), 12).review.issue, 1600);
+ // --issue belongs to the review request only.
+ refused(() => step(several, "move", mv(12, "blocked", { reason: "x", issue: 1508 }), "darkwing"), /--issue applies only to in-progress→in-review/);
+});
+
+test("the row schema refuses a review with a null issue", () => {
+ const r = structuredClone(row(review(row12Started([1508])), 12));
+ validateRow(r);
+ r.review.issue = null;
+ refused(() => validateRow(r), /review issue must be a positive integer/);
+});
+
+// R1: every state × target × actor class against 8.7's table, written from
+// the spec rather than from queue.mjs. Row 12 is owned by darkwing; the
+// gate owner is filbert, jason or the owner; rocko is any other seat.
+const ACTORS = ["darkwing", "filbert", "rocko", "sage", "jason"];
+const PRIV = new Set(["sage", "jason"]);
+
+function specAllows({ from, prev, to, by, gateOwner, required }) {
+ const own = by === "darkwing";
+ const priv = PRIV.has(by);
+ if (from === "done") return false;
+ if (to === "blocked") return ["queued", "briefed", "in-progress", "in-review", "waiting-on-jason"].includes(from) && (own || priv);
+ if (from === "blocked") return to === prev && (own || priv);
+ const edge = `${from}→${to}`;
+ switch (edge) {
+ case "queued→briefed": return priv;
+ case "briefed→in-progress": return own;
+ case "in-progress→in-review": return own; // the claimant is the owner
+ case "in-review→in-progress": case "in-review→waiting-on-jason": return own || priv;
+ case "waiting-on-jason→done": return by === "jason" || by === "sage"; // sage with evidence, which the case supplies
+ case "in-review→done": return gateOwner !== "jason" && (by === gateOwner || priv);
+ case "queued→parked": case "briefed→parked": return by === "jason" && !required;
+ case "parked→queued": return by === "jason";
+ default: return false; // in-progress→briefed is `release`, not `move`
+ }
+}
+
+function matrixStates(gateOwner, required) {
+ let q = step(genesisDoc(), "add", {
+ piece: "M", gate: "g", brief: "docs/plans/brief-b.md#Queue", issues: [1508], owner: "darkwing", gateOwner, required,
+ }, "sage", { brief: brief() });
+ const out = [{ from: "queued", doc: q }];
+ if (!required) out.push({ from: "parked", doc: step(q, "move", mv(12, "parked"), "jason") });
+ const b = step(q, "move", mv(12, "briefed"), "sage");
+ const p = step(b, "move", mv(12, "in-progress"), "darkwing");
+ const r = review(p);
+ const w = step(r, "move", mv(12, "waiting-on-jason"), "sage");
+ out.push({ from: "briefed", doc: b }, { from: "in-progress", doc: p }, { from: "in-review", doc: r }, { from: "waiting-on-jason", doc: w });
+ out.push({ from: "done", doc: step(w, "move", mv(12, "done"), "jason") });
+ for (const s of [...out]) {
+ if (!["done", "parked"].includes(s.from)) out.push({ from: "blocked", prev: s.from, doc: step(s.doc, "move", mv(12, "blocked", { reason: "r" }), "sage") });
+ }
+ return out;
+}
+
+function matrixArgs(from, to) {
+ const extra = {};
+ if (to === "blocked") extra.reason = "r";
+ if (from === "in-progress" && to === "in-review") extra.candidate = "x";
+ if (to === "done" && from === "in-review") extra.evidence = `comment=1,round=1,candidate=${CAND.digest}`;
+ if (to === "done" && from === "waiting-on-jason") extra.evidence = "Jason approved in thread X";
+ return mv(12, to, extra);
+}
+
+test("matrix R1: every state × target × actor class matches 8.7, gate owner jason or not, required or not", () => {
+ let allowed = 0;
+ let refusals = 0;
+ for (const gateOwner of ["filbert", "jason", "darkwing"]) {
+ for (const required of [false, true]) {
+ for (const { from, prev = null, doc } of matrixStates(gateOwner, required)) {
+ assert.equal(row(doc, 12).state, from);
+ for (const to of STATES) {
+ for (const by of ACTORS) {
+ const c = { from, prev, to, by, gateOwner, required };
+ const label = JSON.stringify(c);
+ let got;
+ try {
+ got = step(doc, "move", matrixArgs(from, to), by, { candidate: CAND });
+ } catch (err) {
+ assert.ok(err instanceof QueueError && err.code === 2, `${label}: ${err.stack}`);
+ assert.equal(specAllows(c), false, `${label} refused: ${err.message}`);
+ refusals++;
+ continue;
+ }
+ assert.equal(specAllows(c), true, `${label} was allowed`);
+ const r = row(got, 12);
+ assert.equal(r.state, to, label);
+ if (to === "blocked") assert.equal(r.previousState, from, label);
+ if (from === "blocked") assert.deepEqual([r.previousState, r.blockedReason], [null, null], label);
+ allowed++;
+ }
+ }
+ // release: the claimant or a privileged actor, from in-progress only.
+ for (const by of ACTORS) {
+ const ok = from === "in-progress" && (by === "darkwing" || PRIV.has(by));
+ const label = JSON.stringify({ release: from, by, gateOwner, required });
+ if (ok) assert.deepEqual([row(step(doc, "release", { id: 12 }, by), 12).state], ["briefed"], label);
+ else assert.throws(() => step(doc, "release", { id: 12 }, by), (err) => err instanceof QueueError && err.code === 2, label);
+ }
+ }
+ }
+ }
+ assert.ok(allowed > 100 && refusals > 1000, `allowed ${allowed}, refused ${refusals}`);
});
test("matrix: blocked keeps the claim and returns only to previousState", () => {
@@ -301,7 +452,7 @@ test("next: resume, then review, then start, then wait, then nothing; lowest id
let d = genesisDoc();
const add = (owner, extra = {}) => ({ piece: `p-${owner}`, gate: "g", brief: "docs/plans/brief-b.md#Queue", owner, ...extra });
d = step(d, "add", add("rocko"), "sage", { brief: brief() }); // 12
- d = step(d, "add", add("rocko", { reviewers: ["darkwing"] }), "sage", { brief: brief() }); // 13
+ d = step(d, "add", add("rocko", { reviewers: ["darkwing"], issues: [1508] }), "sage", { brief: brief() }); // 13
d = step(d, "add", add("darkwing"), "sage", { brief: brief() }); // 14
for (const id of [12, 13, 14]) d = step(d, "move", mv(id, "briefed"), "sage");
const rows = () => loadDoc(Buffer.from(serialize(d))).state.rows;
diff --git a/packages/queue/tests/lock.test.mjs b/packages/queue/tests/lock.test.mjs
index 9b6b2174..f1b109d9 100644
--- a/packages/queue/tests/lock.test.mjs
+++ b/packages/queue/tests/lock.test.mjs
@@ -89,6 +89,26 @@ test("a link error other than EEXIST refuses", (t) => {
assert.deepEqual(readdirSync(d), []);
});
+test("an error after the link releases the lock: unreadable gate, failing temp stat", (t) => {
+ const d = dir(t);
+ writeFileSync(join(d, GATE_NAME), record({ verb: "unlock", op: null }));
+ const denied = { ...realIo, readFile: (p) => (p.endsWith(GATE_NAME) ? (() => { throw Object.assign(new Error("denied"), { code: "EACCES" }); })() : realIo.readFile(p)) };
+ refused(() => acquire({ gitDir: d, io: denied, verb: "move" }), /cannot check the unlock gate .*EACCES; lock released$/, 1);
+ assert.deepEqual(readdirSync(d), [GATE_NAME]);
+ rmSync(join(d, GATE_NAME));
+ const badStat = { ...realIo, stat: () => { throw Object.assign(new Error("io"), { code: "EIO" }); } };
+ refused(() => acquire({ gitDir: d, io: badStat, verb: "move" }), /cannot write the lock record .*EIO; no lock taken/, 1);
+ assert.deepEqual(readdirSync(d), []);
+ // A stat that fails from its second call on: publish stats once, before the
+ // link, so nothing after the link can fail and strand the lock.
+ let stats = 0;
+ const lateStat = { ...realIo, stat: (p) => { if (++stats > 1) throw Object.assign(new Error("io"), { code: "EIO" }); return realIo.stat(p); } };
+ const h = acquire({ gitDir: d, io: lateStat, verb: "move" });
+ assert.equal(stats, 1);
+ assert.equal(release(h, realIo), null);
+ assert.deepEqual(readdirSync(d), []);
+});
+
test("a paused holder: another writer waits 10 s, then refuses naming it live", async (t) => {
const d = dir(t);
const child = spawn(process.execPath, [join(HERE, "fixtures", "lock-child.mjs"), d, "hold"], { stdio: ["ignore", "pipe", "ignore"] });
@@ -150,6 +170,22 @@ test("a writer publishing during an unlock, gate first: the writer releases and
assert.deepEqual(readdirSync(d), []);
});
+test("a gate swapped while held is left in place and reported, on success and on refusal (N1)", (t) => {
+ const d = dir(t);
+ const gate = join(d, GATE_NAME);
+ // A copy renamed over the gate: same bytes, a new inode.
+ const swap = () => { writeFileSync(`${gate}.copy`, readFileSync(gate)); renameSync(`${gate}.copy`, gate); };
+ const out = unlock({ gitDir: d, io: realIo, hook: (name) => { if (name === "gate-held") swap(); } });
+ assert.match(out, /^no queue lock present; nothing removed\nwarning: lock .*mosaic-queue\.unlock is not the one this process took; left in place$/);
+ assert.equal(existsSync(gate), true);
+ rmSync(gate);
+ writeFileSync(join(d, LOCK_NAME), record({}));
+ refused(() => unlock({ gitDir: d, io: realIo, hook: (name) => { if (name === "gate-held") swap(); } }),
+ /owner is live: .*; unlock refuses\nwarning: lock .*mosaic-queue\.unlock is not the one this process took; left in place$/);
+ assert.equal(existsSync(gate), true);
+ assert.equal(existsSync(join(d, LOCK_NAME)), true);
+});
+
test("a reused pid within one boot is mismatch; unlock removes the lock and never signals the process", async (t) => {
const d = dir(t);
const s = await sleeper(t);
@@ -177,6 +213,12 @@ test("a foreign host is unknown whatever the local pid says; unlock refuses", as
refused(() => acquire({ gitDir: d, io: realIo, verb: "move", waitMs: 0 }), /unknown: .*recorded on host some-other-host; unlock refuses this too/);
refused(() => unlock({ gitDir: d, io: realIo }), /owner is unknown: .*; unlock refuses/);
assert.equal(existsSync(join(d, LOCK_NAME)), true);
+ // A real foreign host has its own boot id. Host is tested before boot, so
+ // this is still unknown, never mismatch, and unlock still refuses.
+ writeFileSync(join(d, LOCK_NAME), record({ pid: await deadPid(), host: "some-other-host", boot: OTHER_BOOT }));
+ assert.equal(classify(readFileSync(join(d, LOCK_NAME)), realProc).state, "unknown");
+ refused(() => unlock({ gitDir: d, io: realIo }), /owner is unknown: .*recorded on host some-other-host/);
+ assert.equal(existsSync(join(d, LOCK_NAME)), true);
});
test("unreadable /proc: classification is unknown and acquire refuses", (t) => {
diff --git a/packages/queue/tests/store.test.mjs b/packages/queue/tests/store.test.mjs
index 7033b076..70425a73 100644
--- a/packages/queue/tests/store.test.mjs
+++ b/packages/queue/tests/store.test.mjs
@@ -166,10 +166,28 @@ test("Rocko's S4 schedule: a lost result, another writer, then the retry opens n
const review = ["move", "9", "in-review", "--candidate", "HEAD", "--op", "review-9-00001"];
cli(repo, review, { by: "darkwing" }); // result lost
ok(cli(repo, ["note", "9", "looking now", "--op", "note-9-000001"], { by: "filbert" }));
- ok(cli(repo, review, { by: "darkwing" }), /round 1 \(already recorded at rev 3\)/);
+ ok(cli(repo, review, { by: "darkwing" }), /round 1 on #1508 \(already recorded at rev 3\)/);
assert.equal(row(repo, 9).review.rounds.length, 1);
});
+test("the review issue and the evidence round through the CLI (lead decision 23, 8.7)", (t) => {
+ const repo = ready(t);
+ ok(cli(repo, ["move", "6", "blocked", "--reason", "paused", "--op", "block-6-00001"], { by: "darkwing" }));
+ ok(cli(repo, ["set", "9", "issues", "1495,1508", "--op", "issues-9-0001"], { by: "sage" }));
+ ok(cli(repo, ["move", "9", "in-progress", "--op", "start-9-00001"], { by: "darkwing" }));
+ const review = (op, ...extra) => ["move", "9", "in-review", "--candidate", "HEAD", "--op", op, ...extra];
+ no(cli(repo, review("review-9-00001"), { by: "darkwing" }), 2, /lists several issues \(#1495, #1508\); name the review's issue with --issue/);
+ no(cli(repo, review("review-9-00001", "--issue", "1495", "--issue", "1508"), { by: "darkwing" }), 4, /one --issue/);
+ no(cli(repo, review("review-9-00001", "--issue", "#1600"), { by: "darkwing" }), 2, /--issue #1600 is not one of row 9's issues/);
+ ok(cli(repo, review("review-9-00001", "--issue", "#1508"), { by: "darkwing" }), /in-progress→in-review round 1 on #1508$/m);
+ const head = repo.g("rev-parse", "HEAD").trim();
+ no(cli(repo, ["move", "9", "done", "--evidence", `comment=7,candidate=${head}`, "--op", "done-9-000001"], { by: "filbert" }), 2, /round=<n>/);
+ ok(cli(repo, ["move", "9", "in-progress", "--op", "changes-9-0001"], { by: "darkwing" }));
+ ok(cli(repo, review("review-9-00002"), { by: "darkwing" }), /round 2 on #1508$/m);
+ no(cli(repo, ["move", "9", "done", "--evidence", `comment=7,round=1,candidate=${head}`, "--op", "done-9-000001"], { by: "filbert" }), 2, /evidence names round 1; row 9 is in round 2/);
+ ok(cli(repo, ["move", "9", "done", "--evidence", `comment=8,round=2,candidate=${head}`, "--op", "done-9-000002"], { by: "filbert" }), /in-review→done round 2$/m);
+});
+
test("claims and add defaults through the CLI; candidates are manifests or reachable commits", (t) => {
const repo = ready(t);
ok(cli(repo, ["add", "--op", "add-by-dewey-1", "--piece", "Mine", "--gate", "tests", "--brief", "docs/plans/brief-b.md#Template"], { by: "dewey" }));
diff --git a/packages/queue/tests/write.test.mjs b/packages/queue/tests/write.test.mjs
index 0d0781c0..35e9829a 100644
--- a/packages/queue/tests/write.test.mjs
+++ b/packages/queue/tests/write.test.mjs
@@ -3,7 +3,7 @@
// racing a writer (8.4, F2).
import assert from "node:assert/strict";
import { spawn, spawnSync } from "node:child_process";
-import { readFileSync, readdirSync, unlinkSync, writeFileSync } from "node:fs";
+import { readFileSync, readdirSync, renameSync, unlinkSync, writeFileSync } from "node:fs";
import { join } from "node:path";
import { test } from "node:test";
import { cli, genesisCommitted, load, scratchRepo } from "./helpers.mjs";
@@ -41,6 +41,7 @@ function faultIo(m, name, match, code) {
return {
...real,
openExcl: (p, mode) => { const fd = real.openExcl(p, mode); paths.set(fd, p); return fd; },
+ openRead: (p) => { const fd = real.openRead(p); paths.set(fd, p); return fd; },
close: (fd) => { paths.delete(fd); real.close(fd); },
write: (fd, b, off, len) => (hit("write", paths.get(fd)) ? (code === "SHORT" ? 0 : fail()) : real.write(fd, b, off, len)),
fsync: (fd) => (hit("fsync", paths.get(fd)) ? fail() : real.fsync(fd)),
@@ -114,6 +115,65 @@ test("a witness write failure: uncertain, durable, exit 3; the view is untouched
assert.match(m.store.mutate(o(repo), note(9, "x", "note-9-000001")).out[0], /already recorded at rev 1/);
});
+test("the .git fsync after the witness rename fails: uncertain, exit 3, the witness says so", async (t) => {
+ const { repo, m } = await ready(t);
+ const io = faultIo(m, "fsyncDir", (d) => d === repo.gitDir, "EIO");
+ throwsCode(() => m.store.mutate(o(repo, { io }), note(9, "x", "note-9-000001")), 3, /^uncertain note-9-000001 rev 1: durable, witness written, its directory fsync failed \(EIO\)$/);
+ assert.equal(revOf(repo), 1);
+ assert.equal(witness(repo).revision, 1);
+ assert.equal(shownRev(repo), 0);
+ assert.match(m.store.mutate(o(repo), note(9, "x", "note-9-000001")).out[0], /already recorded at rev 1/);
+});
+
+test("confirming a tail fsyncs queue.json and docs/plans before the witness; either failure changes nothing", async (t) => {
+ const { repo, m } = await ready(t);
+ const docs = join(repo.root, "docs/plans");
+ throwsCode(() => m.store.mutate(o(repo, { io: faultIo(m, "fsyncDir", (d) => d === docs, "EIO") }), note(9, "x", "note-9-000001")), 3, /uncertain/);
+ for (const io of [faultIo(m, "fsync", (p) => p === repo.queuePath, "EIO"), faultIo(m, "fsyncDir", (d) => d === docs, "EIO")]) {
+ throwsCode(() => m.store.sync(o(repo, { io })), 1, /^cannot confirm rev 1 durable \(EIO\); nothing changed$/);
+ assert.equal(witness(repo).revision, 0);
+ }
+ assert.match(m.store.sync(o(repo)).out.join("\n"), /durable now, never acknowledged: note-9-000001/);
+ assert.equal(witness(repo).revision, 1);
+});
+
+test("the docs/plans fsync after the view rename fails: the op stands, the view is written, a warning says so", async (t) => {
+ const { repo, m } = await ready(t);
+ const docs = join(repo.root, "docs/plans");
+ let calls = 0;
+ const io = { ...m.io.realIo, fsyncDir: (d) => { if (d === docs && ++calls === 2) throw Object.assign(new Error("EIO"), { code: "EIO" }); m.io.realIo.fsyncDir(d); } };
+ const r = m.store.mutate(o(repo, { io }), note(9, "x", "note-9-000001"));
+ assert.equal(calls, 2);
+ assert.match(r.out[0], /^ok note-9-000001 rev 1/);
+ assert.match(r.err.join("\n"), /the view is written but not confirmed durable \(EIO\); the op stands/);
+ assert.equal(shownRev(repo), 1);
+ assert.deepEqual(tmps(repo), []);
+});
+
+test("a lock swapped while held is left in place and reported, on a receipt and on a refusal", async (t) => {
+ const { repo, m } = await ready(t);
+ const lock = join(repo.gitDir, "mosaic-queue.lock");
+ // Another inode with the same bytes, as a delayed unlock and relock would leave.
+ const swap = (name) => { if (name === "locked") { writeFileSync(`${lock}.copy`, readFileSync(lock)); renameSync(`${lock}.copy`, lock); } };
+ const done = m.store.mutate(o(repo, { hook: swap }), note(9, "x", "note-9-000001"));
+ assert.match(done.out[0], /^ok note-9-000001 rev 1/);
+ assert.match(done.err.join("\n"), /warning: lock .* is not the one this process took; left in place/);
+ unlinkSync(lock);
+ throwsCode(() => m.store.mutate(o(repo, { hook: swap }), note(9, "y", "note-9-000002", "rocko")), 2,
+ /may note row 9[^]*\nwarning: lock .* is not the one this process took; left in place$/);
+ unlinkSync(lock);
+});
+
+test("unlock prints a swapped gate's warning on stderr, the result on stdout", async (t) => {
+ const { repo, m } = await ready(t);
+ const gate = join(repo.gitDir, "mosaic-queue.unlock");
+ const swap = (name) => { if (name === "gate-held") { writeFileSync(`${gate}.copy`, readFileSync(gate)); renameSync(`${gate}.copy`, gate); } };
+ const r = m.store.unlock(o(repo, { hook: swap }));
+ assert.deepEqual(r.out, ["no queue lock present; nothing removed"]);
+ assert.match(r.err.join("\n"), /^warning: lock .*mosaic-queue\.unlock is not the one this process took; left in place$/);
+ unlinkSync(gate);
+});
+
test("a view write that fails keeps the op and reports a stale view", async (t) => {
const { repo, m } = await ready(t);
const io = faultIo(m, "rename", (p) => p === repo.viewPath, "EIO");
+39
View File
@@ -0,0 +1,39 @@
#!/usr/bin/env bash
# N13 check (#1508): a failing nested `node --test` must fail test-foundation.sh
# and test-discord.sh even when a parent runner's NODE_TEST_CONTEXT is set.
# Runs in a scratch clone only: n13-check.sh CLONE (a clone of HEAD with
# node_modules linked). It copies this checkout's two suites into the clone,
# plants a failing test in each suite's test directory, and restores after.
set -uo pipefail
SRC="$(cd "$(dirname "$0")/../../../.." && pwd)"
CLONE="${1:?usage: n13-check.sh CLONE}"
[ "$(cd "$CLONE" && pwd)" != "$SRC" ] || { echo "refusing to run in the canonical checkout" >&2; exit 4; }
cd "$CLONE" || exit 1
PLANT='import { test } from "node:test";
import assert from "node:assert/strict";
test("N13 planted failure", () => assert.equal(1, 2));'
FAILS=0
expect() { # NAME WANT GOT
if [ "$2" = "$3" ]; then echo "ok $1 (rc $3)"; else echo "FAIL $1 (want rc $2, got $3)"; FAILS=$((FAILS+1)); fi
}
run() { # SUITE -> rc, output in /tmp/n13-<suite>-<tag>.txt
NODE_TEST_CONTEXT=child-v8 NO_COLOR=1 "scripts/test-$1.sh" >"/tmp/n13-$1-$2.txt" 2>&1
}
for pair in foundation:scripts/foundation discord:packages/discord/tests; do
suite=${pair%%:*} dir=${pair#*:}
git checkout -q -- "scripts/test-$suite.sh"
printf '%s\n' "$PLANT" >"$dir/zz-n13-planted.test.mjs"
run "$suite" head-planted; expect "$suite at HEAD, planted failure, parent context set" 0 $?
cp -p "$SRC/scripts/test-$suite.sh" "scripts/test-$suite.sh"
run "$suite" fixed-planted; expect "$suite fixed, planted failure, parent context set" 1 $?
rm -f "$dir/zz-n13-planted.test.mjs"
run "$suite" fixed-clean; expect "$suite fixed, no planted failure, parent context set" 0 $?
sed -i 's/node_tests() { env -u NODE_TEST_CONTEXT node --test/node_tests() { node --test/' "scripts/test-$suite.sh"
run "$suite" mutant-clean; expect "$suite with the clearing removed, no planted failure" 1 $?
git checkout -q -- "scripts/test-$suite.sh"
done
git status --short
echo "n13 check: $FAILS failed"
[ "$FAILS" -eq 0 ]
+59
View File
@@ -0,0 +1,59 @@
# N13: nested `node --test` in two suites (#1508)
Darkwing, 2026-09-26. Filbert's note N13, put in DEFERRED by Sage as a
small reviewed item before A2. Nothing is committed, staged or pushed.
## The defect
`scripts/test-foundation.sh:76` and `scripts/test-discord.sh:142` start a
nested `node --test` without clearing `NODE_TEST_CONTEXT`. Under a parent
test runner the nested run reports to that runner and exits 0 whatever its
tests do. I reproduced it at HEAD 40a02d2b: with a failing test planted in
each suite's test directory and `NODE_TEST_CONTEXT=child-v8` set, both
suites exit 0. The check line reads `OK node --test ... (summary
missing)`. In a control run of foundation without the variable, the same
planted test fails the suite, so only a run under a parent runner is
blind. I ran that control for foundation only.
## The change
`n13.patch` (sha256 00868b2f) changes only the two suites, +20 −2 lines.
Each suite now has a `node_tests` function that runs `env -u
NODE_TEST_CONTEXT node --test`, and uses it for its test run. Each also
gets one new check. It writes a failing test into the sandbox, runs it
through `node_tests` with `NODE_TEST_CONTEXT=child-v8` set, and requires
exit 1 and `✖ planted failure` in the output. If someone drops the `env
-u`, that check fails on every run, not only under a parent runner.
| File | sha256 |
|---|---|
| `scripts/test-foundation.sh` | 60f04822 |
| `scripts/test-discord.sh` | 2ad3be74 |
| `n13-check.sh` | 6a231759 |
## The check
`n13-check.sh CLONE` runs in a scratch clone only and refuses the
canonical checkout. For each suite it plants a failing test in the real
test directory and runs the whole suite with `NODE_TEST_CONTEXT=child-v8`:
| Suite | Case | Exit |
|---|---|---|
| foundation | HEAD, planted failure | 0 (the defect) |
| foundation | fixed, planted failure | 1 |
| foundation | fixed, no planted failure | 0 |
| foundation | fixed but `env -u` removed, no planted failure | 1 |
| discord | HEAD, planted failure | 0 (the defect) |
| discord | fixed, planted failure | 1 |
| discord | fixed, no planted failure | 0 |
| discord | fixed but `env -u` removed, no planted failure | 1 |
In the fixed runs with the planted failure, the suite prints `FAIL node
--test ...` with the pass count and the planted test's name. The last row
of each is the mutation: the new check alone fails the suite.
All nine suites pass at 40a02d2b with this change and the queue A1
candidate: foundation 44 and discord 64, one more check each than before.
I found no other nested `node --test` in the suites. `test-queue.sh` and
`queue-commit.sh` already clear the variable.
+44
View File
@@ -0,0 +1,44 @@
diff --git a/scripts/test-discord.sh b/scripts/test-discord.sh
index 6044100c..9ec4ddda 100755
--- a/scripts/test-discord.sh
+++ b/scripts/test-discord.sh
@@ -139,7 +139,16 @@ else
fi
# --- the seven offline groups ---
-node --test --test-reporter=spec packages/discord/tests/ >"$SANDBOX/node-test.log" 2>&1
+# A nested `node --test` inherits a parent runner's NODE_TEST_CONTEXT, reports
+# to that runner and exits 0 whatever its tests do, so the suite clears it
+# (#1508 N13). The planted failing test proves a failure still fails here.
+node_tests() { env -u NODE_TEST_CONTEXT node --test "$@"; }
+mkdir -p "$SANDBOX/planted"
+printf '%s\n' 'import { test } from "node:test";' 'import assert from "node:assert/strict";' 'test("planted failure", () => assert.equal(1, 2));' >"$SANDBOX/planted/planted.test.mjs"
+NODE_TEST_CONTEXT=child-v8 node_tests "$SANDBOX/planted/" >"$SANDBOX/planted.log" 2>&1
+[ $? -eq 1 ] && grep -q '^✖ planted failure' "$SANDBOX/planted.log"
+check "a failing nested test fails the run under a parent runner's NODE_TEST_CONTEXT" $?
+node_tests --test-reporter=spec packages/discord/tests/ >"$SANDBOX/node-test.log" 2>&1
NODE_RC=$?
check "node --test packages/discord/tests/ ($(grep -E '^ℹ pass' "$SANDBOX/node-test.log" | tr -d '\n' || echo 'summary missing'))" $NODE_RC
if [ "$NODE_RC" -ne 0 ]; then
diff --git a/scripts/test-foundation.sh b/scripts/test-foundation.sh
index 251eb674..b94dab5c 100755
--- a/scripts/test-foundation.sh
+++ b/scripts/test-foundation.sh
@@ -73,7 +73,16 @@ done
check "checked-in demo bundles equal a fresh generation" $DEMO_OK
# --- unit, CLI, privacy, non-effect and fixture-index tests ---
-node --test scripts/foundation/ >"$SANDBOX/node-test.log" 2>&1
+# A nested `node --test` inherits a parent runner's NODE_TEST_CONTEXT, reports
+# to that runner and exits 0 whatever its tests do, so the suite clears it
+# (#1508 N13). The planted failing test proves a failure still fails here.
+node_tests() { env -u NODE_TEST_CONTEXT node --test "$@"; }
+mkdir -p "$SANDBOX/planted"
+printf '%s\n' 'import { test } from "node:test";' 'import assert from "node:assert/strict";' 'test("planted failure", () => assert.equal(1, 2));' >"$SANDBOX/planted/planted.test.mjs"
+NODE_TEST_CONTEXT=child-v8 node_tests "$SANDBOX/planted/" >"$SANDBOX/planted.log" 2>&1
+[ $? -eq 1 ] && grep -q '^✖ planted failure' "$SANDBOX/planted.log"
+check "a failing nested test fails the run under a parent runner's NODE_TEST_CONTEXT" $?
+node_tests scripts/foundation/ >"$SANDBOX/node-test.log" 2>&1
NODE_RC=$?
check "node --test scripts/foundation/ ($(grep -E '^ℹ pass' "$SANDBOX/node-test.log" | tr -d '\n' || echo 'summary missing'))" $NODE_RC
[ "$NODE_RC" -ne 0 ] && grep -E "^✖|AssertionError" "$SANDBOX/node-test.log" | head -20
+159
View File
@@ -0,0 +1,159 @@
# Queue A1 (#1508), revision r1
Darkwing, 2026-09-26. This answers Filbert's review
(`agents/filbert/work/queue-a1-review-2026-09-26.md`, sha256 6933b885) and
Sage's lead decision 23 (40a02d2b). Nothing is committed, staged or pushed.
## Files
| File | sha256 | What it is |
|---|---|---|
| `delta-r1.patch` | b733b894 | the change on top of `build.patch` |
| `build-manifest-r1.sha256` | 85a8a453 | all 20 files after the delta |
The delta changes 11 of the 20 files, +410 −49 lines. It adds no file and changes no mode.
In a fresh clone at 3a209eea, `build.patch` and then `delta-r1.patch` apply
cleanly, and the result matches the new manifest 20/20.
## Required changes
**R1, the matrix.** `data.test.mjs` has a new test, "matrix R1". Its
oracle, `specAllows`, is written from 8.7's table, not from `queue.mjs`.
The test replays a row into every reachable state and tries every target
in `STATES` as five actors: darkwing (the row's owner), filbert, rocko,
sage and jason. It does this with the gate owner as filbert, jason and the
row's owner, and with the row required and not. Release gets the same
treatment. That is 273 allowed moves and 2487 refusals, each compared with
what `applyOp` does. Your R1 mutation and the agent's two survivors each
fail it (R1a to R1c below).
**R2, the review issue, as lead decision 23 rules.** `move in-review`
refuses a row with no issues. A row with one issue uses it. A row with
several needs `--issue N`, and N must be one of them. A later round keeps
the previous round's issue unless `--issue` names another. `--issue` is
accepted only on in-progress→in-review, and giving it twice is a usage
error (exit 4). The receipt now ends `round N on #ISSUE`.
The ruling didn't cover one case: a later round whose kept issue the row
no longer lists, after a `set issues`. I refuse it until `--issue` names
one of the row's issues. Falling back to the first issue would be the
silent choice decision 23 replaced.
The row schema now refuses `review.issue: null`, so a hand-built file
can't hold one either. `set piece` and `set gate` stay privileged only
(decision 23, point 2); the code already did that, and nothing changed.
Tests: "review issue, lead decision 23" in `data.test.mjs` covers the four
cases and a refused `--issue` that isn't in the row. "the review issue and
the evidence round through the CLI" in `store.test.mjs` runs the same
through the CLI. "the row schema refuses a review with a null issue"
checks `validateRow`. Mutations D1 to D5.
**R3, the evidence round.** The format is now
`comment=<id>,round=<n>,candidate=<digest>`. `move done` compares the
round with the current one and refuses `evidence names round 1; row 12 is
in round 2`. The J5 test refuses evidence with no round and with a wrong
round, then sends a row back and re-requests it with the same candidate:
round-1 evidence is refused and round-2 evidence closes it. Mutation E1.
## Notes I took
- **N1.** `withLock` appends the release warning to a refusal's message.
`unlock` now does the same for the gate: a swapped gate is left in place
and reported, on a refusal and on success. On success the CLI prints the
warning on stderr and the result on stdout. Mutations W1, G1, G2, U1.
- **N2.** In `acquire`, an error checking the gate releases the lock, then
refuses with `cannot check the unlock gate ...; lock released`. In
`publish`, the temp file's stat now runs before the link, so nothing
that can fail runs between a successful link and the return. The test
covers an unreadable gate, a stat that fails before the link, and a stat
that fails from its second call on. The last case came late. L1 (a
second stat after the link) survived my first mutation run, so I added
it; L1 is now caught.
- **N3.** `write.test.mjs` has three fault tests: the `queue.json` and
`docs/plans` fsyncs in `confirmTail` (sync exits 1, nothing changes);
the `.git` fsync after the witness rename (exit 3); the `docs/plans`
fsync after the view rename (the op stands, a warning says the view
isn't confirmed durable). Mutations F1 to F4.
- **N4.** The foreign-host test adds a record with another boot id. It
must still classify `unknown`. Mutation H1 swaps the two checks.
- **N6.** The message now says `durable, witness written, its directory
fsync failed` when the rename happened, and `witness not updated` only
when it didn't.
- **N9.** `BRIEF-TEMPLATE.md`: `briefed` when a privileged actor (jason or
sage) accepts it; the owner can't.
- **N14.** You were right. The paused-editor test is now `pausedCommit(t,
form)`, which checks whether `index.lock` exists at the pause and
asserts on that: exit 3 and `another git process holds .git/index.lock`
when held, exit 0 when free. Either way the paused commit then fails
with `cannot lock ref 'HEAD': is at C but expected H`. Two forms run on
git 2.55.0: plain `commit -e` (the lock was free, exit 0) and `commit -e
-- src.txt` (the lock was held, exit 3). The test no longer pins a git
version, and the stale comment is gone.
## A correction to build.md
build.md says "`unlock` works on a missing or invalid lock file". That's
wrong, as you said. It means a missing or invalid `queue.json`. `unlock`
refuses an invalid lock. build.md stays as sent, since your review pins
it.
## Notes not taken
N5, N7, N8, N10, N11, N12, N15 and N16. They don't block, and none is in
the files this round had to touch for a reason. N8 (the `\` escape in
`cell()`) and N11 (replay looser than the CLI on op ids) are cheapest
before genesis. That's Sage's call; I can take them in A2.
## Verification
At 3a209eea plus the candidate (`/tmp/qa1-verify`): `test-queue.sh` 19
checks, `node --test` 107/107 (data 22, lock 19, store 19, write 25,
commit 22), `verify` skipped because HEAD has no `queue.json`.
At today's HEAD, 40a02d2b, plus the candidate and the N13 change
(`/tmp/n13-verify`): config 24, task 90, foundation 44, conductor 17,
release 14, auth 15, discord 64, extension-package 18, queue 19 with
107/107. Foundation and discord each gained one check from N13. I reran
the queue suite there after the last change; the other suites can't reach
`packages/queue`.
The canonical `.git` is unchanged: `.git/hooks` holds only samples, no
`mosaic-queue*` file, and `git config --show-scope --get-all
core.hooksPath` returns nothing (rc 1). Every run was in a `--shared`
clone under `/tmp`.
### Mutations
Each mutation went into the verify clone, the queue tests ran, and the
file was restored from the candidate. All 20 files matched the candidate
after each run. The number is how many tests failed.
| Id | Mutation | Failing tests |
|---|---|---|
| R1a | drop the owner check on unblock | 1 |
| R1b | let the owner move waiting-on-jason→in-progress | 1 |
| R1c | check the owner on block only when not queued | 1 |
| D1 | allow review with no issues | 1 |
| D2 | require `--issue` with one issue | 9 |
| D3a | take the first of several issues | 2 |
| D3b | accept an `--issue` the row doesn't list | 2 |
| D4a | never keep the previous round's issue | 2 |
| D4b | keep an issue the row no longer lists | 1 |
| D5 | schema allows a null review issue | 1 |
| E1 | ignore the evidence round | 2 |
| L1 | stat the temp again after the link | 1 |
| L2 | don't release the lock when the gate check fails | 1 |
| F1 | drop `confirmTail`'s `queue.json` fsync | 1 |
| F2 | drop `confirmTail`'s `docs/plans` fsync | 1 |
| F3 | drop the `.git` fsync after the witness rename | 1 |
| F4 | drop the `docs/plans` fsync after the view rename | 1 |
| W1 | drop the release warning on a refusal | 1 |
| H1 | check boot before host | 1 |
| G1 | drop the gate warning on a refusal | 1 |
| G2 | drop the gate warning on success | 2 |
| U1 | print the gate warning nowhere in the CLI | 1 |
R1a to E1 ran before the last lock and store changes, which touch neither
`queue.mjs` nor the tests that caught them. L1 to U1 ran on the final
candidate. L1 first survived with 0 failures, as noted under N2.

Some files were not shown because too many files have changed in this diff Show More