Files
stack/agents/darkwing/work/relaunch-activity/rerun-2026-09-26.md
T
jason.woltjeandClaude Opus 5.5 af4203ca92 feat(board): session attention, Discord rows, task attribution and relaunch activity (rows 18, 22, #1511, #1512)
One cumulative control-board, webui and seat state. The four rows edit the
same files (scan.mjs, page.html, README.md, app.js), so they land together,
each on its own receipt:

- Row 18, Discord connector rows on the board (#1509): R3 approved by
  Darkwing and Dewey, Gitea comment 26257, manifest 254403b8. Jason
  accepted the visual test.
- Row 22, board attention status (#1503): Filbert approved R1, comment
  26248, manifest e40b58ec; restart receipt 26249.
- #1511, task attribution (row 6 code phase): R2 approved by Filbert and
  Dewey, manifest d4c96395. docs/TOOLS.md carries the approved --by usage
  line (tools-usage.patch 86bcba3c).
- #1512, relaunch activity (row 6 pilot): R1 approved by Darkwing and
  Dewey, candidate manifest 47769fad. All seven source files match it.

Row 16, internal development bootstrap (#1510): the seven files outside
shared records match Filbert's R1 pins, receipt 26204 (agents/researcher/*,
scripts/test-darkwing-launch.mjs, the bootstrap plan).

packages/webui/src/public/app.js is committed at its #1512 R1 pin ce7d79a4.
The working copy holds Dewey's unreviewed return-flow candidate on top of
that, and it stays uncommitted.

Also: the four row briefs and Darkwing's evidence records under
agents/darkwing/work, including the 2026-09-26 tree manifest and the #1512
re-run against 21e3e908. Serial acceptance command: 397/397, three runs.
The failures that only show when tests run concurrently are in #1509 engine
tests, and they reproduce on clean HEAD.

Suites on the exact staged tree: config 24, task 90, foundation 43,
conductor 17, release 14, auth 15, discord 63; package union 397/397
(serial); test-darkwing-launch 5/5.

Shared records (BUILD-LOG, QUEUE, CURRENT, DEFERRED, SESSIONS, AGENTS.md,
agents/README.md) follow in Sage's records commit.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 14:54:18 -05:00

5.3 KiB

#1512 R1 candidate re-run against HEAD, 2026-09-26

Requested by Sage over T3 on 2026-09-26. Integration was held because of the row 21 engine test race, and 1685deb4 fixed that race. This record is new evidence. The R1 files next to it are unchanged.

Snapshot

  • /tmp/relaunch-activity-r1-rerun-oupRiqIC: git archive 21e3e908, plus the working packages/control-board, packages/webui and packages/seat (node_modules excluded), plus a symlink to the root node_modules.
  • CANDIDATE.sha256 in the snapshot hashes to 47769fad…, which is the R1 candidate manifest. All 7 candidate pins match.
  • 71 of the 89 R1 dependency pins are unchanged. The other 18 are packages/discord/** and scripts/test-discord.sh, and they equal HEAD. engine-pi.mjs, engine.test.mjs, fake-pi.mjs and helpers.mjs in the snapshot are byte-identical to HEAD.
  • Node v26.8.1. Test union: control-board, webui, seat, mosaic, ledger and discord tests. The count rose from 351 at R1 to 397 because rows 23 to 25 added Discord tests.

Results

Run Flags Result Log (/tmp/) sha256 prefix
A1 --test-concurrency=1 --test-timeout=15000 (R1 form) 397/397, rc 0, 19.9 s darkwing-1512-rerun-A1-bounded.txt 8b34a64a660d5948
A2 same 397/397, rc 0, 19.7 s darkwing-1512-rerun-A2-bounded.txt 5c8b96209f5b8b2c
A3 same 397/397, rc 0, 19.4 s darkwing-1512-rerun-A3-bounded.txt e9829f454bf2b63f
B1 --test-concurrency=1 397/397, rc 0, 19.0 s darkwing-1512-rerun-B1-serial-notimeout.txt bb6eaedf6250c410
C1 default concurrency 395/397, rc 1, hung until I ended a leaked child darkwing-1512-rerun-C1-default-concurrency.txt 5200188e1bcd8b15
C2 default concurrency 395/397, rc 1, hung until I ended a leaked child darkwing-1512-rerun-C2-default-concurrency.txt 1f244a4835000e31

In C1 and C2, the same two tests failed, both in packages/discord/tests/engine.test.mjs:

  • Test 205 (line 100), "a prompt while streaming is held…": ENOENT on commands.jsonl.
  • Test 208 (line 146), "tool events from a run that outlived its timeout…": engine refused prompt: agent is streaming; specify streamingBehavior.

Both runs stopped making progress after every test had finished and the after() hook had removed the temp root. The only thing left was a fake-pi.mjs child. I sent that child SIGTERM (C1 pid 3636242, C2 pid 3668771). The file then exited and the runner printed its summary. Apart from that, I didn't touch either run.

Control: clean HEAD, no #1512 files

Snapshot /tmp/darkwing-head-21e3e908-zinKkQ (git archive 21e3e908 only). Same union, default concurrency:

Run Result Log (/tmp/) sha256 prefix
1 rc 124 at the 240 s cap, 178 ok, interrupted in seat and webui files darkwing-head-default-1.txt 799acad299d0b20e
2 rc 124 at the 240 s cap, same shape darkwing-head-default-2.txt e9028b28c02ffc86
3 369/370, rc 1, test 187 = the line 146 engine failure darkwing-head-default-3.txt fcd85bd474f2c210
discord only 162/162, rc 0 darkwing-head-discord-only.txt 5f068557caee7fc7
seat + webui only 21/21, rc 0 darkwing-head-seat-webui.txt 9952b57707577a92

engine.test.mjs alone in the candidate snapshot passed 5 of 5 times, 11/11 each run (darkwing-engine-alone-{1..5}.txt).

Dewey was running packages/webui/tests/ from the canonical checkout during part of this window. That added load. It doesn't explain the HEAD control, and the two engine failures reproduce without it.

Reading

The #1512 candidate passes its R1 acceptance command 3 of 3 times at 397/397, and it passes the serial run with no timeout. Every default-concurrency failure is in committed #1509 Discord engine code, and clean HEAD fails the same way. I found nothing that implicates the seven candidate files.

Two engine defects are exposed under load. They belong to #1509, not to #1512:

  1. Test race plus leak, engine.test.mjs:100. The test reads the fake's command log 20 ms after the second prompt. Under load the fake hasn't written the file yet, so the read fails with ENOENT. The test has no try/finally, so engine.stop() never runs, the fake-pi child stays alive, and the test file never exits. That is the hang. The fix is the pattern 1685deb4 already uses elsewhere: wait with until() and stop in finally. Six other tests in the file also stop without finally.
  2. Engine race, engine-pi.mjs, seen through engine.test.mjs:146. busy is state.busy || pending.some((t) => !t.done). Suppose a turn times out before the engine has read pi's agent_start. failTurn marks it done, state.busy is still false, and busy reads false. The next prompt then goes straight to pi, and pi refuses it because it's still streaming. This fails closed: the prompt errors and nothing is misattributed. Production timeouts are long, so this should be rare there. It's still a real race. One option is to count a sent turn toward busy until pi's own end or settle event arrives, not only until the client gives up on it.

Rocko's finding C, a follow-up sent after a timeout, looks superseded. Since 1685deb4, the engine never sends a follow-up (held prompts), and the line 100 test asserts that streamingBehavior is absent. Defect 2 is the timeout race that remains.