fix(discord): engine tests stop in finally; a turn pi never started stops pi instead of guessing (#1509)

Every engine test stops its engine in finally, and commands() tolerates a
log that doesn't exist yet, so a failed assertion no longer leaks a fake pi
and hangs the six-package union. A turn that failed client-side stays at
the front of the queue and holds the next prompt. If pi has sent no
agent_start ABORT_GRACE_MS (30 s) after the failure, the engine marks
itself wedged, fails held prompts with engine-wedged, refuses new ones with
engine-down, and stops pi. The exit reaches onExit, the connector exits 1,
and the unit restarts it. Pi's events carry no prompt id, so R1's approach,
dropping the turn and sending on, let a late run answer the next prompt.
Rocko rejected R1 and approved R2.

Tests: engine 17/17 (R1 fails 4, HEAD fails 5). Union 408/408 and the eight
suites green on the committed index. Record:
agents/darkwing/work/discord-engine-busy/.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
This commit is contained in:
2026-09-26 15:50:37 -05:00
co-authored by Claude Opus 5.5
parent f55b94e866
commit 3edb15eb96
11 changed files with 1843 additions and 102 deletions
+49
View File
@@ -2997,3 +2997,52 @@ captured the proxy's real headers: Host 127.0.0.1:<port>, no Origin. One low
finding is open: the notes' wording on raw Host syntax is broader than the
code (see DEFERRED Done). The running board keeps the old code until Sage
restarts it.
## 2026-09-26: Discord engine test cleanup and the busy gap (#1509, Darkwing, reviewed by Rocko)
Before, the engine tests could leak a fake pi. Seven tests stopped the engine
outside `finally`, and one read the fake's `commands.jsonl` 20 ms after a
prompt, which got ENOENT under load. A failed assertion then left the fake
running and the test file never exited. On clean HEAD the six-package union
(control-board, webui, seat, mosaic, ledger, discord) hung at the 240 s cap at
199 ok in four of four runs. In `engine-pi.mjs`, `busy` ignored a turn that
timed out before its `agent_start` was read, so the next prompt went to a pi
that was still running and pi refused it as streaming.
After, every engine test stops its engine in `finally`, and `commands()`
tolerates a log that doesn't exist yet. A failed turn stays at the front of
the queue and holds the next prompt, so late events land on it. If pi has
sent no `agent_start` 30 s after the turn failed (`ABORT_GRACE_MS`), the
engine marks itself wedged. It fails held prompts with `engine-wedged`,
refuses anything new with `engine-down`, and stops pi (SIGTERM, then SIGKILL
after 5 s). The exit reaches `onExit`, `cli.mjs` exits 1, and the unit's
`Restart=on-failure` starts a new connector. A run pi did start still ends
on its `agent_end` or a settle, as before.
Correction: R1 handled the grace by dropping the failed turn and sending the
next prompt. Rocko rejected it (F1, High). Pi's events carry no prompt id,
so if the old run came late, its answer and tool records went to the new
prompt. With a real child, R1 answered "after stall" with "echo: stalled".
R1's README claimed late events would be dropped. That was wrong.
Evidence: `packages/discord/tests/engine.test.mjs` 17/17. R1's engine fails 4
of these tests and HEAD's fails 5. Three mutations of R2 (no busy check, no
wedge, held prompts not failed) each fail. On `git archive` of 401cc850 plus
the three files, the union passed 406/406 three times, 23 to 24 s each, with
no fake pi left running. The eight suites passed at 24/90/43/17/14/15/63/18
on that snapshot. On the committed index (base f55b94e8, which adds two board
tests), the union passed 408/408 and the eight suites passed again. Rocko approved
R2 (engine 77077b7f, tests 47a99817, fake a8e54cc3) in
`agents/rocko/work/discord-engine-busy-r2-review-2026-09-26.md`. Record:
`agents/darkwing/work/discord-engine-busy/`.
Known limit: the unit's start limit (5 starts in 600 s) does not bound
repeated wedges. The live binding has turnTimeoutSeconds 600, so one wedge
cycle takes at least 645 s, which is longer than the window. A pi that wedges
on every prompt restarts once per prompt and never hits the limit. It happens
only when someone sends a message, and they get an error reply. Sage accepted
this and is recording two follow-ups in DEFERRED: a cumulative wedge stop, and
a startup journal line with HEAD, the scoped dirty state and a runtime file
digest. Also unchanged from HEAD: a run pi started that never ends holds
prompts until each one times out. No connector restart and no push were done
here.