Files
stack/agents/darkwing/work/discord-engine-busy
jason.woltjeandClaude Opus 5.5 3edb15eb96 fix(discord): engine tests stop in finally; a turn pi never started stops pi instead of guessing (#1509)
Every engine test stops its engine in finally, and commands() tolerates a
log that doesn't exist yet, so a failed assertion no longer leaks a fake pi
and hangs the six-package union. A turn that failed client-side stays at
the front of the queue and holds the next prompt. If pi has sent no
agent_start ABORT_GRACE_MS (30 s) after the failure, the engine marks
itself wedged, fails held prompts with engine-wedged, refuses new ones with
engine-down, and stops pi. The exit reaches onExit, the connector exits 1,
and the unit restarts it. Pi's events carry no prompt id, so R1's approach,
dropping the turn and sending on, let a late run answer the next prompt.
Rocko rejected R1 and approved R2.

Tests: engine 17/17 (R1 fails 4, HEAD fails 5). Union 408/408 and the eight
suites green on the committed index. Record:
agents/darkwing/work/discord-engine-busy/.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 15:50:37 -05:00
..

Discord engine: guaranteed test cleanup and the timeout gap in busy (#1509), R2 candidate

Sage assigned this on 2026-09-26 as 6b, source only. Rocko reviews. Base is HEAD 401cc850. Not committed. The live connector runs from this checkout, so Sage is holding its restart until this is approved and committed. Nobody should restart it from a working copy.

Defects (DEFERRED Open, "Discord engine: leaked fake pi…")

(a) engine.test.mjs read the fake's commands.jsonl 20 ms after a prompt and got ENOENT under load. Seven tests stopped the engine outside finally, so a failed assertion left the fake pi running and the test file never exited.

(b) busy was state.busy || pending.some((t) => !t.done). If a turn timed out before its agent_start was read, it was done while state.busy was still false. The next prompt then went straight to pi, which refused it as streaming.

R1 and Rocko's finding

R1 held later prompts behind a failed turn. If pi had sent no agent_start for it within a grace period, R1 dropped that turn from the queue and sent the next prompt. Rocko rejected it (F1, High), in agents/rocko/work/discord-engine-busy-r1-review-2026-09-26.md, sha256 047dbd8f.

Pi's events carry no prompt id. The engine attributes them to the front of its queue. Silence until the grace ends does not prove the old run will never come. If pi then runs it, its events land on the new prompt, which R1 had just put at the front. Rocko's reproducer got the old run's answer and its old.md tool record back as the new prompt's result. My R1 README said such events "find no live head and are dropped". That was wrong.

The R1 files stay here as r1-manifest.sha256 and r1.patch.

Change (R2)

packages/discord/src/engine-pi.mjs:

  • engineBusy() is state.busy || state.pending.length > 0. A failed turn still in the queue holds the next prompt back, and stays at the front, so any late events for it land on it. prompt(), sendHeld() and the busy getter use it. This part is unchanged from R1.
  • The bound is now a stop, not a drop. When a turn fails while it is still in the queue, failTurn starts a timer, abortGraceMs (default ABORT_GRACE_MS, 30 s, an engine option, not binding config). When it fires:
    • If pi has sent agent_start (state.busy), nothing happens. That run ends on its agent_end or a settle, as on HEAD.
    • Otherwise wedge() sets state.wedged, fails every held prompt with code engine-wedged, and stops pi: stdin closed, SIGTERM, then SIGKILL after 5 s. The failed turn stays at the front until the exit.
  • While wedged, nothing is written to that child. write(), sendHeld() and prompt() refuse, and a new prompt fails at once with engine-down. The exit runs the usual failAll and onExit.
  • stop()'s body moved into stopChild(), which both stop() and wedge() call.
  • release() clears the timer wherever a turn leaves the queue: agent_end, settle, a refused send, and process exit. As in R1, the settle handler removes turns before failing them.

What recovery looks like live: cli.mjs handles onExit with shutdown(1). The unit's Restart=on-failure starts a new connector and a new pi 15 s later, within its limit of five tries in ten minutes. This change doesn't touch the unit or the restart policy. A wedge now costs one connector restart. R1 would have kept the same pi and risked a wrong answer.

packages/discord/tests/fake-pi.mjs:

  • mute: accepted and never run.
  • stall <ms>: accepted, then the fake reads nothing for <ms>, runs the stalled prompt, and only then reads what came in meanwhile. This is the order in Rocko's case.

packages/discord/tests/engine.test.mjs:

  • withEngine() stops the engine in finally. Every test that starts an engine uses it, or has its own try/finally in the exit test.
  • commands() returns [] until the fake creates its log. The held-prompt test waits for the first prompt with until() instead of a 20 ms sleep, and its first prompt is slow 300.
  • The manual-timer test fires the turn timer before any pi event is read. The next prompt must wait for the settle and get its own answer.
  • New or changed for R2:
    • mute with abortGraceMs: 150. The held prompt fails with engine-wedged after the grace, a later prompt fails with engine-down, onExit fires, and pi saw only mute and abort.
    • stall 400 with the same grace, which is Rocko's case with a real child. The held prompt fails with engine-wedged, pi exits, and "after stall" never reaches pi.
    • Rocko's reproducer as an in-memory test, run twice. The old prompt's response comes either before its timeout or only with the late events. After the grace, the old run's start, tool pair, answer, end and settle arrive while pi is still exiting. The held prompt stays failed with engine-wedged and a later prompt fails with engine-down. Pi saw only old and abort, then SIGTERM, then SIGKILL at 5 s. Only the exit reaches onExit. The test reads recorded outcomes after a tick instead of awaiting, so a regression fails instead of hanging.
    • late 400 with the same grace. Pi started that run, so the grace does not stop pi, and the next prompt gets its own answer when the run ends.

Evidence

  • engine.test.mjs: 17/17.
  • R1's engine (d5bf24b5, from r1.patch) against these tests fails 4: mute, stall, and both in-memory runs. In that run a probe shows R1 answering "after stall" with "echo: stalled". An earlier draft of the in-memory test awaited the held prompt and hung on R1 until the 120 s cap. It now fails in milliseconds.
  • HEAD's engine against these tests fails 5: the manual-timer test and the same four.
  • Mutations of R2:
    • Without the state.busy check, the late 400 test fails.
    • Without the wedge() call, 4 fail.
    • Without failing held prompts in wedge(), 4 fail.
  • Test union (control-board, webui, seat, mosaic, ledger, discord) at default concurrency on git archive of 401cc850 plus the three files: 406/406 three times, 23 to 24 s each. No fake pi was left running.
  • Eight suites green on that snapshot: config 24, task 90, foundation 43, conductor 17, release 14, auth 15, discord 63, extension-package 18.

Logs: /tmp/dw-6b-r2-conc-{1,2,3}.txt. R1's evidence runs: /tmp/dw-6b-conc-{1,2,3}.txt, /tmp/dw-6b-serial.txt. HEAD's hang control: /tmp/dw-1509-headctl-{1,2,3}.txt, /tmp/dw-1509-ef00-1.txt. There, HEAD hit the 240 s cap at 199 ok under the union's load.

Not covered

  • A run pi started and never ends, even after abort, still holds prompts until pi settles or exits. Each held prompt fails at its own timeout ("while waiting for the engine"). HEAD behaves the same way through state.busy, and Rocko did not block on it. Only a pi restart clears it.
  • A wedge ends the connector process, and the recovery is systemd's restart. Nothing here changes the unit, and the restart limit still applies.
  • No live restart, and no change to the binding schema.

Frozen files

r2-manifest.sha256 holds the three R2 hashes. r2.patch is git diff packages/discord at freeze time.

Review

Rocko, R1, 2026-09-26: request changes, F1 High, as described above. Report: agents/rocko/work/discord-engine-busy-r1-review-2026-09-26.md, sha256 047dbd8f.

Rocko, R2, 2026-09-26: approved the three pinned files. Report: agents/rocko/work/discord-engine-busy-r2-review-2026-09-26.md, sha256 ed5510a0. He checked the manifests before and after, ran 17/17 himself, and read the CLI shutdown path, connector.stop and the unit template. Sage asked him three operational questions:

  • A wedge exits 1, never 3. Exit 3 remains the supervised startup refusal.
  • The unit's start limit (5 starts in 600 s) is a rate limit. It does not bound repeated wedges. With the default 180 s turn timeout, the 30 s grace and the 15 s restart delay, a cycle takes at least 225 s. That stays under the limit, so a pi that wedges every time could restart indefinitely. Stopping for good after repeated wedges would need a separate policy. This change does not add one.
  • He recommends, as a nonblocking follow-up, that the connector journal record at startup: HEAD, dirty state scoped to runtime source, and a digest of the runtime files. A wedge restart loads whatever the checkout holds.

This section was added after approval, so the README hash no longer matches the one Rocko pinned (69350f29). The three source files are unchanged.