Files
stack/packages/discord/README.md
T
jason.woltjeandClaude Fable 5.1 436ba6ed6b feat(discord): systemd user service with a supervised run; brakes exit 3 and are never retried (#1509)
QUEUE row 17, MVP iteration 2. scripts/discord-service.sh renders and
installs mosaic-discord@<binding> from packages/discord/systemd/. The
unit's main process is `run --supervised`, which applies the new recover
policy first: a lock whose owner is gone is cleared and only the STOP
written for that is removed; an operator STOP or a held binding refuses
with exit 3, which RestartPreventExitStatus never retries. `recover` is
also a CLI verb. First cut used ExecStartPre and looped live, since systemd
honours the never-retry status only from the main process; replaced and
re-verified before any message traffic. Suite 40/40, 95 node tests.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-13 14:39:11 -05:00

11 KiB

discord

The Discord connector: one seat's conversation reachable from listed channels of one Discord server, chat only. Issue #1509, brief docs/plans/2026-09-13_discord-connector-pilot.md. Plain ESM, no dependencies, Node 24 or newer, built-in WebSocket and fetch.

A Discord channel is one more interface onto a seat's conversation, the same class of thing as the terminal and the WebUI session chat. The connector owns a pi --mode rpc engine today; that module is the one to swap when the CHAT-03 conversation controller exists.

Commands

scripts/discord.sh check <binding>
scripts/discord.sh run   <binding> [--supervised]
scripts/discord.sh stop  <binding>
scripts/discord.sh unlock <binding>
scripts/discord.sh recover <binding>
scripts/discord-service.sh render | install | uninstall | status <binding>

<binding> names <dataRoot>/discord/<binding>.json. The wrapper passes --repo for you; the CLI also takes --config PATH.

  • check validates the binding, the token file (0600, regular, not a symlink), the context files and the pi install; reads the bot, the guild and every listed channel over REST; opens one gateway connection, waits for READY, closes it. Nothing is sent to a channel. Close code 4014 is reported as the message-content intent not being granted in the developer portal.
  • run refuses when STOP exists or an unresolved delivery cannot be reconciled. Otherwise it starts pi, connects, and serves turns until SIGTERM, SIGINT or stop. Run it under the service unit below, or by hand in a tmux window.
  • stop writes STOP and sends SIGTERM to the process in run.lock, only when that pid is alive and both its start time and the boot id match the recorded ones; a reused pid, a pid from a previous boot, or a pid whose identity cannot be read right now is never signaled. The current turn finishes or times out, then the process exits. Remove STOP to run again.
  • unlock writes STOP, then removes a run.lock whose owner is gone (crash, reboot, a start interrupted before it published its record). It refuses while the owner is running; use stop for that. It also refuses, removing nothing, when the pid is alive and its identity cannot be established: the record predates the boot id (an upgrade over a running connector), a recorded value is not a start tick or a boot id, or /proc cannot be read right now. Once that pid is dead, unlock clears it. A record file that exists but cannot be parsed is never removed or claimed over; inspect it by hand. STOP is the gate that serializes cleanup with starts: a run re-checks STOP after publishing its record and releases itself if it is there, so nothing that starts during an unlock can hold the binding. run never reclaims a stale lock on its own; it refuses and names this command. Remove STOP to run again.
  • recover is the supervised pre-start; run --supervised performs it first, in the same process, and that is the form the service unit uses, because systemd honours a never-retry exit status only from the main process. It refuses, with exit 3 and touching nothing, while STOP is present or the binding is held by a live process, an alive pid whose identity cannot be verified, or an unreadable record. A lock whose owner is gone, or that has no record, it clears the way unlock does, then it removes the STOP it wrote for that so the run that follows can claim. It removes only a STOP that consists of the one line it wrote itself; a brake an operator wrote at any point, even during the recovery, stays and the start is refused. Nothing automatic ever removes an operator's STOP.

Exit codes: 0 ok, 1 operation failed, 2 invalid data or configuration, 3 refused by a brake (STOP present or the binding held; a supervisor must not retry), 4 usage.

Service unit

scripts/discord-service.sh install renders packages/discord/systemd/[email protected] with the repository path and the directory of node, writes it to ~/.config/systemd/user/[email protected] (temp file, then rename; --dir DIR for another place, --no-reload to skip daemon-reload) and prints the commands that follow. One instance per binding:

systemctl --user enable --now mosaic-discord@<binding>   start now and at login
systemctl --user status mosaic-discord@<binding>
journalctl --user -u mosaic-discord@<binding> -f          the log (stderr of `run`)
systemctl --user stop mosaic-discord@<binding>            SIGTERM; the turn in flight finishes; restartable
scripts/discord.sh stop <binding>                         the brake: writes STOP; the unit stays down until STOP is removed
scripts/discord-service.sh status <binding>               unit state, STOP, run.lock

What the unit does: ExecStart runs run --supervised, Restart=on-failure with 15 seconds between tries and at most five in ten minutes, and RestartPreventExitStatus=3 so a brake is never retried. A crash (any other non-zero exit, a signal, an engine that died) restarts, and the supervised run clears the dead lock on its way in. systemctl --user stop sends SIGTERM only; the connector exits 0 and no STOP is written, so the next start needs no hand. The stop timeout is 3700 seconds, the largest turn timeout a binding may set plus margin; a normal stop takes as long as the turn in flight. The unit never reads the binding or the token; run does, at runtime, as before. Surviving logout and reboot needs loginctl enable-linger. render prints the unit without writing it; uninstall refuses while an instance is active. The log goes to journald: run writes no message text to stderr, only ids, counts and state.

The binding

Deployment policy for one seat on one server. It carries Discord ids of real people, so it lives under the data root at mode 0600 and is never committed. fixtures/binding.example.json is the shape with placeholder ids; the schema is src/binding.mjs.

Field Meaning
bindingVersion 1
name, seat binding name (matches the file name) and the seat it serves
guildId, guildName, botUserId the one server and the bot identity check confirms
tokenFile absolute path to the bot token, 0600, read into memory at start, never printed or journaled
channels[] {id, name, mode}; open answers every message, mention only when the bot is mentioned; threads inherit the parent's mode
users[] {id, name}; the only authors that get a turn
engine provider, model, thinking for pi
limits turnsPerDay (200), turnTimeoutSeconds (180), replyChunkChars (1900), inboundMaxChars (4000)
context.files[] files appended to pi's system prompt in order, repository-relative and inside the repository (no absolute paths, .. or symlinks); the Discord block is added after them

Unknown keys, missing fields, wrong types, empty allowlists and a bot listed as a user all refuse with exit 2.

What happens to a message

  1. The gateway delivers MESSAGE_CREATE. authorize drops it unless the guild matches, the author is listed and is not a bot, webhook or the bot itself, the channel or the thread's parent is listed, and in mention mode the bot is in mentions (@everyone does not count). A drop is one line in drops.jsonl and no reply.
  2. The message id is appended to inbox.jsonl before anything else. On start the inbox is read back; a replayed id is dropped as duplicate. That is the restart guard.
  3. STOP, an oversize message and the daily ceiling are checked next. Over size gets one fixed line. Over the ceiling gets one fixed line per UTC day, then silence until midnight UTC; the process stays up.
  4. The prompt is an envelope, one bracketed line naming server, channel, thread, author id and message id, then the text. The system prompt says that text is data. A message that arrives during a turn is queued in pi as a follow-up, so it is neither lost nor run concurrently. As soon as the turn is admitted the connector reacts to the inbound message with eyes as a read receipt; a typing indicator follows every 8 seconds while the turn runs. A reaction Discord refuses is logged and recorded in the turn; it never fails the turn.
  5. The reply is split at 1900 characters on paragraph boundaries. Each chunk is posted with nonce and enforce_nonce: true, allowed_mentions empty, and the first chunk as a reply to the inbound message. An intent line goes to outbox.jsonl before the POST and a confirmed, refused or unknown line after it. A chunk that is not confirmed stops the rest of that reply.
  6. One write-once record per turn lands in turns/<message id>.json: ids, timing, usage, read-receipt outcome, delivery outcome, and the error on a failed turn. A failed turn posts one fixed line, never model output.

On start, every intent or unknown delivery is reconciled by sending the same nonce again; Discord returns the existing message instead of posting twice. An intent older than five minutes is outside Discord's dedupe window and is marked refused rather than re-sent, because a re-send could post a second reply. If anything is still unknown after that, run refuses to start and names the nonces.

Runtime data

<dataRoot>/discord/<binding>.json          the binding, 0600
<dataRoot>/discord/<binding>/inbox.jsonl    accepted message ids
<dataRoot>/discord/<binding>/outbox.jsonl   delivery intent and receipts, by nonce
<dataRoot>/discord/<binding>/drops.jsonl    one line per dropped or refused message
<dataRoot>/discord/<binding>/admissions.jsonl  one line per admitted turn, before the engine runs
<dataRoot>/discord/<binding>/turns/<id>.json  write-once turn records
<dataRoot>/discord/<binding>/launches/      context snapshot and sha256 per run
<dataRoot>/discord/<binding>/STOP           stop switch
<dataRoot>/discord/<binding>/notices.jsonl  once-per-day fixed lines already attempted
<dataRoot>/discord/<binding>/run.lock/      ownership directory (atomic mkdir) with owner.json {pid, start, boot}; stale ones need `unlock`
<dataRoot>/sessions/discord-<binding>/      the pi session, continued across runs

Directories are 0700, files 0600. Logs are append-only; turn records are written with O_EXCL and never rewritten.

Tests

scripts/test-discord.sh or node --test packages/discord/tests/. All offline: fake WebSocket and timers for the gateway, fake fetch for REST, a scripted stand-in for pi over stdio, a disposable data root. Groups: binding, authorization table, gateway (hello, identify, heartbeat, missed ack, op 7, op 9, close 4014), delivery and reconcile, engine (follow-up, timeout, malformed line), restart replay, stop and ceiling.

Not in this piece

Tools, repository writes, announcements, attachments, slash commands, DMs, per-thread sessions, more than one server or seat, a control-board row. Section 8 of the brief keeps the list.