Files
stack/docs/plans/2026-09-12_control-board-mvp.md
T
jason.woltjeandClaude Fable 5.1 6ec253de20 Pin the mid-tool-call rule in the control board scanner (#1503)
Acceptance rule (plan page, bf641e22): a seat mid-tool-call is working,
never waiting. The scanner already met it through pi's stopReason values;
deriveState now also checks the content for a toolCall block (working),
after the error stop reasons and before "stop" (waiting). Thinking blocks
do not keep a text turn from being waiting. Three JSONL fixture tests and
three state-table cases pin the rule. Live check on the real board:
orch-01 and rev-code-01 mid-tool-call are working, velma's finished
text-only turn is waiting. Sonnet review: APPROVED.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
2026-09-12 09:29:19 -05:00

187 lines
8.2 KiB
Markdown

# Control board MVP — plan
Tracking: Gitea issue #1503 (https://git.mosaicstack.dev/mosaicstack/stack/issues/1503)
## Why
Jason is building blind. He has many agent sessions running across projects
and no single place to see what they are doing or who is waiting on him.
This plan replaces the registry line as the next thing to build. Registry
increment 3 stays parked; it does not resume by inertia. Decision record:
`MOSAIC-STACK-D-001` (Jason, 2026-09-12).
## What "done" looks like for the MVP
One web page lists my running agent sessions across projects, shows each
one's current status, and flags which ones are waiting on me.
Acceptance:
- Jason opens one page and sees every live pi agent, grouped by project.
- Each agent shows a state in plain words (working, waiting, error,
offline, idle, unknown).
- A "waiting on you" section sits at the top of the page.
- The page refreshes itself. Jason does not have to reload it by hand.
- Every state on the page comes from a status file written by the scanner,
not from a guess made in the page itself.
## Three steps
### Step 1 (today): plan, ticket, scanner
- Write this plan and open the Gitea ticket.
- Build the status scanner script with tests. Running it writes one status
file per agent under `<dataRoot>/board/`.
- This step does not touch any page or launcher. It only produces files.
### Step 2: the web page
- Build a page that reads the status files the scanner writes and shows
them grouped by project, with the waiting-on-you section on top.
- The page refreshes itself on a short timer. The small local server behind
it may re-run the scanner on each refresh so the files stay fresh. No
separate daemon is needed for this step.
### Step 3: daily use and fixes
- Jason uses the page every day for real work.
- He says what is wrong or missing.
- Fix that one thing, then repeat. No new scope beyond what he actually hits.
### Step 3 log
**2026-09-12 — liveness follows the tmux pane, not just the session.**
Jason found agents marked "waiting" that were actually dead: their tmux
session still existed but no longer ran `pi` (it had exited to a shell or
something else). The scanner now checks which program each tmux pane is
running (`tmux list-panes -s -t '=<session>' -F '#{pane_current_command}'`) and only counts an
agent alive if a pane is running `pi`. A tmux session with no `pi` pane is
now "offline" instead of "waiting".
**2026-09-12 — a "Seen" mark for rows that don't need a reply.** Jason
noticed most "waiting" rows were completion reports, not real asks, and
they kept cluttering "Waiting on you". The page now has a "Seen" button on
waiting/error rows; clicking it stores the row's `lastActivity` in
`<dataRoot>/board/seen.json` (via `POST /api/seen`) and drops the row out
of "Waiting on you" while keeping its state visible. The mark clears
itself the moment the agent writes anything new, and "Unsee" puts a row
back by hand. `seen.json` is only ever written by that click, never by a
scan.
## How the scanner decides state
The scanner reads the newest pi session log for each agent, plus whether
the agent's tmux session is still alive. It does not guess; it only reports
what the log and tmux say.
| State | Plain-words meaning |
|---------|----------------------|
| working | The agent is in the middle of a turn: thinking or running a tool. |
| waiting | The agent finished its turn. It is your move now. |
| error | The agent's last turn ended in an error, was aborted, or was cut off. Go look at it. |
| offline | There is no live tmux session for this agent right now. |
| idle | The agent is live but has not had a conversation yet. |
| unknown | The scanner could not ask tmux (missing or not answering). It does not assume the agent is alive. |
Two kinds of agents are scanned today:
- Repo agents: session logs live under `.pi/state/<agent>/sessions` in a
project checkout, on the default tmux socket.
- Fleet agents: session logs live under
`~/.mosaic/fleet/agents/<agent>/.pi/agent/sessions`, on the tmux socket
named `mosaic-fleet`.
## Where files go
- One file per agent: `<dataRoot>/board/sessions/<project>/<agent>.json`
- One summary file for the whole board: `<dataRoot>/board/index.json`
- `dataRoot` comes from `~/.config/mosaic-dev/config.json`. If that config
is missing or broken, the scanner refuses to run. It does not guess a
fallback location.
Board files are derived. They can be deleted and rebuilt at any time by
running the scanner again. They are NOT run records and they are not
evidence under the repository's write-once rules.
## Commands
Scan once and print the table:
```
node packages/control-board/src/cli.mjs scan --print
```
Start the page (step 2), then open http://127.0.0.1:7331/ in a browser:
```
node packages/control-board/src/cli.mjs serve
```
Tests:
```
node --test packages/control-board/tests/
```
## Boundaries
- No changes to any launcher script.
- No new files at the repository root.
- No secrets read, stored, or printed.
- No daemon yet. In step 1 the scanner is run by hand. In step 2 the page
may trigger it on refresh; nothing runs it on a schedule.
- No changes to `packages/mosaic`.
- This plan does not authorize push or merge beyond whatever the existing
refactor-branch plan already allows.
**2026-09-12 — a place to find what you marked Seen.** After using the
button, Jason noted the board had no way to show seen rows again. The page
now has a collapsed "Seen (N)" section between "Waiting on you" and "By
project" that lists every marked row with an Unsee button. It stays open or
closed across refreshes.
**2026-09-12 — "Hide seen" per project.** Jason suggested a checkbox like
"Hide offline" so seen rows stop cluttering the fleet table now that the
Seen section exists. Each project header has both boxes, on by default,
and the note under the table reads "N offline hidden · N seen hidden".
**2026-09-12 — Mid-tool-call rule checked and pinned.** Requested by the
professor session on Jason's behalf. Finding: the scanner already honoured
the rule, but only through pi's `stopReason` (`toolUse` and tool results
were `working`, `stop` was `waiting`; `deriveState` in `scan.mjs`). It now
also looks at the content: an assistant message with a `toolCall` block is
`working` whatever its stop reason; an error stop reason still wins; a
`thinking` block does not stop a text turn from being `waiting`. Three
fixture tests (tool call after a question-looking text, tool result last,
finished text-only turn) plus three state-table cases. Live check at
14:27Z: orch-01 mid-tool-call → working, rev-code-01 tool result last →
working, velma finished text-only turn → waiting; all 13 waiting seats had
text-only (or thinking+text) last messages, none had a tool call.
**2026-09-12 — Header count "N of N".** Jason pointed out that "fleet (38)"
sat above a table showing 13 rows. The header now reads "fleet (13 of 38)"
while a checkbox hides something and falls back to "fleet (38)" when nothing
is hidden. The note under the table stays: the header says how many, the
note says why. One static test.
**2026-09-12 — Parked until the board passes its gate.** Jason's ruling:
nothing else starts until he has used the board in place of cycling tmux.
Parked, not cancelled, and not to be resumed by inertia:
- Dewey's dashboard mockup selection (task D05, `agents/dewey/work/wui/`).
It gets chosen against a board Jason has used, not before.
- Registry increment 3 (already parked above; restated so it stays parked).
- New seat directories. The roster is frozen at its current seats until the
rails pass condition is written (phase C of the 48-hour plan).
**2026-09-12 — Rule: a seat mid-tool-call is `working`, never `waiting`.**
Added as an acceptance rule, not a fix log entry, because it is the second
way the top section can lie (the first was the dead-session case above).
If the newest session-log entry is an assistant message that contains a
tool call, or a tool result with no assistant text after it, the state is
`working` even if the last text the agent printed looks like a question.
`waiting` requires the last entry to be an assistant message whose content
is text only and whose turn has ended. Verify with one live seat caught
mid-tool-call and one that has finished a turn; add a fixture test for
each so the rule cannot regress silently.