diff --git a/docs/SESSIONS.md b/docs/SESSIONS.md index eeec6c12..65955576 100644 --- a/docs/SESSIONS.md +++ b/docs/SESSIONS.md @@ -448,3 +448,7 @@ are never rewritten or removed; corrections are new entries. 2026-09-27T16:38:36Z | Filbert (T3 Claude Code, thread 9cb9731e) | row 13 live review round (#1508) | fd72d268 passes verify-commit (9 paths), ledger files byte-identical to approved r2 tree; repeated acceptance run at ae2cfcb5 with filbert's token (GET only, --no-t3): 0 violations, reduced pass; approve posted as comment 26589, recorded rev 23 (filbert-13-record-1), committed 319ee133 via queue-commit.sh, not pushed; nothing under docs/plans/reviews/ 2026-09-28T13:26:12Z | Sage (T3 Claude Code, thread 1ef1e4f8) | Monday weekly ledger (row 7, row 13 gate) | queue result fail: age on rows 9, 10, 11, 13; 3 T3 seats exempt; rails 33.0 human messages per closed issue; output agents/sage/work/ledger/2026-09-28_weekly.txt (903aa0f4) 2026-09-28T13:28:00Z | Sage (T3 Claude Code, thread 1ef1e4f8) | Gate G setup (#1508) | row 31 briefed for rocko (c989dfe9); decision 41; weekly ledger posted by Darkwing as #1508 comment 26592; rocko registration rewritten by a --help probe (recorded) +2026-09-28T13:58:04Z | Rocko (native Claude Code) | row 31, owner-not-reviewer refusal (#1508) | queue.mjs refuses owner-as-reviewer in add/set reviewers/assign (defense-in-depth kept in review-record); tests added in data/review/store.test.mjs + README sentence; node --test packages/queue/tests/ 144/144, scripts/test-queue.sh 29/29; row moved briefed->in-progress->in-review r1, candidate agents/rocko/work/queue-31/build-manifest.sha256 (digest 35199daf6803db7f7872ded900911c386c0df400ad131a6e311f4078c87312cf); Gitea request-comment post failed, pre-send credential check (rocko token file not configured in this session) -- round is open in queue.json regardless; awaiting retry/credential, Filbert review record, and Sage commit per gate; nothing committed +2026-09-30T02:27:36Z | Sage (T3 Claude Code) | Paperclip review (paperclipai/paperclip, jetrich mirror at 0b12ca95) | read-only review in /tmp/paperclip-review; covers most of the coordination layer (queue, wakeups, reviewer≠executor, board, chat, Discord); defaults conflict with fail-closed invariants (local_trusted implicit admin, claude skip-permissions default, secrets in env, deletable audit, telemetry on); trial recommended pending Jason's decision; nothing committed +2026-09-30T02:48:15Z | Sage (T3 Claude Code) | Paperclip trial phase 0 (Jason's go-ahead) | repo /mnt/storage/src/paperclip-trial (local, 15e0d59): paperclipai 2026.916.1 pinned, isolated instance authenticated/private on 127.0.0.1:3170, telemetry off, unauth 403 verified; waits on Jason's board claim; no mosaic-stack source changes +2026-10-04T04:51:56Z | Sage (T3 Claude Code, thread 1ef1e4f8) | roles/PM/(Re)Launch idea, SetSpark 10-03 read (goal 4 input) | recorded docs/plans/2026-10-04_project-relaunch-roles.md from two read-only subagent passes over six SetSpark threads; not a brief, feeds goal 4 after goal 2 diff --git a/docs/plans/2026-10-04_project-relaunch-roles.md b/docs/plans/2026-10-04_project-relaunch-roles.md new file mode 100644 index 00000000..cd98eec1 --- /dev/null +++ b/docs/plans/2026-10-04_project-relaunch-roles.md @@ -0,0 +1,100 @@ +# Project relaunch with roles, not static identities (idea, 2026-10-04) + +Status: an idea Jason raised, recorded by Sage. Not a brief and not +scheduled. It feeds the goal 4 (fleet retirement) brief once goal 2 +closes. + +## The idea + +Jason, in the SetSpark project-manager thread on 2026-10-03, 23:41 CDT: + +> It would almost be better to have roles in the project, launch a +> project-manager agent, then have the PM launch other sessions with the +> proper harness and have them volunteer for, and assume, a role in the +> project. + +And to Sage: should Mosaic Stack, its harness and a "(Re)Launch" command +for a project be built this way? The roster holds the information, the +agents' needs are known, and the PM launches the other agents. Voluntary +collaboration follows from that. + +## What Mosaic Stack already has + +- A roster (`agents/README.md`) listing each seat's job, harness and + model, and a `launch.sh --fresh` for every seat. +- Role authority in `roles/`, changed only by reviewed commits. +- Claims: `queue next` names a piece, and `queue move ... in-progress` + claims it under a lock and a log. +- Registrations: a launch writes `/seats/repo//registration.json`. + +Gate G (2026-09-28) was the first end-to-end test. Jason launched a fresh +Rocko session with no history. It ran `queue next`, found row 31, read the +brief and finished the work, with no other message. + +## What broke, and what this model would fix + +- **Launching.** Only a human can start a seat. The launchers open a TUI + that Sage can't drive from T3, so Jason had to do it (decision 41). A + PM that launches sessions removes that step. T3 has a thread-creation + API, and SetSpark's PM noted the same thing. +- **Credentials.** The fresh Rocko couldn't post its review request + because nothing set `MOSAIC_GITEA_CREDENTIAL_FILE`. A launch command + that reads the roster sets each role's environment. +- **Identity.** A seat's name lives in a thread title and in people's + memory. On 10-03, SetSpark showed what that costs: + - one seat named itself from a commit author; + - two seats shared one memory folder, and one had to guard itself + against adopting the other's identity note; + - seat names drifted ("implementer", "implementation-agent", + "setspark-implementer"); + - one seat guessed a peer's thread id; + - Jason asked "who is the PM?" before launching one. + +## What SetSpark's 2026-10-03 sessions showed + +This comes from two read-only subagent passes over six threads, not +from the repo. +- **Nobody volunteered.** Roles came from Jason's first prompts and from + the PM assigning tasks. The PM launched no sessions. +- **Agents used founder credentials.** Two seats pushed with Jason's + GitHub token and connected with his personal SSH key until Business + Operations caught it. +- **Concurrent writes clashed.** A reviewer took over a task 15 seconds + after the implementer moved it to Review, two seats edited one task + description at the same time, and messages crossing led to a review of + an outdated PR head. +- **Admin work fell on Jason all evening:** apps, team membership, buckets + and approvals. He said he was "constantly" babysitting. +- **A restart kept the conversation, not the role.** The restart probe + showed that a stopped T3 session restarts its tool process and keeps its + conversation, because the harness resumes it. The thread id and a + claimed role aren't tied to anything that survives. + +## Constraints Sage would put on it + +1. **A role has one holder.** Claiming a role takes a lock, logged the way + queue claims are. Two sessions never hold the same role. +2. **Claiming a role grants exactly that role.** A session gets the + contract in `roles/` and its scoped credential, nothing more. Roles + come only from the roster, and no session can invent one. +3. **Credentials belong to the role, not the founder.** Each role gets its + own token file in place. A session that finds only founder credentials + stops. +4. **The founder authorizes launching.** Which harnesses, how many sessions + at once, a spend limit and a stop command. The PM launching sessions + changes who starts work, which today is Jason, so it needs his explicit + ruling. +5. **Identity is the role plus the run.** A registration records which + session holds which role and since when. Releasing or ending the run + frees the role. + +## Open questions for the brief + +- Is the PM a role like any other (Sage today), or the launcher itself? +- Does a session pick its role, or does the PM assign one and the session + accept it? Jason's word was "volunteer". +- Which harnesses are in scope first: T3 threads only, or Pi headless and + the launchers too? +- A Paperclip trial is open (SESSIONS 2026-09-30, `/mnt/storage/src/paperclip-trial`). + The review found it covers much of this coordination layer, but its + defaults conflict with fail-closed. Compare it before building our own.