From 8bcf079a58be0910d02141a363ac74a14495219a Mon Sep 17 00:00:00 2001 From: Jason Woltje Date: Sun, 4 Oct 2026 13:45:43 -0500 Subject: [PATCH] docs(plans): foundation direction proposal after Jason's KPI message Answers the five questions, maps the brain dump against the repo, and proposes one end-to-end slice before measures. Not ratified. Co-Authored-By: Claude Opus 5.5 --- docs/SESSIONS.md | 3 + docs/plans/2026-10-04_foundation-direction.md | 240 ++++++++++++++++++ 2 files changed, 243 insertions(+) create mode 100644 docs/plans/2026-10-04_foundation-direction.md diff --git a/docs/SESSIONS.md b/docs/SESSIONS.md index 8881e161..c443924c 100644 --- a/docs/SESSIONS.md +++ b/docs/SESSIONS.md @@ -460,3 +460,6 @@ are never rewritten or removed; corrections are new entries. 2026-10-04T06:01:02Z | Filbert (T3 Claude Code, thread 9cb9731e) | row 33 handoff (#1508) | REVIEW-REQUEST to Darkwing (round 1, comment 26648, manifest fe7da3ef, build notes agents/filbert/work/queue-33/build.md); INFO to Sage (queue commit 930d2757 not pushed, source uncommitted, rev 44 race note); waiting on Darkwing's verdict 2026-10-04T06:10:28Z | Filbert (T3 Claude Code, thread 9cb9731e) | row 33 approved (#1508) | Darkwing r1 approve, comment 26651, rev 48 (862f865a, not pushed); verify --current ok at rev 48; manifest fe7da3ef still matches the working tree; candidate left unchanged for Sage's commit; n1 (owner-second genesis case) and n2 (idle cold runs) noted, not acted on 2026-10-04T06:26:07Z | Sage (T3 Claude Code, thread 1ef1e4f8) | row 33 commit and close (#1508) | candidate 6fb50cc0 after Darkwing's r1 approval (26651); suites green except one first-run queue node failure (147/1), 10 reruns 148/148, cold-start item reopened in DEFERRED; row 33 done rev 49 +2026-10-04T06:26:40Z | Filbert (T3 Claude Code, thread 9cb9731e) | row 33 landed (#1508) | Sage committed 6fb50cc0, row done rev 49, pushed through 1c724958; n1 checked on the committed code: parseMigrationMap refuses the owner first, second or third among reviewers, so a coverage gap and not a bug, left as Sage directed; cold-start item reopened by Sage (147/1 on an export, test name lost), my closing entry stands +2026-10-04T08:11:41Z | Dewey (T3 Claude Code) | CHAT-03 I1 build, row 5 (#1507) | candidate frozen: manifest agents/dewey/work/chat-03/I1-manifest.sha256 (1404341e, 27 files), packet BUILD-I1.md; all acceptance suites green, 41/41 mutants killed after 7 kill tests; row 5 in-review round 1, request comment 26671 (rev 52); reviewers darkwing, filbert; nothing committed or pushed +2026-10-04T18:45:43Z | Sage (T3 Claude Code, thread 1ef1e4f8) | foundation direction (Jason's KPI message) | proposal docs/plans/2026-10-04_foundation-direction.md: five questions answered, brain dump mapped against what exists, Vikunja integrate-not-annex, slice 1 order; row 32 park recommended; needs Jason's ratification before any goal moves; row 5 I1 reviews pinged to darkwing and filbert diff --git a/docs/plans/2026-10-04_foundation-direction.md b/docs/plans/2026-10-04_foundation-direction.md new file mode 100644 index 00000000..02601c8e --- /dev/null +++ b/docs/plans/2026-10-04_foundation-direction.md @@ -0,0 +1,240 @@ +# Foundation direction: build the system, then measure it (proposal, 2026-10-04) + +Status: a proposal from Sage. Not ratified. It would change the +2026-09-27 goals page (`docs/plans/2026-09-27_goals-review.md`), which +Jason ratified, so it needs his ruling before any goal moves. + +## Why this page exists + +On 2026-10-04 Jason rejected the KPI line of work, verbatim: + +> Let's put some intelligence behind the question of the KPIs. Reduction +> of necessary human interaction is good. But: +> - Reduced in what way? +> - What automations do the agents handle without interaction? +> - How do the agents know where to steer? +> - How do we establish the north star with clear direction? +> - Who makes the decisions about conflicts? +> +> We are chasing the reduction of human interactions without building the +> system properly to keep the direction true. + +The brain dump in the same message, verbatim: + +> Mosaic Stack must provide separation of duties for agents. Project +> Manager, CEO, CFO, CTO, CMO, CIO, COO, etc. must all be declarable and +> have strict adherence to protocol. +> +> The stack must: +> - provide the ability to mint credentials +> - provide credentials +> - declare variables per system, business, workspace/project, agent +> - enable inter-agent communications +> - enable agent-level, proactive decision making +> - enable surfacing of outstanding decisions +> - track open tasks in a project management platform (Vikunja) with +> agentic capability to fully manage the tasks, dates, assignments, and +> all aspects of the system. Perhaps just annex the Vikunja code INTO +> Mosaic Stack?? +> - reduce the human interaction by intelligently handling things that +> don't need the human to make a decision +> - use PRDY to establish the PRD and north star +> - adapt to the needs of the user to provide options when appropriate, +> and avoid the presentation of interruptions where not needed +> - provide a meta-harness for Claude, Codex, Pi, and other providers with +> the skills, system prompt, hooks, etc. to establish the system +> +> If we are able to focus on the establishment of the system, establish +> the controls, tools, skills, system prompts, etc., we will necessarily +> address the KPIs over time. Then we can isolate an area of deficiency +> and optimize. +> +> I don't have a functional system yet for Mosaic Stack. We have been able +> to surface some level of interaction within the system, but not yet +> providing a full webUI for interaction. Actually, I don't have a cli for +> interaction either. Therefore, we should be focused on establishing the +> basal systems, interface, messaging, cli/tui, webui, task management, +> profiles, and build upon those foundations. Without a system to measure, +> we can't iterate and improve. + +Sage agrees. The ledger brief (`2026-10-04_ledger-kpi-guideposts.md`, +row 32) tries to sort Jason's chat messages by hand into "decision" and +"overhead". It has to work that way because the system has no decision +object to count. That's a missing piece of the system, not a missing +metric. + +## The five questions + +### 1. Reduced in what way? + +Jason's messages fall into three kinds. + +- **Decisions only he can make:** direction, money, authority, outward + communication, taste. These stay. They should cost less: batched, each + with a recommendation, answerable in one word, and never asked twice. +- **Admin the system should do:** launching seats, relaying between them, + finding a credential, asking for status, correcting drift, and asking + "did you actually start?". These go to zero by building the missing + parts, not by asking agents to try harder. +- **Interruptions that aren't decisions:** questions an agent could have + settled under policy. These go to zero by declaring, per role, what it + may decide. + +Fewer messages follows from those three. It isn't the target. + +### 2. Which automations run without Jason? + +That depends on the decision's class and on the role that holds it. It's +declared in policy, not left to each agent's judgment. + +| Class | Examples | Who decides | +|---|---|---| +| Routine | naming, approach, tests, docs, retries | the agent doing the work | +| Within a role, reversible | assigning and scheduling tasks, reviews, pushing reviewed work to the working branch | the role holder; logged | +| Across roles | two roles want opposite things, or one role's work blocks another's | the arbiter role named for that business | +| Gated | spend, minting credentials, outward communication, merge to main, deploy, policy or role changes, the PRD | the human | + +This table is prose in AGENTS.md today. It has to become data in each +role's file under `roles/`, so a harness hook can enforce it and the +decision inbox can sort by it. + +### 3. How do agents know where to steer? + +A chain every piece of work must sit in: PRD (objective, non-goals, +success measures) → goals → milestones → tasks. A task names the PRD +requirement or goal it serves, and a task with no parent can't be created. +That check fails closed. An agent choosing work reads up the chain. + +Today the chain is loose. The goals page leads to queue rows, and rows +lead to briefs. No row cites a requirement id, because no PRD exists. + +### 4. How is the north star set? + +With PRDY. In v1 it's a package (`v1/packages/prdy`, about 1,500 lines of +TypeScript, 20 tests) plus a skill (`v1/skills/mosaic-prdy`). It holds: +- a guided interview that asks a few questions with lettered options; +- section templates (problem, scope and non-goals, requirements, + acceptance criteria, risks, milestones, success measures); +- a lifecycle of draft → review → approved → archived; +- successor versions on conflict, so an approved PRD is never edited in + place. + +The north star of a business or project is its approved PRD's objective, +non-goals and success measures. Only the human approves a PRD. + +Mosaic Stack needs no code to start this. Sage can run the PRDY interview +with Jason by hand, using v1's template, and write the first PRD for +Mosaic Stack itself. The port to code comes later, once the shape is +proven. + +### 5. Who decides conflicts? + +Each business declares an order. A conflict inside one role's domain goes +to that role's holder. A conflict between domains goes to the named +arbiter (CEO for business questions, PM for delivery order, CTO for +technical ones). Anything in the gated class, or anything that would +change the PRD, goes to the human. Escalation happens as a decision object +in the inbox, never as a chat message. + +For Mosaic Stack today: Sage by Jason's ruling of 2026-09-26, then Jason. +When two of Jason's own rulings conflict, Sage puts both to him side by +side and doesn't pick one. + +## What exists against the brain dump + +| Need | Today | Gap | +|---|---|---| +| Declarable roles with protocol | `roles/` has `conductor-policy.json` and `researcher.json`; `agents/README.md` lists seats | No schema for business roles (PM, CEO, CTO and the rest), no decision classes, no role lock | +| Mint credentials | none | Gitea can create tokens for bot users with an admin token; nothing in the stack does it | +| Provide credentials | `scripts/gitea-api.sh` reads a token file in place; `auth.sh` for pi | Per seat by hand; a launch doesn't set a role's credential | +| Variables per system, business, project, agent | one system config, `~/.config/mosaic-dev/config.json` | No business, project or agent layers | +| Inter-agent messages | T3 messages addressed by thread id; the Discord connector (#1509) | No bus addressed by role; the thread id is the address | +| Proactive decisions | prose rules in AGENTS.md | Not data, not enforced | +| Surfacing outstanding decisions | the waiting-on-jason state in the queue; "Input needed:" lines | No decision object, no inbox | +| Task management | the queue (`packages/queue`): rows, owners, reviews, a log, a lock | Built for agents; no dates, no UI, nothing for Jason to work in; Vikunja not connected | +| PRD and north star | the ratified goals page | No PRD; PRDY exists only in v1 | +| Options vs. no interruption | none | Depends on decision classes and the inbox | +| Meta-harness | `adapters/pi`, `adapters/mock`; Claude deferred (LAYERS L4) | No Claude or Codex adapter; skills, prompts and hooks aren't generated from roles | +| CLI or TUI | `scripts/mosaic launch`, `seat task`, `queue`; per-seat TUI launchers | No front door for Jason | +| WebUI | `packages/webui` (control board screen), CHAT-03 chat in review (row 5) | No tasks, decisions or conversation view in one place | + +## Vikunja: integrate, don't annex + +Recommendation: reach Vikunja through its REST API behind a task interface +in Mosaic Stack, and keep the agent-side logic in the stack. + +- Vikunja is Go with a Vue front end, under the AGPL-3.0. Annexing it + means a Go codebase in a Node repository, every upstream release merged + by hand, and AGPL terms on the combined work if Mosaic Stack is ever + offered to others over a network. +- Its API has scoped tokens, which helps the credential plan: each role + gets a token limited to what it does. +- v1 tried the other direction, a native kanban that would replace + Vikunja (`v1/docs/requirements/native-kanban-sot.md`). Its canon review + came back NO-GO with 8 blockers before any code shipped. +- Integration can be undone. A fork can't really. + +The queue keeps doing what it does well, which is the agents' lock, review +and audit log. Whether Vikunja becomes where tasks live and the queue +mirrors it, or the reverse, is the first design question for that piece. + +## Proposed order + +Building the layers one at a time, each finished before the next starts, +risks the failure Jason named: machinery with no working system. Sage +recommends one thin path through every layer first, then widening. + +**Slice 1, the first working system.** Jason opens `mosaic` in a terminal +and talks to the PM role. The PM files a task in the tracker, linked to a +PRD requirement, and assigns it to a coder role. The stack launches that +agent with the role's contract, variables and credential. The agent works. +A gated question shows up in Jason's decision inbox in the CLI. He answers +in one word. The task is reviewed and closed, and the CLI and the WebUI +both show the trail. Every step is a logged event, so the measures come +from data and not from hand tagging. + +What slice 1 needs, in build order: + +1. **PRD for Mosaic Stack.** A PRDY interview with Jason, run by hand. It + replaces the goals page as the top of the chain. +2. **Profiles, roles and variables.** A role schema with decision classes + and a lock; variable layers for system, business, project and agent; + one business (Mosaic Stack) and four roles (PM, CTO, coder, reviewer). +3. **Decision objects and messages.** A decision record (who raised it, + class, options, recommendation, who resolved it, when) and messages + addressed by role. T3 and Discord become transports. +4. **Tasks.** The tracker interface with the Vikunja adapter, or the queue + alone if Jason picks annexing and that becomes a later piece. +5. **The `mosaic` CLI and TUI.** One front door: talk to a role, see the + inbox, tasks and agents. +6. **WebUI.** The same data through the console; CHAT-03 carries on as + part of it. +7. **Credentials per role.** Provided at launch first. Minting stays + gated and comes after slice 1. +8. **Meta-harness.** Generate the system prompt, skills and hooks for Pi + from role and variables, then add Claude and Codex adapters. + +After slice 1, each layer widens: more roles and businesses, minting, +more harnesses, PRDY in code. Measures come last and pick one weak spot at +a time, which is what Jason described. + +## What happens to work in flight + +- **Row 5, CHAT-03 I1:** continues; it's WebUI work slice 1 uses. +- **Row 32, ledger guideposts:** park it. Only Jason can park. +- **Row 7, the weekly ledger:** keep the Monday run as a raw baseline. It + costs a command, and no new ledger work. +- **Goal 2, rows 11 and 13:** the queue stays as the agents' lock and log. + Row 11's second accepted brief can be the first slice brief. +- **Goal 3, Gate E:** unchanged. +- **Goal 4, retiring the fleet:** it needs a working system to move to, + so it folds into slice 1. The relaunch idea + (`2026-10-04_project-relaunch-roles.md`) becomes the roles piece. + +## Open questions + +- Storage. The canon keeps config and run records in files. Tasks, + messages and decisions at this scale may want a database. That's an + invariant question for later, and it doesn't block the PRD interview. +- Which roles a business must declare, and which are optional. +- Whether Sage holds PM or CTO for Mosaic Stack once roles exist.