docs(plans): foundation direction proposal after Jason's KPI message

Answers the five questions, maps the brain dump against the repo, and
proposes one end-to-end slice before measures. Not ratified.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
This commit is contained in:
2026-10-04 13:45:43 -05:00
co-authored by Claude Opus 5.5
parent 1c72495815
commit 8bcf079a58
2 changed files with 243 additions and 0 deletions
+3
View File
@@ -460,3 +460,6 @@ are never rewritten or removed; corrections are new entries.
2026-10-04T06:01:02Z | Filbert (T3 Claude Code, thread 9cb9731e) | row 33 handoff (#1508) | REVIEW-REQUEST to Darkwing (round 1, comment 26648, manifest fe7da3ef, build notes agents/filbert/work/queue-33/build.md); INFO to Sage (queue commit 930d2757 not pushed, source uncommitted, rev 44 race note); waiting on Darkwing's verdict
2026-10-04T06:10:28Z | Filbert (T3 Claude Code, thread 9cb9731e) | row 33 approved (#1508) | Darkwing r1 approve, comment 26651, rev 48 (862f865a, not pushed); verify --current ok at rev 48; manifest fe7da3ef still matches the working tree; candidate left unchanged for Sage's commit; n1 (owner-second genesis case) and n2 (idle cold runs) noted, not acted on
2026-10-04T06:26:07Z | Sage (T3 Claude Code, thread 1ef1e4f8) | row 33 commit and close (#1508) | candidate 6fb50cc0 after Darkwing's r1 approval (26651); suites green except one first-run queue node failure (147/1), 10 reruns 148/148, cold-start item reopened in DEFERRED; row 33 done rev 49
2026-10-04T06:26:40Z | Filbert (T3 Claude Code, thread 9cb9731e) | row 33 landed (#1508) | Sage committed 6fb50cc0, row done rev 49, pushed through 1c724958; n1 checked on the committed code: parseMigrationMap refuses the owner first, second or third among reviewers, so a coverage gap and not a bug, left as Sage directed; cold-start item reopened by Sage (147/1 on an export, test name lost), my closing entry stands
2026-10-04T08:11:41Z | Dewey (T3 Claude Code) | CHAT-03 I1 build, row 5 (#1507) | candidate frozen: manifest agents/dewey/work/chat-03/I1-manifest.sha256 (1404341e, 27 files), packet BUILD-I1.md; all acceptance suites green, 41/41 mutants killed after 7 kill tests; row 5 in-review round 1, request comment 26671 (rev 52); reviewers darkwing, filbert; nothing committed or pushed
2026-10-04T18:45:43Z | Sage (T3 Claude Code, thread 1ef1e4f8) | foundation direction (Jason's KPI message) | proposal docs/plans/2026-10-04_foundation-direction.md: five questions answered, brain dump mapped against what exists, Vikunja integrate-not-annex, slice 1 order; row 32 park recommended; needs Jason's ratification before any goal moves; row 5 I1 reviews pinged to darkwing and filbert
@@ -0,0 +1,240 @@
# Foundation direction: build the system, then measure it (proposal, 2026-10-04)
Status: a proposal from Sage. Not ratified. It would change the
2026-09-27 goals page (`docs/plans/2026-09-27_goals-review.md`), which
Jason ratified, so it needs his ruling before any goal moves.
## Why this page exists
On 2026-10-04 Jason rejected the KPI line of work, verbatim:
> Let's put some intelligence behind the question of the KPIs. Reduction
> of necessary human interaction is good. But:
> - Reduced in what way?
> - What automations do the agents handle without interaction?
> - How do the agents know where to steer?
> - How do we establish the north star with clear direction?
> - Who makes the decisions about conflicts?
>
> We are chasing the reduction of human interactions without building the
> system properly to keep the direction true.
The brain dump in the same message, verbatim:
> Mosaic Stack must provide separation of duties for agents. Project
> Manager, CEO, CFO, CTO, CMO, CIO, COO, etc. must all be declarable and
> have strict adherence to protocol.
>
> The stack must:
> - provide the ability to mint credentials
> - provide credentials
> - declare variables per system, business, workspace/project, agent
> - enable inter-agent communications
> - enable agent-level, proactive decision making
> - enable surfacing of outstanding decisions
> - track open tasks in a project management platform (Vikunja) with
> agentic capability to fully manage the tasks, dates, assignments, and
> all aspects of the system. Perhaps just annex the Vikunja code INTO
> Mosaic Stack??
> - reduce the human interaction by intelligently handling things that
> don't need the human to make a decision
> - use PRDY to establish the PRD and north star
> - adapt to the needs of the user to provide options when appropriate,
> and avoid the presentation of interruptions where not needed
> - provide a meta-harness for Claude, Codex, Pi, and other providers with
> the skills, system prompt, hooks, etc. to establish the system
>
> If we are able to focus on the establishment of the system, establish
> the controls, tools, skills, system prompts, etc., we will necessarily
> address the KPIs over time. Then we can isolate an area of deficiency
> and optimize.
>
> I don't have a functional system yet for Mosaic Stack. We have been able
> to surface some level of interaction within the system, but not yet
> providing a full webUI for interaction. Actually, I don't have a cli for
> interaction either. Therefore, we should be focused on establishing the
> basal systems, interface, messaging, cli/tui, webui, task management,
> profiles, and build upon those foundations. Without a system to measure,
> we can't iterate and improve.
Sage agrees. The ledger brief (`2026-10-04_ledger-kpi-guideposts.md`,
row 32) tries to sort Jason's chat messages by hand into "decision" and
"overhead". It has to work that way because the system has no decision
object to count. That's a missing piece of the system, not a missing
metric.
## The five questions
### 1. Reduced in what way?
Jason's messages fall into three kinds.
- **Decisions only he can make:** direction, money, authority, outward
communication, taste. These stay. They should cost less: batched, each
with a recommendation, answerable in one word, and never asked twice.
- **Admin the system should do:** launching seats, relaying between them,
finding a credential, asking for status, correcting drift, and asking
"did you actually start?". These go to zero by building the missing
parts, not by asking agents to try harder.
- **Interruptions that aren't decisions:** questions an agent could have
settled under policy. These go to zero by declaring, per role, what it
may decide.
Fewer messages follows from those three. It isn't the target.
### 2. Which automations run without Jason?
That depends on the decision's class and on the role that holds it. It's
declared in policy, not left to each agent's judgment.
| Class | Examples | Who decides |
|---|---|---|
| Routine | naming, approach, tests, docs, retries | the agent doing the work |
| Within a role, reversible | assigning and scheduling tasks, reviews, pushing reviewed work to the working branch | the role holder; logged |
| Across roles | two roles want opposite things, or one role's work blocks another's | the arbiter role named for that business |
| Gated | spend, minting credentials, outward communication, merge to main, deploy, policy or role changes, the PRD | the human |
This table is prose in AGENTS.md today. It has to become data in each
role's file under `roles/`, so a harness hook can enforce it and the
decision inbox can sort by it.
### 3. How do agents know where to steer?
A chain every piece of work must sit in: PRD (objective, non-goals,
success measures) → goals → milestones → tasks. A task names the PRD
requirement or goal it serves, and a task with no parent can't be created.
That check fails closed. An agent choosing work reads up the chain.
Today the chain is loose. The goals page leads to queue rows, and rows
lead to briefs. No row cites a requirement id, because no PRD exists.
### 4. How is the north star set?
With PRDY. In v1 it's a package (`v1/packages/prdy`, about 1,500 lines of
TypeScript, 20 tests) plus a skill (`v1/skills/mosaic-prdy`). It holds:
- a guided interview that asks a few questions with lettered options;
- section templates (problem, scope and non-goals, requirements,
acceptance criteria, risks, milestones, success measures);
- a lifecycle of draft → review → approved → archived;
- successor versions on conflict, so an approved PRD is never edited in
place.
The north star of a business or project is its approved PRD's objective,
non-goals and success measures. Only the human approves a PRD.
Mosaic Stack needs no code to start this. Sage can run the PRDY interview
with Jason by hand, using v1's template, and write the first PRD for
Mosaic Stack itself. The port to code comes later, once the shape is
proven.
### 5. Who decides conflicts?
Each business declares an order. A conflict inside one role's domain goes
to that role's holder. A conflict between domains goes to the named
arbiter (CEO for business questions, PM for delivery order, CTO for
technical ones). Anything in the gated class, or anything that would
change the PRD, goes to the human. Escalation happens as a decision object
in the inbox, never as a chat message.
For Mosaic Stack today: Sage by Jason's ruling of 2026-09-26, then Jason.
When two of Jason's own rulings conflict, Sage puts both to him side by
side and doesn't pick one.
## What exists against the brain dump
| Need | Today | Gap |
|---|---|---|
| Declarable roles with protocol | `roles/` has `conductor-policy.json` and `researcher.json`; `agents/README.md` lists seats | No schema for business roles (PM, CEO, CTO and the rest), no decision classes, no role lock |
| Mint credentials | none | Gitea can create tokens for bot users with an admin token; nothing in the stack does it |
| Provide credentials | `scripts/gitea-api.sh` reads a token file in place; `auth.sh` for pi | Per seat by hand; a launch doesn't set a role's credential |
| Variables per system, business, project, agent | one system config, `~/.config/mosaic-dev/config.json` | No business, project or agent layers |
| Inter-agent messages | T3 messages addressed by thread id; the Discord connector (#1509) | No bus addressed by role; the thread id is the address |
| Proactive decisions | prose rules in AGENTS.md | Not data, not enforced |
| Surfacing outstanding decisions | the waiting-on-jason state in the queue; "Input needed:" lines | No decision object, no inbox |
| Task management | the queue (`packages/queue`): rows, owners, reviews, a log, a lock | Built for agents; no dates, no UI, nothing for Jason to work in; Vikunja not connected |
| PRD and north star | the ratified goals page | No PRD; PRDY exists only in v1 |
| Options vs. no interruption | none | Depends on decision classes and the inbox |
| Meta-harness | `adapters/pi`, `adapters/mock`; Claude deferred (LAYERS L4) | No Claude or Codex adapter; skills, prompts and hooks aren't generated from roles |
| CLI or TUI | `scripts/mosaic launch`, `seat task`, `queue`; per-seat TUI launchers | No front door for Jason |
| WebUI | `packages/webui` (control board screen), CHAT-03 chat in review (row 5) | No tasks, decisions or conversation view in one place |
## Vikunja: integrate, don't annex
Recommendation: reach Vikunja through its REST API behind a task interface
in Mosaic Stack, and keep the agent-side logic in the stack.
- Vikunja is Go with a Vue front end, under the AGPL-3.0. Annexing it
means a Go codebase in a Node repository, every upstream release merged
by hand, and AGPL terms on the combined work if Mosaic Stack is ever
offered to others over a network.
- Its API has scoped tokens, which helps the credential plan: each role
gets a token limited to what it does.
- v1 tried the other direction, a native kanban that would replace
Vikunja (`v1/docs/requirements/native-kanban-sot.md`). Its canon review
came back NO-GO with 8 blockers before any code shipped.
- Integration can be undone. A fork can't really.
The queue keeps doing what it does well, which is the agents' lock, review
and audit log. Whether Vikunja becomes where tasks live and the queue
mirrors it, or the reverse, is the first design question for that piece.
## Proposed order
Building the layers one at a time, each finished before the next starts,
risks the failure Jason named: machinery with no working system. Sage
recommends one thin path through every layer first, then widening.
**Slice 1, the first working system.** Jason opens `mosaic` in a terminal
and talks to the PM role. The PM files a task in the tracker, linked to a
PRD requirement, and assigns it to a coder role. The stack launches that
agent with the role's contract, variables and credential. The agent works.
A gated question shows up in Jason's decision inbox in the CLI. He answers
in one word. The task is reviewed and closed, and the CLI and the WebUI
both show the trail. Every step is a logged event, so the measures come
from data and not from hand tagging.
What slice 1 needs, in build order:
1. **PRD for Mosaic Stack.** A PRDY interview with Jason, run by hand. It
replaces the goals page as the top of the chain.
2. **Profiles, roles and variables.** A role schema with decision classes
and a lock; variable layers for system, business, project and agent;
one business (Mosaic Stack) and four roles (PM, CTO, coder, reviewer).
3. **Decision objects and messages.** A decision record (who raised it,
class, options, recommendation, who resolved it, when) and messages
addressed by role. T3 and Discord become transports.
4. **Tasks.** The tracker interface with the Vikunja adapter, or the queue
alone if Jason picks annexing and that becomes a later piece.
5. **The `mosaic` CLI and TUI.** One front door: talk to a role, see the
inbox, tasks and agents.
6. **WebUI.** The same data through the console; CHAT-03 carries on as
part of it.
7. **Credentials per role.** Provided at launch first. Minting stays
gated and comes after slice 1.
8. **Meta-harness.** Generate the system prompt, skills and hooks for Pi
from role and variables, then add Claude and Codex adapters.
After slice 1, each layer widens: more roles and businesses, minting,
more harnesses, PRDY in code. Measures come last and pick one weak spot at
a time, which is what Jason described.
## What happens to work in flight
- **Row 5, CHAT-03 I1:** continues; it's WebUI work slice 1 uses.
- **Row 32, ledger guideposts:** park it. Only Jason can park.
- **Row 7, the weekly ledger:** keep the Monday run as a raw baseline. It
costs a command, and no new ledger work.
- **Goal 2, rows 11 and 13:** the queue stays as the agents' lock and log.
Row 11's second accepted brief can be the first slice brief.
- **Goal 3, Gate E:** unchanged.
- **Goal 4, retiring the fleet:** it needs a working system to move to,
so it folds into slice 1. The relaunch idea
(`2026-10-04_project-relaunch-roles.md`) becomes the roles piece.
## Open questions
- Storage. The canon keeps config and run records in files. Tasks,
messages and decisions at this scale may want a database. That's an
invariant question for later, and it doesn't block the PRD interview.
- Which roles a business must declare, and which are optional.
- Whether Sage holds PM or CTO for Mosaic Stack once roles exist.