From 07dd2030722b7221ed833cc951840f2141944651 Mon Sep 17 00:00:00 2001 From: Jason Woltje Date: Sun, 4 Oct 2026 00:22:56 -0500 Subject: [PATCH] docs(plans): brief, ledger guideposts beyond messages per closed issue (#1514) Co-Authored-By: Claude Opus 5.5 --- .../plans/2026-10-04_ledger-kpi-guideposts.md | 97 +++++++++++++++++++ 1 file changed, 97 insertions(+) create mode 100644 docs/plans/2026-10-04_ledger-kpi-guideposts.md diff --git a/docs/plans/2026-10-04_ledger-kpi-guideposts.md b/docs/plans/2026-10-04_ledger-kpi-guideposts.md new file mode 100644 index 00000000..9bc7ffbc --- /dev/null +++ b/docs/plans/2026-10-04_ledger-kpi-guideposts.md @@ -0,0 +1,97 @@ +# Ledger guideposts beyond messages per closed issue + +Brief for one queue row. Written by Sage on 2026-10-04 after Jason's +ruling on row 7 (lead decision 42, item 7): "We are chasing a metric that +may be useless. A good starting point perhaps, but we will need to augment +the guideposts over time." It's also the second brief toward row 11's gate. + +## Ledger: Jason's overhead per finished piece, unattended pieces, rework and wait time + +### Problem + +The ledger's one success number is human messages per closed issue +(`packages/ledger/src/ledger.mjs`, `humanMessagesPerClosedIssue`). The +2026-09-28 run gave 33.0 against row 7's gate of under 10. The number +counts every message Jason sends to a mosaic-stack T3 thread the same way. +A ruling he alone can make ("1. pass" on Gate G) counts the same as a +message he shouldn't have had to send: launching a seat, relaying between +seats, finding a credential, correcting an agent, or asking "Did you +actually start?" after Sage said it had. The goal is for Jason's time to go +to decisions, not admin, and the current number can't tell the two apart. +It also divides by Gitea issues, while work is done and accepted per queue +row. + +### Owner and reviewer + +- Owner: darkwing, who built the ledger's queue section (Piece E). +- Reviewer: filbert. + +### Files owned + +- `packages/ledger/src/ledger.mjs`, `packages/ledger/src/cli.mjs`, + `packages/ledger/src/t3.mjs`, `packages/ledger/src/queue-checks.mjs` +- a new `packages/ledger/src/guideposts.mjs` if the owner prefers to keep + this apart +- `packages/ledger/tests/` (new or changed test files and fixtures) +- `packages/ledger/README.md` (the guideposts section and the weekly + routine) + +### What ships + +All four measures cover the run's date range. Each prints next to the +current totals line, which stays unchanged. + +1. **Overhead per finished piece (the headline).** A tag file named with + `--tags FILE` lists Jason's in-range mosaic-stack messages by T3 + message id. + - Each entry is `{ "id": "", "kind": "decision" | "overhead", "row": }`. + An overhead entry also has `"why": "launch" | "relay" | "credential" | "correction" | "status" | "other"`. + - The file holds ids and tags only, never message text. The ledger + prints no transcript, and the tag file doesn't either. + - Overhead per finished piece is overhead messages divided by rows that + moved to `done` in range, read from the queue log. + - Every in-range human message id the ledger counts must be tagged. An + untagged id, a duplicate, an unknown id or an unknown kind makes the + measure `incomplete` and names the ids. It never guesses. + - Without `--tags`, the measure prints `not tagged`. +2. **Unattended pieces.** Of the rows done in range, the share with no + overhead message tagged to them. Overhead tagged `row: null` counts in + measure 1 but not here, and the output prints that count. +3. **Rework.** For rows done in range, the review rounds and the + `changes` verdicts in the queue log, per row and as a total. +4. **Waiting on Jason.** For each row that sat in `waiting-on-jason` + during the range, the hours it spent there, from the queue log's + moves. This one is reported, not scored. It shows where the queue + stalls, not a mark against Jason. + +The `--json` output carries the same fields. Tests use fixture queue logs, +fixture T3 databases and fixture tag files covering: +- every refusal above; +- a row closed and reopened in range; +- a range with no done rows, where the per-piece measures print `none` + rather than dividing by zero. + +`node --test packages/ledger/tests/` and `bash scripts/test-queue.sh` pass. + +The README's weekly routine gains one step. Sage writes the week's tag +file at `agents/sage/work/ledger/_tags.json`, commits it with the +run, and Jason corrects any tag by saying so. Sage commits the correction +as a new file, never by editing the old one. + +### Out of scope + +- Row 7's gate and whether these measures replace the current number. + That stays Jason's decision once a few weeks of runs exist. +- Automatic tagging by a model. Sage tags by hand at first, and the + ledger only checks the file. +- Corrections recorded in BUILD-LOG and fixes after a commit. Neither has + a machine-readable form yet, so they wait on that. +- Messages in other T3 projects. The ledger already scopes T3 to the + mosaic-stack project. + +### Gate + +Filbert approves through the queue (`queue review record`). Then Sage +commits the candidate, and the suites above must pass on it. Before +closing, Sage runs the ledger with `--tags` on the 2026-09-27 to +2026-10-03 week and posts that run on #1514 next to the plain run.