docs(slice1): darkwing row 38 S3 round 2 packet

B1 truncates due dates to the second; B2 records what landed when the
final read fails; tests for V9, V21 and S5. Candidate manifest
e10e30e3, 28 files over c9ef7ee4. Mutants 46/63. Live provider cases
in test-release and test-task failed on a zai usage limit (429).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
This commit is contained in:
2026-10-08 19:27:17 -05:00
co-authored by Claude Opus 5.5
parent c9ef7ee49f
commit b144a03d47
7 changed files with 7470 additions and 0 deletions
@@ -0,0 +1,28 @@
ef568876a3420597a4dd5bc3a1e9a869cb32da620856fb43404f8ef8f2df0938 packages/bus/src/broker.mjs
72b53e03e3290be3d065c5e72a929cf9f8c1eabed259c19faefe3085e91f5d3e packages/bus/src/index.mjs
c3ed2eb8c690f43e4b0383e0d50c7a7cf6c59ce14e67b09b44cad6be2bd40f20 packages/bus/src/process.mjs
265b74444523e29cca6f76c9ea7b128ddf5890145b2f713de658c0144a44211f packages/bus/src/runtime.mjs
3e449bb037d8f1693237dd576dcaecb811e146280c94f2dd5ea2024754fa35bd packages/bus/src/server.mjs
a36f3492841b2541b2e43c05dc94400397013a69230f2cf66f153ba53622290e packages/bus/tests/tasks.test.mjs
9c12aa51012a65806443c8e9597028686619e338545c13a2e75ef20e9dd0aaf7 packages/tasks/README.md
aa79e805d782fe8a566c4160f0757efcf6edd3c0d7a7145a4764ab5911725856 packages/tasks/deploy/vikunja/README.md
9f31242e34de26f28677d6b6496923d6d4ee4077cfeb31a86cd2185e67a575b3 packages/tasks/deploy/vikunja/compose.yaml
72ea76d1d73e947a0acace06642ed5dac1fdd43fddd98fcebcae2b192344f165 packages/tasks/live/README.md
343c9a65a4e28792b70986ec73f35df4fc748dfab64ce8d524bf82e5c37d1e79 packages/tasks/live/run.mjs
269161444d253da7886d83f7b2f7bd79792b6c6b798c6e72f05f69bb99ad5295 packages/tasks/package.json
97478e907b15915e75e40e41df74fdeeefdacb390b402c92674bbade7b74153f packages/tasks/src/adapter.mjs
b5580224e931198c7024326135db5bd4f660dfa42834f19e4cd3c08550845bc9 packages/tasks/src/digest.mjs
a9c494b9cfa8e95b1d401523dedc5543ea2680206fb2455f28748e6fd6b00209 packages/tasks/src/fake.mjs
883cbdd868c8d1a4f33ca5133b5324d1d8cd0b77a0293e864f0ae4e7c0bae81f packages/tasks/src/index.mjs
281029b955a64c18276bdace809e64fb74b152717ff38ccd383da9c07692f083 packages/tasks/src/startup.mjs
45298f168fa96670d3694f58bd8e644beba1c2d50c59813ec6b85695a8ad7af7 packages/tasks/src/sync.mjs
6b7e749af2e98c25153adcf3246b2776d8b28b8bdcb516c1ce876a21af6b1f32 packages/tasks/src/verbs.mjs
47fce7d5180a0a9a7d6c763f92016c0d10e2c012920d995803a9a713487b121d packages/tasks/src/vikunja.mjs
1ebf6e769e91e917c2c473b956fa3a2a10e6488e10ec9e553cf88e4999f652e2 packages/tasks/tests/adapter.test.mjs
57b7dcbe5af76bbbc08ff9552b831d80e99449d295c8f2df23a36c1ad1291a66 packages/tasks/tests/deploy.test.mjs
df4412b870a1521182080ccae51e4ff16eab9e8ba5b9b58daf93ad889710a607 packages/tasks/tests/fixtures/v2-shapes.json
99dc77c72beccd421c9ea9058d29aaac2718c2093dbbdbf5ac4480222cf3c7c9 packages/tasks/tests/shapes.test.mjs
fb233bb220a4b777d259ed8d632a3486912bde6b93696356f92edcf27cd7f700 packages/tasks/tests/startup.test.mjs
215dc6503f4235cc0b511ace3fb4726103be529ec7206d54cf3745078774b2aa packages/tasks/tests/sync.test.mjs
edb7989f22f03a5ec33f4c4c63638b7031635c39a23afd83073c976e3db96c91 packages/tasks/tests/verbs.test.mjs
1659ba927506c15382a545a57733d2223e69ed8a6f9a942e9cd979f38c93d8c8 packages/tasks/tests/world.mjs
+175
View File
@@ -0,0 +1,175 @@
# Row 38, S3: round 2
Darkwing, 2026-10-09. Issue #1520, reviewer Filbert. Round 1 review:
`agents/filbert/work/slice1-s3-review/review-r1.md` (comment 26850,
rev 183). Scope from Sage (rev 184): B1, B2, and tests for V9, V21 and
S5. Nothing in the candidate is committed, staged or pushed.
Base is c9ef7ee4. Nothing under `packages/`, `scripts/` or `tools/`
changed between bee89d10, where the candidate was built, and c9ef7ee4.
In a fresh worktree at c9ef7ee4 the patch applies and the result matches
the manifest 28/28.
## Files
- `build-r2.patch` (sha256 `71ce87e631b04d35591e2df966aae565d5f665ac2a9d5d252322ce1b0d5be69c`): the whole candidate against
c9ef7ee4, the same 28 files as round 1.
- `build-r2-manifest.sha256` (sha256 `e10e30e3d68c12177afba438ac02a06c208ea0370c7198af1f9d9cf163d637f4`) pins all 28.
- `r1-to-r2.diff`: what changed since round 1. Five files under
`packages/tasks`: `src/verbs.mjs`, `src/fake.mjs`, `README.md`,
`tests/verbs.test.mjs`, `tests/sync.test.mjs`. Nothing in
`packages/bus` changed.
- `r2-filbert-probes.txt`, `r2-mutants.txt`, `r2-gate-summary.txt`:
evidence, below.
## B1: due dates with milliseconds
Sage's ruling: truncate in `due()`, don't refuse. `due()` now keeps the
first 19 characters and appends `.000Z`. Every use of the value comes
after that: the PATCH body, the expected fields for the self snapshot,
the digest handed back, and the check an uncertain PATCH's re-read runs.
`task.create` goes through the same `due()`.
`FakeVikunja` now stores a due date to the second, the way the pinned
image does in Filbert's `r1-vk-due-probe.txt`. A write's answer echoes
the value sent, in Go's form, as the image does. The probe can't tell
whether Vikunja floors or rounds. The adapter only sends whole seconds,
so that doesn't matter to it.
The README says due dates have one-second precision, with an example.
Test: "a due date with milliseconds is written to the second" covers a
create with `.789Z` and a schedule with `.456Z`. Then:
- the fake holds whole seconds;
- the self snapshot holds `.000Z`, and its digest is the one returned;
- a tick records nothing;
- the same schedule again sends no PATCH;
- the returned digest works as `expect`;
- a PATCH whose answer is lost settles on the re-read.
Filbert's probe D1 against this tree shows no external event after the
tick and no second PATCH. `expect` set to the returned digest gets past
the conflict check and stops at `decision-not-found`, the next check,
because the probe passes a made-up decision.
## B2: a write lands and the final read fails
`handle()` counts the writes that land. The final read's refusal is no
longer passed on:
- Every write landed: the verb succeeds and records its event and a
self snapshot of `work.expect(before.fields)`. `updated` is the
compare-read's, the newest one read. Filbert offered this or a null
`updated`. Keeping the compare-read's value means `recordTask` didn't
have to change.
- Some landed and a later write failed: `write-uncertain`, and nothing is
recorded. The poll records what the tracker holds.
- None landed: the failed write's own code, as in round 1.
Filbert's fix allowed refusing `write-uncertain` or returning when every
write landed. I return. The outcome isn't uncertain: each write was
confirmed, and the record is what a successful read would have written,
apart from `updated`. A refusal would send the caller back to a verb
that then refuses `task-assigned`, which is the W1 problem again.
Probe W1 against this tree: the caller sees ok, one `task.assigned`
event, one new self snapshot, no external event on the poll, and a
retry refuses `task-assigned`.
One addition beyond Filbert's list. `create()` had the same shape: the
create, labels and relations land, the final read fails, and the
caller got `tracker-unavailable`. That reads as "nothing happened" and
invites a second create. Now, when every write landed and the final
read failed, the snapshot is built from the create's answer plus the
labels written, in the default bucket. `task.created` is recorded and the
verb succeeds. If a label or relation write failed, the verb refuses with
that write's code, as in round 1, and `task.created` is still recorded.
Tests:
- "every write landed and the final read failed: the verb succeeds and
records what it wrote" (assign, Filbert's W1 as a test).
- "some writes landed and the final read failed: write-uncertain, and
nothing is recorded" (schedule: the PATCH lands, the label write gets
403, the final read fails).
- "a create whose final read fails succeeds and records task.created".
A tick afterwards records nothing, so the fallback snapshot matches
what the poll reads.
## T1, T2, T3
- T1 (V9): "an edit between the last write and the final read shows as
external on the next poll". A UI title edit lands between the assignee
write and the final read. The self snapshot keeps the old title, and
the tick records `changed: ["title"]`.
- T2 (V21): "task.created is recorded when a later label write fails".
- T3 (S5): "the first look at a task counts only comments inside the
window". The fake's clock is moved back two hours for the create and
one comment, then forward for a second comment. The tick counts 1.
## Do the tests catch the round 1 code?
With round 1's `src/verbs.mjs` and `src/fake.mjs` under the round 2
tests, four tests fail: the due date test, both final-read tests for
`handle()`, and the create final-read test. T1, T2 and T3 pass there,
since they test behaviour round 1 already had. The mutants show they
hold it.
## Mutants
Filbert's `mutants.sh`, unchanged, all 63, on a copy of this tree:
46 of 63 killed. V9, V21, S5 and G1 are killed. Round 1 was 42/63.
The 17 survivors are round 1's survivors less those four: V12, V19, V20,
V23, V24, S3, S6, S7, S9, S10, U5, U6, U7, K1, K3, G2 and B8. Filbert
called each of them a test gap or near-equivalent, and none was in this
round's scope.
## Gate
In the worktree at c9ef7ee4 with the patch applied, run one after
another, each to its own file:
| Suite | Result |
|---|---|
| `node --test 'packages/tasks/tests/*.test.mjs'` | 51/51 |
| `node --test 'packages/bus/tests/*.test.mjs'` | 67/67 |
| `node --test 'packages/business/tests/*.test.mjs'` | 60/60 |
| `node --test 'packages/discord/tests/*.test.mjs'` | 173/173 |
| `test-auth.sh` | 15/15 |
| `test-conductor.sh` | 17/17 |
| `test-config.sh` | 24/24 |
| `test-discord.sh` | 64/64 |
| `test-extension-package.sh` | 18/18 |
| `test-foundation.sh` | 44/44 |
| `test-queue.sh` | 27/27, the canonical-root check skipped in a scratch tree |
| `test-release.sh` | 11 pass, 3 fail |
| `test-task.sh` | 90 pass, 8 fail |
The 11 failures are every case that runs a real model turn: the release
health check and the live, fork and recall cases in test-task. The zai
account had hit its 5-hour usage limit. A sandboxed `release.sh activate`
in the same tree two minutes later left a run whose stderr ends in a 429,
"Usage limit reached for 5 hour". The mock-adapter cases in test-task all
pass.
I tried a clean base at c9ef7ee4 as a control and it didn't work as one.
Docker couldn't create its compose network ("all predefined address pools
have been fully subnetted"), so every container case failed there, mock
ones included: test-task 67 pass, 31 fail, twice. The candidate tree's
network already existed. I didn't remove other networks to make room.
`r2-gate-summary.txt` has the case names and the error lines.
So the live cases are unproven this round, not failed by the patch.
Neither suite loads `packages/tasks` or `packages/bus`, nothing under
`scripts/` changed since round 1, and round 1 had 14/14 and 98/98. I'll
rerun both suites in the candidate tree after the limit resets and add
the result.
## Not in this round
Filbert's notes 1 to 14 stay as they were, except note 9 (the fake kept
millisecond due dates), which B1 closes. Note 7, `secretCheck` on the
result after a write, has B2's shape. It fails closed, and the fix is a
change to the broker's result check, which is outside this round's scope.
I'd take it with the follow-ups listed in round 1.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,335 @@
diff -ruN r1tree/packages/tasks/README.md gate2/packages/tasks/README.md
--- r1tree/packages/tasks/README.md 2026-10-08 19:13:37.288575219 -0500
+++ gate2/packages/tasks/README.md 2026-10-08 19:13:37.275629637 -0500
@@ -122,6 +122,24 @@
create can't be checked that way, so an uncertain create is reported as
`write-uncertain` and a person checks the board.
+On success the self snapshot holds the fields the verb set, not what the
+final read saw. If a person edits the task between the last write and
+that read, the next poll still records the edit as theirs.
+
+The final read can fail after the writes landed. Its refusal would hide
+them, so it isn't passed on:
+
+- Every write landed: the verb succeeds and records its event and a
+ snapshot of the fields it set. The snapshot's `updated` is the
+ compare-read's, since nothing newer was read.
+- Some landed and a later one failed: `write-uncertain`, and nothing is
+ recorded. The poll records what the tracker holds.
+- None landed: the failed write's own code.
+- A create: the create's answer, plus the labels it wrote, in the
+ default bucket. `task.created` is recorded and the verb succeeds. A
+ refusal here would read as "nothing happened" and invite a second
+ create.
+
| Verb | Who | Arguments | Event |
|---|---|---|---|
| `task.create` | pm | `title`, `request`, `requirement`; optional `description`, `due_date`, `priority`, `labels`, `relations` | `task.created` |
@@ -137,6 +155,13 @@
`vikunja:<project>/<task>`, and a ref in another project refuses with
`task-project`.
+Due dates have one-second precision. Vikunja v2.7.0 stores `due_date`
+to the second, though a write's answer echoes the milliseconds it was
+sent. The adapter drops the milliseconds before it writes, so
+`2026-12-01T09:00:00.456Z` is stored, recorded and digested as
+`2026-12-01T09:00:00.000Z`. Without that, the poll would read the bot's
+own write back as a person's edit.
+
Details that aren't obvious from the table:
- `task.create` needs `request`, the id of a `human.input` event of this
diff -ruN r1tree/packages/tasks/src/fake.mjs gate2/packages/tasks/src/fake.mjs
--- r1tree/packages/tasks/src/fake.mjs 2026-10-08 19:13:37.288781975 -0500
+++ gate2/packages/tasks/src/fake.mjs 2026-10-08 19:13:37.282873539 -0500
@@ -49,6 +49,11 @@
const bad = (detail, code) => ({ title: 'Bad Request', status: 400, detail, ...(code ? { code } : {}) });
const unprocessable = (detail) => ({ title: 'Unprocessable Entity', status: 422, detail });
const second = (ms) => new Date(Math.floor(ms / 1000) * 1000).toISOString().replace('.000Z', 'Z');
+// Vikunja v2.7.0 keeps a due date to the second; a write's answer echoes what was sent, in Go's form
+// (review r1, r1-vk-due-probe.txt). The probe can't tell floor from rounding; the adapter sends
+// whole seconds, so it doesn't matter to it.
+const dueIn = (x) => (x && x !== NULL_DATE ? second(Date.parse(x)) : null);
+const dueEcho = (x) => (x && x !== NULL_DATE ? new Date(x).toISOString().replace(/\.?0+Z$/, 'Z') : NULL_DATE);
const fine = (ms) => new Date(ms).toISOString();
const page = (list, q) => {
const per = Math.min(Math.max(Number(q.get('per_page')) || 50, 1), 50),
@@ -480,7 +485,7 @@
description: body.description ?? '',
done: false,
doneAt: null,
- due: body.due_date && body.due_date !== NULL_DATE ? body.due_date : null,
+ due: dueIn(body.due_date),
priority: body.priority ?? 0,
percent: body.percent_done ?? 0,
labels: [],
@@ -497,6 +502,7 @@
this.#tasks.set(t.id, t);
const out = this.#taskJson(t, { empty: [] });
out.related_tasks = null;
+ if ('due_date' in body) out.due_date = dueEcho(body.due_date);
out.created = fine(now);
out.updated = fine(now);
return { status: 201, body: out };
@@ -516,7 +522,7 @@
before = JSON.stringify(this.#taskJson(t));
if ('title' in body) t.title = body.title;
if ('description' in body) t.description = body.description;
- if ('due_date' in body) t.due = body.due_date && body.due_date !== NULL_DATE ? body.due_date : null;
+ if ('due_date' in body) t.due = dueIn(body.due_date);
if ('priority' in body) t.priority = body.priority;
if ('percent_done' in body) t.percent = body.percent_done;
if ('done' in body && body.done !== t.done) {
@@ -529,7 +535,9 @@
// bucket_id is echoed and not applied (probes.md, "Writes and moves").
if (JSON.stringify(this.#taskJson(t)) === before && !('bucket_id' in body)) return { status: 304 };
if (JSON.stringify(this.#taskJson(t)) !== before) this.#touch(t);
- return { status: 200, body: this.#taskJson(t, { empty: [], bucket: body.bucket_id ?? 0 }) };
+ const out = this.#taskJson(t, { empty: [], bucket: body.bucket_id ?? 0 });
+ if ('due_date' in body) out.due_date = dueEcho(body.due_date);
+ return { status: 200, body: out };
}
_remove(user, [id]) {
const x = this.#taskFor(user, id, true);
diff -ruN r1tree/packages/tasks/src/verbs.mjs gate2/packages/tasks/src/verbs.mjs
--- r1tree/packages/tasks/src/verbs.mjs 2026-10-08 19:13:37.288829710 -0500
+++ gate2/packages/tasks/src/verbs.mjs 2026-10-08 19:13:37.282945050 -0500
@@ -39,9 +39,11 @@
/^\d{4}-\d\d-\d\dT\d\d:\d\d:\d\d\.\d{3}Z$/.test(x) &&
Number.isFinite(Date.parse(x)) &&
new Date(x).toISOString() === x;
+// Vikunja keeps due dates to the second (review r1, B1), so the milliseconds are dropped before the
+// write. The snapshot, the event and an uncertain write's re-read then all hold what Vikunja holds.
const due = (x) => {
if (x !== null && !canonical(x)) refuse('invalid-request');
- return x;
+ return x === null ? null : `${x.slice(0, 19)}.000Z`;
};
const priority = (x) => {
if (!Number.isSafeInteger(x) || x < 0 || x > 5) refuse('invalid-request');
@@ -133,29 +135,36 @@
const work = plan.check(before, id);
ctx.broker.authorize(cap, verb, { decision: plan.decision, target: args.task_ref });
let failure = null;
+ let landed = 0;
try {
- for (const [method, path, body, done] of work.steps) await step(role, method, path, body, id, done);
+ for (const [method, path, body, done] of work.steps) {
+ await step(role, method, path, body, id, done);
+ landed++;
+ }
} catch (e) {
failure = e instanceof BusError ? e : new BusError('adapter-failed');
}
- let after;
+ let after = null;
try {
after = await read(id);
- } catch (e) {
- if (failure) throw failure;
- throw e instanceof BusError ? e : new BusError('adapter-failed');
+ } catch {
+ // handled below: a failed final read must not hide a write that landed (review r1, B2)
}
- // Expected fields on success, so a concurrent edit still shows as external on the next poll.
- // On a failure the snapshot is what the tracker holds now.
+ // The final read failed after some writes and not others: what the tracker holds isn't known.
+ if (!after && failure) throw landed ? new BusError('write-uncertain') : failure;
+ // Expected fields on success, so a concurrent edit still shows as external on the next poll. When
+ // every write landed but the final read failed, that is still what the writes set, and `updated`
+ // is the compare-read's. On a failure the snapshot is what the tracker holds now.
const fields = failure ? after.fields : work.expect(before.fields);
+ const seen = after ?? before;
ctx.broker.recordTask({
cap,
business: ctx.business,
- snapshots: [{ task_ref: args.task_ref, updated: after.updated, etag: after.etag ?? null, digest: digest(fields), fields }],
+ snapshots: [{ task_ref: args.task_ref, updated: seen.updated, etag: after?.etag ?? null, digest: digest(fields), fields }],
events: failure || !work.event ? [] : [{ ...work.event, subject: args.task_ref }],
});
if (failure) throw failure;
- return { task_ref: args.task_ref, digest: digest(fields), updated: after.updated };
+ return { task_ref: args.task_ref, digest: digest(fields), updated: seen.updated };
};
function prepare(verb, args, role, mine) {
const common = ['task_ref', 'decision', 'expect'];
@@ -362,8 +371,14 @@
let after = null;
try {
after = await read(id);
- } catch (e) {
- failure ??= e instanceof BusError ? e : new BusError('adapter-failed');
+ } catch {
+ // handled below, as in handle()
+ }
+ // Every write landed but the final read failed: the create's answer plus the labels written, in
+ // the default bucket. A refusal here would read as "nothing happened" and invite a second create.
+ if (!after && !failure) {
+ const fields = { ...taskFields(r.json, ctx.view.ids.todo), labels: sorted(plan.labels) };
+ after = { fields, digest: digest(fields), etag: null, updated: normal(r.json.updated) };
}
// task.created is recorded even when a later label or relation write failed: the task exists.
ctx.broker.recordTask({
diff -ruN r1tree/packages/tasks/tests/sync.test.mjs gate2/packages/tasks/tests/sync.test.mjs
--- r1tree/packages/tasks/tests/sync.test.mjs 2026-10-08 19:13:37.288953627 -0500
+++ gate2/packages/tasks/tests/sync.test.mjs 2026-10-08 19:13:37.283110988 -0500
@@ -196,3 +196,17 @@
(e) => e.code !== undefined && !e.message.includes(w.tokens.coder),
);
});
+// The first look at a task counts only comments inside the window, so a restart doesn't replay
+// every comment ever made as new.
+test('the first look at a task counts only comments inside the window', async (t) => {
+ let shift = -2 * 3600000;
+ const w = await ready(t, { fake: { now: () => Date.now() + shift } });
+ const { task_ref } = await create(w);
+ await w.ui('POST', '/tasks/1/comments', { comment: 'two hours ago' });
+ shift = 0;
+ await w.ui('POST', '/tasks/1/comments', { comment: 'just now' });
+ await sleep(20);
+ await w.adapter.tick('demo');
+ const [e] = external(w, task_ref);
+ assert.equal(e.body.comments, 1);
+});
diff -ruN r1tree/packages/tasks/tests/verbs.test.mjs gate2/packages/tasks/tests/verbs.test.mjs
--- r1tree/packages/tasks/tests/verbs.test.mjs 2026-10-08 19:13:37.288964639 -0500
+++ gate2/packages/tasks/tests/verbs.test.mjs 2026-10-08 19:13:37.283122508 -0500
@@ -1,6 +1,7 @@
import test from 'node:test';
import assert from 'node:assert/strict';
import { world, sleep } from './world.mjs';
+import { FakeVikunja } from '../src/fake.mjs';
const ready = async (t, edit) => {
const w = await world(t, edit);
await w.adapter.start('demo');
@@ -280,3 +281,129 @@
assert.deepEqual(w.fake.task(1).labels, [w.label]);
assert.deepEqual(w.fake.task(1).assignees, [w.bots.coder]);
});
+// Vikunja keeps due dates to the second (review r1, B1). The adapter drops the milliseconds before it
+// writes, so its own snapshot, the digest it returns and the next poll all agree.
+test('a due date with milliseconds is written to the second', async (t) => {
+ const w = await ready(t);
+ const pm = w.agent('pm');
+ const a = await task(w, { due_date: '2026-11-01T00:00:00.789Z' });
+ assert.equal(w.fake.task(1).due, '2026-11-01T00:00:00Z');
+ assert.equal(w.snapshots(a.task_ref).at(-1).fields.due_date, '2026-11-01T00:00:00.000Z');
+ const r = await w.call(pm, 'task.schedule', { task_ref: a.task_ref, due_date: '2026-12-01T09:00:00.456Z' });
+ assert.equal(w.fake.task(1).due, '2026-12-01T09:00:00Z');
+ const s = w.snapshots(a.task_ref).at(-1);
+ assert.equal(s.fields.due_date, '2026-12-01T09:00:00.000Z');
+ assert.equal(s.digest, r.digest);
+ await sleep(20);
+ assert.deepEqual(await w.adapter.tick('demo'), { snapshots: 0, events: 0 });
+ const patches = () => w.fake.requests.filter((x) => x.method === 'PATCH').length;
+ const n = patches();
+ await w.call(pm, 'task.schedule', { task_ref: a.task_ref, due_date: '2026-12-01T09:00:00.456Z' });
+ assert.equal(patches(), n);
+ // The digest handed back is the one the next verb has to name.
+ await w.call(pm, 'task.schedule', { task_ref: a.task_ref, due_date: '2026-12-02T09:00:00.456Z', expect: r.digest });
+ // A landed PATCH whose answer was lost settles on the re-read.
+ w.fake.fault('PATCH', /^\/tasks\/1$/, 500, { apply: true });
+ await w.call(pm, 'task.schedule', { task_ref: a.task_ref, due_date: '2026-12-03T09:00:00.456Z' });
+ assert.equal(w.fake.task(1).due, '2026-12-03T09:00:00Z');
+ await sleep(20);
+ assert.deepEqual(await w.adapter.tick('demo'), { snapshots: 0, events: 0 });
+});
+// Arms a fault on the final read: after `trigger` answers, the next GET of task 1 fails.
+const thenFailRead = (trigger) => {
+ const fake = new FakeVikunja();
+ let armed = false;
+ const fetch = async (url, init = {}) => {
+ const res = await fake.fetch(url, init);
+ if (armed && trigger(init.method ?? 'GET', new URL(url).pathname)) {
+ armed = false;
+ fake.fault('GET', /^\/tasks\/1$/, 0);
+ }
+ return res;
+ };
+ return { edit: { fake, fetch }, arm: () => (armed = true) };
+};
+test('every write landed and the final read failed: the verb succeeds and records what it wrote', async (t) => {
+ const f = thenFailRead((m, p) => m === 'POST' && p.endsWith('/tasks/1/assignees'));
+ const w = await ready(t, f.edit);
+ const pm = w.agent('pm');
+ const { task_ref } = await task(w);
+ f.arm();
+ const r = await w.call(pm, 'task.assign', { task_ref, role: 'coder' });
+ assert.deepEqual(w.fake.task(1).assignees, [w.bots.coder]);
+ assert.equal(w.events('task.assigned').length, 1);
+ const s = w.snapshots(task_ref).at(-1);
+ assert.equal(s.source, 'self');
+ assert.deepEqual([...s.fields.assignees], [w.bots.coder]);
+ assert.equal(s.digest, r.digest);
+ await sleep(20);
+ await w.adapter.tick('demo');
+ assert.equal(w.events('task.changed.external').length, 0);
+ await assert.rejects(w.call(pm, 'task.assign', { task_ref, role: 'coder' }), refusedWith('task-assigned'));
+});
+test('some writes landed and the final read failed: write-uncertain, and nothing is recorded', async (t) => {
+ const f = thenFailRead((m, p) => m === 'POST' && p.endsWith('/tasks/1/labels'));
+ const w = await ready(t, f.edit);
+ const pm = w.agent('pm');
+ const { task_ref } = await task(w);
+ const n = w.snapshots(task_ref).length;
+ w.fake.fault('POST', /^\/tasks\/1\/labels$/, 403);
+ f.arm();
+ await assert.rejects(
+ w.call(pm, 'task.schedule', { task_ref, due_date: '2026-12-01T09:00:00.000Z', labels: { add: [w.label] } }),
+ refusedWith('write-uncertain'),
+ );
+ assert.equal(w.fake.task(1).due, '2026-12-01T09:00:00Z');
+ assert.equal(w.snapshots(task_ref).length, n);
+});
+test('a create whose final read fails succeeds and records task.created', async (t) => {
+ const f = thenFailRead((m, p) => m === 'POST' && /\/projects\/\d+\/tasks$/.test(p));
+ const w = await ready(t, f.edit);
+ w.agent('pm');
+ f.arm();
+ const r = await task(w, { due_date: '2026-11-01T00:00:00.000Z' });
+ assert.equal(w.events('task.created').length, 1);
+ const [s] = w.snapshots(r.task_ref);
+ assert.equal(s.source, 'self');
+ assert.equal(s.digest, r.digest);
+ await sleep(20);
+ await w.adapter.tick('demo');
+ assert.equal(w.events('task.changed.external').length, 0);
+});
+// The success snapshot holds what the verb wrote, not what the final read saw, so an edit that lands
+// between the last write and that read is still a person's edit on the next poll.
+test('an edit between the last write and the final read shows as external on the next poll', async (t) => {
+ const fake = new FakeVikunja();
+ let w,
+ armed = false;
+ w = await ready(t, {
+ fake,
+ fetch: async (url, init = {}) => {
+ const res = await fake.fetch(url, init);
+ if (armed && init.method === 'POST' && new URL(url).pathname.endsWith('/tasks/1/assignees')) {
+ armed = false;
+ await w.ui('PATCH', '/tasks/1', { title: 'renamed in the ui' });
+ }
+ return res;
+ },
+ });
+ const pm = w.agent('pm');
+ const { task_ref } = await task(w);
+ armed = true;
+ await w.call(pm, 'task.assign', { task_ref, role: 'coder' });
+ assert.equal(w.snapshots(task_ref).at(-1).fields.title, 'a task');
+ await sleep(20);
+ await w.adapter.tick('demo');
+ const ext = w.events('task.changed.external');
+ assert.equal(ext.length, 1);
+ assert.deepEqual(ext[0].body.changed, ['title']);
+});
+test('task.created is recorded when a later label write fails', async (t) => {
+ const w = await ready(t);
+ w.agent('pm');
+ w.fake.fault('POST', /^\/tasks\/\d+\/labels$/, 500);
+ await assert.rejects(task(w, { labels: [w.label] }), refusedWith('write-uncertain'));
+ assert.notEqual(w.fake.task(1), undefined);
+ assert.deepEqual(w.fake.task(1).labels, []);
+ assert.equal(w.events('task.created').length, 1);
+});
@@ -0,0 +1,53 @@
D1 self snapshot due_date self 2026-12-01T09:00:00.000Z returned digest 1c30ad1187e1
D1 external events after one tick []
D1 same schedule again sends PATCH 0
D1 expect=returned digest -> decision-not-found
D2 external events 0
W1 caller sees ok
W1 coder assigned in tracker true
W1 action.allowed task.assign 1
W1 task.assigned events 1 self snapshots 2
W1 poll external events []
W1 caller retries assign -> task-assigned
H1 human task.assign agent-required
H2 reader task.assign read-only
H3 non-holder coder update not-holder
R1 other task id 2 schedule -> ok relations on task 1 [{"kind":"related","other":2}]
C1 reviewer (no scope.change authority) wrong expect -> task-conflict task.conflict events 1
C1 reviewer right digest -> decision-required
M1 tick -> tracker-unavailable external 0 missing 0
T1 next tick -> ok external 0 missing 1
✖ D1 schedule with milliseconds: the poll reports the bot's own due date as an external change (247.676624ms)
✔ D2 whole seconds: no external change (control) (197.418492ms)
✔ W1 write lands, final read fails: refusal, no self record, poll calls it external (205.701114ms)
✔ H1 a human, H2 a reader, and a non-holder cannot call a task verb (221.27354ms)
✔ R1 a relation to another project's task id under this project's ref (195.93214ms)
✔ C1 a conflict event costs no authority; a role without the verb can still write task.conflict (172.85982ms)
✔ M1 one task with a 500 in missing() stalls the tick; T1 the tick after recovers (221.482298ms)
ℹ tests 7
ℹ suites 0
ℹ pass 6
ℹ fail 1
ℹ cancelled 0
ℹ skipped 0
ℹ todo 0
ℹ duration_ms 1559.406301
✖ failing tests:
test at ../probe-r2/probe.test.mjs:30:1
✖ D1 schedule with milliseconds: the poll reports the bot's own due date as an external change (247.676624ms)
AssertionError [ERR_ASSERTION]: false external change
0 !== 1
at TestContext.<anonymous> (file:///home/jwoltje/darkwing-scratch/s3/probe-r2/probe.test.mjs:48:10)
at async Test.run (node:internal/test_runner/test:1409:7)
at async startSubtestAfterBootstrap (node:internal/test_runner/harness:387:3) {
generatedMessage: false,
code: 'ERR_ASSERTION',
actual: 0,
expected: 1,
operator: 'strictEqual',
diff: 'simple'
}
@@ -0,0 +1,53 @@
Row 38 S3 round 2 gate. Darkwing.
Tree: detached worktree at c9ef7ee4, build-r2.patch applied, sha256sum -c build-r2-manifest.sha256 28/28 OK.
Suites run one after another, each teed to its own file.
Start 2026-10-09T00:18:27Z, done 00:23:15Z, load 7 to 15.
suite result
node --test packages/tasks/tests/*.test.mjs 51/51
node --test packages/bus/tests/*.test.mjs 67/67
node --test packages/business/tests/*.test.mjs 60/60
node --test packages/discord/tests/*.test.mjs 173/173
scripts/test-auth.sh 15/15
scripts/test-conductor.sh 17/17
scripts/test-config.sh 24/24
scripts/test-discord.sh 64/64
scripts/test-extension-package.sh 18/18
scripts/test-foundation.sh 44/44
scripts/test-queue.sh 27/27 (canonical-root check skipped: scratch worktree)
scripts/test-release.sh 11 pass, 3 fail
scripts/test-task.sh 90 pass, 8 fail
Failures, all cases that call the live provider:
test-release: healthy activation succeeds; pointer written with valid fields;
repeat activation succeeds (log grows).
test-task: live hello task succeeds with exact marker; result.json contents are
correct; wrong expectExact (reason exit-nonzero); fork base: teach succeeds;
fork child recalls ancestor context; forked child recalled ancestor code word;
user recall run succeeds; recalled user name.
Cause. A sandboxed `release.sh activate` in the same tree at 00:25:13Z left a
run whose stderr.txt ends:
429: {"code":"1308","message":"Usage limit reached for 5 hour. Your limit will reset at 2026-10-09 10:48:02"}
The zai account had hit its 5-hour usage limit. The health check and the live
test-task cases all run a real model turn, so all of them fail with
exit-nonzero. The mock-adapter cases in the same suite pass (mock, mission,
seat, policy, workspace and user-layer cases).
Base comparison. Clean worktree at c9ef7ee4, no patch, run twice
(00:23 and 00:26):
test-release 11 pass, 3 fail, the same three cases
test-task 67 pass, 31 fail, both runs
The base fails more because Docker refused to create its compose network:
failed to create network base2_default: Error response from daemon:
all predefined address pools have been fully subnetted
gate2_default already existed from the build, so the candidate's containers
started. The same mock task run in both trees at 00:26:23Z: base exit 1
(network error), candidate exit 0, response MOSAIC_HELLO_OK.
So the base run says nothing about the live cases, and I didn't remove
other networks to make room.
Neither suite loads packages/tasks or packages/bus. Round 1's gate had
test-release 14/14 and test-task 98/98 on the same scripts. Nothing under
scripts/ changed between bee89d10 and c9ef7ee4. I'll rerun both suites in the
candidate tree after the provider limit resets and add the result.
@@ -0,0 +1,63 @@
V1 killed (2)
V2 killed (1)
V3 killed (1)
V4 killed (1)
V5 killed (1)
V6 killed (1)
V7 killed (2)
V8 killed (2)
V9 killed (2)
V10 killed (2)
V11 killed (1)
V12 SURVIVED
V13 killed (1)
V14 killed (1)
V15 killed (1)
V16 killed (1)
V17 killed (1)
V18 killed (1)
V19 SURVIVED
V20 SURVIVED
V21 killed (1)
V22 killed (1)
V23 SURVIVED
V24 SURVIVED
S1 killed (1)
S2 killed (8)
S3 SURVIVED
S4 killed (2)
S5 killed (1)
S6 SURVIVED
S7 SURVIVED
S8 killed (1)
S9 SURVIVED
S10 SURVIVED
S11 killed (1)
U1 killed (1)
U2 killed (3)
U3 killed (1)
U4 killed (2)
U5 SURVIVED
U6 SURVIVED
U7 SURVIVED
U8 killed (1)
A1 killed (1)
A2 killed (1)
A3 killed (2)
A4 killed (1)
A5 killed (1)
A6 killed (1)
K1 SURVIVED
K2 killed (1)
K3 SURVIVED
K4 killed (5)
G1 killed (2)
G2 SURVIVED
B1 killed (1)
B2 killed (1)
B3 killed (1)
B4 killed (1)
B5 killed (1)
B6 killed (1)
B7 killed (1)
B8 SURVIVED