agenthropic

Testing & quality

This page is the QA reference for agenthropic: how the golden real-session fixture corpus is captured, promoted, and labeled with ground truth; the three P0 release-blocker tests that must be green before Phase 3 is considered done; the 12-scenario negative-test catalogue; and the coverage gate — specified at >90%, shipped at 100 — that is live from Phase 1, not a hardening pass bolted on at the end. The key takeaway up front: this project treats test infrastructure as a first-class, phase-spanning engineering problem, not overhead — “you cannot unit-test ‘the dashboard’; the four real units are ingest correctness, tree correctness, cost correctness, live-flow correctness, each with a distinct harness. The golden real-session fixture corpus is the #1 QA investment — without it, >90% coverage is high-coverage tests of synthetic happy-paths (false safety)” (concept-analysis-v2 §4.3). Everything below traces to development-plan.md Track X (WP-X1…WP-X5) and the quantified acceptance criteria in concept-analysis-v2.md §6.

Update — 2026-08 (as built). This page was written before a single test existed, so the sections below still speak in the future tense about gates that have since been built, superseded by something stricter, or blocked on a human act that has not happened. This note is the verified state of the suite as of 2026-08-15, measured by running pnpm -r --workspace-concurrency=1 run test on a clean tree and reading the per-package coverage/coverage-summary.json it writes — with the merge-gating bullet re-verified on 2026-08-25, the day main became branch-protected. Where a section below disagrees with this note, this note is the current truth and the section is the historical intent.

1. Four units, not “the dashboard”

The QA lens’s starting move is refusing to treat “test the dashboard” as one problem. It decomposes into four real units, each needing its own harness because each fails in a structurally different way (concept-analysis-v2 §4.3):

Unit What it proves Fails when
Ingest correctness A hook event and a JSONL line for the same fact collapse into one events_raw row; nothing is lost or double-counted. Dual-write dedup is wrong, or a crash mid-ingest drops data.
Tree correctness The reconstructed subagent hierarchy in orchestration_edges matches reality. A parent→child edge is missing, mis-attributed, or the tree is only “almost correct” — which the QA lens judges worse than no graph because it manufactures false trust (concept-analysis-v2 §4.3).
Cost correctness Every displayed dollar traces to ground-truth tokens × a dated, versioned price, including across a compaction. A stale price, a silent zero-cost default, or a PreCompact repricing bug.
Live-flow correctness The realtime status board reflects reality without a permanent false “working” state. A missing SubagentStop never resolves to a watchdog “unknown”.

On top of these four units, the project’s consolidated test model layers a shape and a priority scheme (concept-analysis-v2 §4.3):

This is scoped as a project-wide, always-on workstream, not a phase: the implementation plan names it WS-Test — “golden real-session fixtures corpus (happy + pathological: deep nesting, missing Stop, mid-session compaction, two concurrent instances); reconciliation/idempotency/compaction/security test suites; CI coverage gate at >90%, blocking merges” — one of four cross-cutting workstreams that “run through every phase” (implementation-plan.md §B.1), alongside security, docs, and ops.

2. The golden real-session fixture corpus

Synthetic fixtures cannot stand in for this corpus: a hand-written happy-path JSONL transcript cannot prove the tree survives a crash it was never built to have, so the corpus has to be captured from real Claude Code sessions, not authored. Capture and permanent promotion are two different work packages in two different phases:

Phase 0 (throwaway spike)                Phase 1 (permanent corpus)
──────────────────────────               ──────────────────────────
WP-S1 paired-capture harness    ──────►  WP-X1 golden fixture corpus
 · ≥3 real sessions                       · promotes the Phase-0 capture
 · paired JSONL + hook log per            · three artifacts: raw, redacted,
   session                                  manifested
 · Ivan-labeled expected tree             · ≥3 real sessions; all four
   per session                              pathologies each represented
 · throwaway hook block                   · a "manifest self-test" — CI
   reverted after capture                   asserts pathology coverage from
                                             the manifest, not a comment

(development-plan §5, WP-S1, WP-X1.) WP-X1 depends on WP-S2, WP-S4, and WP-F1 — the two Phase-0 probes that already exercise the corpus, plus the monorepo scaffold — and is scheduled at wave 6, labeled there as “corpus promotion” (development-plan §4, wave 6). It is not stood up in isolation: WP-X1 runs alongside WP-X5 (the coverage gate config) and WP-X7 (the docs-site build) in the same wave, all gated on the Phase-0 GO/CONDITIONAL-GO verdict (WP-S7) that unblocks WP-F1 in the first place (development-plan §1, §4).

The three tiers

WP-X1’s own Done-when names three artifacts, not one (development-plan §5, WP-X1):

The four pathologies

Both WP-S1’s capture harness and WP-X1’s promoted corpus are required to represent the same four pathological session shapes (development-plan §5, WP-S1; concept-analysis-v2 §6):

# Pathology Why it’s in the corpus
1 Crashed, no Stop Proves the missing-SubagentStop → watchdog “unknown” rule, and that a crash doesn’t leave a permanent false “working” state.
2 Deep nesting Stresses the DAG-rebuild-from-JSONL path against multi-level subagent-of-subagent chains, not just a flat parent→child pair.
3 Mid-session PreCompact The one EXPANDED’s own negative-test catalogue omitted — proves the tree and the cost baseline both survive a context compaction (concept-analysis-v2 §4.3, “its catalogue dropping the compaction case”).
4 Two concurrent instances Proves instance/host_id correctly partitions two simultaneously running Claude Code sessions instead of merging their trees.

The corpus size floor is ≥3 real sessions (development-plan §5, WP-S1/WP-X1; concept-analysis-v2 §6), and it doubles as the population the ≥95% hierarchy correctness gate is measured against in Phase 3 (concept-analysis-v2 §6) — see ingest & reconciliation §10 for that gate in full.

3. Labeled ground truth: expected/*.json and the typed loader

A corpus of real sessions is only useful to CI if there is a machine-checkable answer key. WP-X2 — labeled ground-truth annotations + typed fixture loader — is that answer key: “every session has an expected/*.json; a test fails if any lacks one” (development-plan §5, WP-X2). Two things make this more than a convention:

Everything downstream — the three P0 tests (§4), the negative catalogue (§5), and the ≥95% hierarchy gate (§2) — is scored against these expected/*.json files, which is why WP-X2 sits on the critical dependency edge into both WP-X3 and WP-X4 (development-plan §5).

3.1 As built: the annotation corpus, the Wilson floor, and a gate that refuses to sign

What shipped is not expected/*.json but something with the same job and a stricter posture about its own authority. The ground truth lives in packages/test-fixtures/annotations/, in a hand-writable markdown format that is diffable and needs no tooling to author: a ## meta block declaring the session, provenance, substrate, labeled-by and labeled-on, then a ## edges block of one line per subagent — <child hex> <- ROOT | ORPHAN | UNKNOWN | <parent hex>, with an optional trailing comment. Everything outside those two blocks is prose the loader ignores, so a labeler can leave notes to themselves in the file. The loader, validator, scorer, report renderer and read-only filesystem adapters are in packages/test-fixtures/src/annotations/; the runner that wires the real parser to them is packages/core/test/hierarchy-gate.test.ts.

Three design choices in that tooling are the point of it, and each exists to stop a number from meaning less than it appears to.

The score is a Wilson lower confidence bound, not a ratio. A naive percentage is silent about sample size — 3/3 and 300/300 both read “100%”, and only one of them is evidence. wilsonLowerBound() computes the one-sided 95% lower bound instead, which returns 0 when there are no observations at all: no data, no confidence. The direct consequence is a hard floor on the sample. Solving n / (n + z²) ≥ 0.95 gives n ≥ 0.95 · 1.6449² / 0.05 = 51.4, so 52 labeled agents is the minimum at which even a flawless run can clear the bar, and roughly 90 are needed to survive a single error. minimumClaimsForThreshold() computes that floor and certifyExitGate() refuses to certify below it no matter how good the raw percentage looks. The two prepared templates — b24be30c (42 agents, dual on-disk layout, deepest observed nesting) and f28af3fd (18 agents, an independent depth-2 population, 5 compactions) — total 60, chosen to clear 52 with a little headroom while covering structurally distinct ground. Labeling only one of them leaves the sample below the floor, and the gate says so.

Provenance is enforced structurally, not by convention. annotations/synthetic/ holds eight annotations, one per fixture, that state the hierarchy each fixture was built to have. They are genuinely useful — they prove the loader, the scorer, the report and every join path work end to end, and they regression-guard the depth-2 case — but they were written by the same side as the parser, so agreement with them proves internal consistency and nothing else. Every annotation must therefore declare provenance: human or provenance: synthetic-by-construction; scoreCorpus() throws if a corpus mixes the two, so a blended figure cannot be produced by accident; certifyExitGate() hard-refuses any non-human corpus and prints the reason; and the report prints an ADMISSIBILITY banner above the numbers so no reader can quote the figure without also reading what it is made of.

Abstention is a first-class answer. UNKNOWN is not scored as a miss and not scored as agreement — it is excluded from the accuracy fraction and reported in its own bucket, because a guess that turns out wrong is strictly worse than an abstention when the exit gate would be signed against it. ORPHAN, by contrast, is a positive claim (“there is no parent to find here, and a parser that invents one is wrong”) and is scored. To stop abstention from becoming a way to launder a number, the report prints label coverage and a worst-case figure — every abstention assumed wrong — next to the headline, so an under-labeled corpus cannot masquerade as a passing one.

What the gate reports today. annotations/human/ is empty. Running pnpm --filter @agenthropic/core exec vitest run test/hierarchy-gate.test.ts therefore prints SUBSTRATE UNAVAILABLE - not measured for each missing substrate and Phase-3 exit clause: NOT MEASURED - no hand-labeled sessions, points at packages/test-fixtures/annotations/README.md, and returns NOT CERTIFIED. The test itself passes in that state, and the code says why in a comment at the branch: nothing to certify on this machine — do not manufacture a verdict. This is the distinction the whole subsystem is built around. A build that fails because a human has not done a manual task teaches a team to route around the check; a build that green-washes an unmeasured gate is worse. Passing while loudly reporting n = 0 and refusing to certify is the only option that is both honest about the state and honest about the number.

Until those templates come back filled in, every hierarchy-accuracy figure anywhere in this corpus is PROVISIONAL and the Phase-3 exit clause is unmet — not failed, unmeasured.

4. The three P0 release-blocker tests

Three tests are named release-blockers: they must be green and merge-blocking in CI before Phase 3 — projection, the DAG moat, reconciliation, cost — can be considered done, and no other feature work substitutes for them (development-plan §3, Phase 3 exit gate; concept-analysis-v2 §4.3). (As built: green, and merge-blocking since 2026-08-25 for anyone who is not the repository owner — read “merge-blocking” here per the note at the top of this page.) The QA lens calls the third one “the make-or-break test both externals omit” (concept-analysis-v2 §4.3):

# Test Proves
1 Σ token_usage == JSONL exact, per session The ground-truth-tokens invariant holds in the projected data, not just the raw log — zero drift, no silent rounding, no double count.
2 Double-replay → byte-identical DB state Replay-on-startup is deterministic and safe to run on every process start, not just theoretically pure.
3 DAG-rebuild from JSONL alone, after a simulated outage The persisted orchestration_edges tree survives the exact failure mode it exists to survive.

(concept-analysis-v2 §4.3, §6; development-plan §5, WP-X3: “Three P0 reconciliation release-blocker tests. Σtoken_usage==JSONL exact; double-replay byte-identical; DAG-rebuild-from-JSONL-alone.”) Ownership is split across two work packages by design: WP-X3 (QA/fixture track, deps X2, IN10, IN7, D1) owns the test bodies run against the golden corpus from §2–3; WP-IN13 (deps IN10, IN9, X1) wires them into CI as “Reconciliation / idempotency / DAG-rebuild suite (P0 blockers). All three P0 tests green in CI and blocking” (development-plan §5, WP-IN13). Both land at wave 16, alongside the operator alerts endpoints — the last thing on the moat’s own critical sub-chain before release (development-plan §4, wave 16 and the “moat spine” note).

Two adjacent, non-P0 bars sharpen what “passing” is allowed to mean, and both are scored against the same corpus:

The full mechanics of why these three tests are structured this way — the events_raw substrate, idempotency keys, replay-on-startup, and the contingent outbox — are the architecture-level deep dive on ingest & reconciliation §10; this page is the QA-process view: who owns the test body, what corpus it runs against, and where it sits in the release gate.

5. The 12-scenario negative-test catalogue

WP-X4 — expanded negative-test catalogue (12 scenarios) — requires that “each maps to a CD/acceptance criterion” (development-plan §5, WP-X4). Its composition is explicit in the consolidated test model (concept-analysis-v2 §4.3):

WP-X4 depends on WP-X2 (the labeled corpus + typed loader, §3), WP-IN6 (the pure Normalizer), and WP-C4 (compaction-baseline repricing), and is scheduled at wave 14, alongside the reconciliation/backfill and session-tree endpoints (development-plan §4, wave 14; §5, WP-X4). That dependency set is itself informative about the catalogue’s shape: it has to exercise the Normalizer’s handling of malformed or unrecognized input (via IN6) and the cost engine’s repricing path (via C4), not only the reconciliation logic the three P0 tests already cover.

The base 10 of the catalogue are now recovered from the source in §5.1 (LOST-7); the table below is the complementary CD-anchored view — the acceptance criteria elsewhere in concept-analysis-v2 §6 and the CD register (§3) name the concrete, individually testable conditions the catalogue has to trace to, each the kind of scenario “mapped to a CD/acceptance criterion” that WP-X4’s Done-when requires:

Traceable condition CD / WP Category
An unknown event_type is stored, not crashed CD-2; development-plan §3, Phase 2 exit gate Ingest robustness
Kill+restart resumes JSONL tail-follow at the persisted offset, zero loss/dup development-plan §3, Phase 2 exit gate Ingest robustness
Missing SubagentStop → explicit “unknown” within the watchdog window concept-analysis-v2 §6; WP-IN12 Live-flow honesty
A PreCompact session reprices correctly against its preserved baseline concept-analysis-v2 §6; WP-C4 Cost correctness
A fixture model with no price row FAILS CI concept-analysis-v2 §6; WP-C6 Cost correctness
events_raw exposes no UPDATE/DELETE path CD-4, CD-7; concept-analysis-v2 §6 Data integrity
Server fails startup when DASHBOARD_TOKEN is unset CD-7 Security
No-spawner grep/static gate fails the build on a child_process import CD-7; WP-F5 Security
An SSRF test proves no outbound dial to a payload-supplied URL CD-7; WP-F5 Security — static half live since 2026-09-26 (gate:spawner refuses outbound primitives in server-process code, tested in apps/server/test/scripts-gates.test.ts); the dynamic test waits for a dispatcher to exist
SSE rejects a cross-origin connection CD-5, CD-7 Security

This table is illustrative of the categories and traceable anchors the 12-scenario catalogue draws from in this project’s own decision set — it is not a claim that these are the literal, final 12 entries. The literal base 10 originate in the external report’s §7.1; they are now recovered verbatim in §5.1 below (finding LOST-7).

5.1 The recovered EXPANDED §7.1 base catalogue (10 scenarios) — LOST-7

The ten literal scenarios below are the external EXPANDED report’s own §7.1 negative-test catalogue, recovered from the EXPANDED external report (internal source material, not published in this repo) per corpus-audit finding LOST-7 (corpus-audit-2026-07-06.md §4.3). They are the base 10 of WP-X4’s 12-scenario catalogue; the remaining two (compaction-mid-session and PreCompact re-pricing) are the project-specific additions described in §5 above (concept-analysis-v2 §6). The expected-behavior column is quoted from the source; the maps to column is this project’s decision anchor. Two scenarios were hardened as they crossed into the plan of record — flagged inline.

# Scenario (EXPANDED §7.1) Expected behavior (source) Maps to (CD / NFR / WP)
1 Duplicate hook event No duplicate normalized event or double token total. CD-2 idempotency-keyed events_raw; NFR-DATA-02; P0 double-replay test (§4).
2 SubagentStop arrives before SubagentStart Raw event stored; normalized anomaly flagged; UI shows uncertain edge. CD-2/CD-3; anomaly + watchdog state (WP-IN12); relates to OPEN-2 ('unknown' is in the agents.status CHECK since migration 4, which closed it).
3 Missing parent id Agent marked orphan / pending reparent; no fake root unless explicitly synthesized. CD-4 self-ref parent_agent_id, orphan-safe self-FK (WP-D6); NFR-DATA-01.
4 Malformed JSON 400 error; no DB mutation except optional audit log. CD-2 ingest schema validation (FR-01). Distinct from the §5 anchor “an unknown event_type is stored, not crashed” — this one rejects, that one accepts-and-stores.
5 Unauthenticated POST 401/403; no raw event stored. CD-7 mandatory DASHBOARD_TOKEN; NFR-SEC-02.
6 Invalid-token timing-attack attempt Timing-safe compare; no differentiated error leak. CD-7 timing-safe DASHBOARD_TOKEN compare; NFR-SEC-02.
7 Connection from a foreign origin (hardened: WS → SSE) Rejected. CD-5 same-origin SSE (transport corrected from the source’s “WebSocket” per CD-5/ADR-0007), CD-7. Overlaps §5’s “SSE rejects a cross-origin connection”.
8 Unknown model pricing (hardened) Token counts shown; cost marked unknown/estimated, not silently zero. CD-3/CD-4; the “estimated, never silent zero” rule governs runtime display. Project hardening: at fixture/build time a model with no price row FAILS CI (WP-C6) — a stricter gate the source did not have; both hold, at different times.
9 Huge payload Rejected or truncated per policy; no UI lockup. Payload-size + redaction policy (CD-10); NFR-PERF-01 (no UI lockup — render budget).
10 Server restart mid-session Session resumes / marks stale via watchdog; raw events intact. CD-1 replay-on-startup; NFR-OPS-01; P0 double-replay test (§4). Refinement: the project’s watchdog separates stale (idle session) from “unknown” (an agent whose SubagentStop never arrived, WP-IN12); the source folded both into “stale”.

Reconciliation with WP-X4’s 12 and §5’s illustrative anchors. WP-X4’s catalogue = these 10 + compaction-mid-session + PreCompact re-pricing (concept-analysis-v2 §6), so the recovered 10 are a strict subset — no conflict, no duplication. Where the recovered scenarios overlap the §5 illustrative anchor table, they are the same control seen from the negative-test angle: #5/#6 ↔ “Server fails startup when DASHBOARD_TOKEN is unset” + the timing-safe compare (one CD-7 control, three distinct assertions — keep all three); #7 ↔ “SSE rejects a cross-origin connection”; #8 ↔ “a fixture model with no price row FAILS CI” (WP-C6, the hardened build-time facet); #10 ↔ “kill+restart resumes JSONL tail-follow at the persisted offset” + P0 test #2.

Where the recovered scenarios fill a gap — EXPANDED entries with no row in §5’s illustrative anchor table, implicit in acceptance criteria but never enumerated as negative scenarios — the value of the recovery concentrates: #1 duplicate-hook dedup, #3 missing-parent-id orphaning, #4 malformed-JSON rejection (as distinct from the accept-and-store event_type case), and #9 huge-payload handling. These four are the scenarios WP-X4 most needs written as explicit test bodies.

6. The coverage gate: specified at >90%, shipped at 100

The coverage bar is a canonical decision, not a style preference: CD-7 states plainly that “the coverage gate [is a] boundary condition from commit one, CI-blocking… >90% coverage blocks merges” (concept-analysis-v2 §3, CD-7), and the build-sequencing principle underneath it is that “the security invariants and the >90% coverage gate are load-bearing from Phase 1 — not a hardening pass at the end” (implementation-plan.md §B.0). Three work packages implement it end to end:

WP Owner What it does
WP-F3 qa Vitest coverage harness + >90% gate config (scope defined). Produces lcov + json-summary output.
WP-F4 devops CI pipeline skeleton (GitHub Actions) with the coverage gate blocking merges.
WP-X5 devops CI coverage gate (>90%, blocking) live from Phase 1. A PR dropping below the threshold is blocked, demonstrated as such.

(development-plan §5, WP-F3/WP-F4/WP-X5.) All three land in wave 6–7, ahead of any ingest feature code (development-plan §4, waves 6–7; §7, “security + coverage go live at Phase 1, never deferred”), and CD-7’s coverage-gate obligation is itself implemented by WP-X5 in the CD-coverage matrix (development-plan §6, CD-7 row).

(As built: the “>90%” in CD-7’s clause and in those three Done-whens is 100 in every package that shipped (§6.1 below), and their “blocks merges” half was intent rather than mechanism until 2026-08-25, when main became branch-protected on the ci check — lowercase ci, the job id in .github/workflows/ci.yml, not its CI display name. A coverage regression now withholds the merge button from a contributor, but not from the repository owner, whose enforce_admins exemption is deliberate in a single-maintainer repository — see the note at the top of this page and the standing correction.)

Scope is resolved explicitly, not left ambiguous. The open scope question the gap analysis raised — whether the shipped UI would be quietly exempted from the bar — is resolved by not exempting it: the web package “counts toward the >90% gate” once Phase 4 ships it (development-plan §3, Phase 4 exit gate), and alerts modules are held to the same “>90% covered” bar once the alert track ships them (development-plan §3, Phase 6 exit gate — post-1.0 per best-path §6.1; WP-A10). The one deliberate carve-out — the labeled-experimental vector-DB stub — no longer exists: WP-X11 was deleted per best-path §6.3 (applied 2026-07-06), so no coverage exemption remains in scope.

A concrete illustration of “coverage from commit one” in practice: WP-F7’s security-invariant contract tests (loopback bind, token compare, SSE origin) are committed intentionally red at wave 8, before the server bootstrap (WP-U0) exists to make them pass at wave 9 — “do not merge WP-F7 as ‘passing’; its DoD is jointly owned with WP-U0” (development-plan §7). The same red-then-green discipline the coverage gate enforces for ordinary unit tests is applied deliberately to the security suite, not relaxed for it — see security model for the control catalogue those contract tests protect.

6.1 As built: why the number is 100, and what stops it being bought

The gate that shipped is stricter than the gate that was specified. All five packages — apps/server, apps/web, packages/core, packages/shared, packages/test-fixtures — run vitest run --coverage and pin lines, branches, functions and statements at 100, and as of 2026-08-15 all five sit at 100 on all four. The reasoning is written into the configs themselves, in a comment repeated in each: a 90% bar on a package sitting at 100% licenses a ten-point regression to pass in silence, which is the opposite of a gate. A threshold is only load-bearing when it is set at the level the code actually holds. Set lower, it does not measure the code; it measures how far the code is allowed to fall before anyone is told.

That raises the obvious objection: a 100% figure is exactly the kind of number that gets manufactured. There are three ways to buy one, and each has a test that reads the source as text and never imports it, so a mock or a stub cannot satisfy it.

The cheat What it actually does The guard
An ignore pragma — /* v8 ignore */, /* c8 ignore */, /* istanbul ignore */ Removes both arms of the operator from the denominator, so the uncovered arm stops existing rather than starts being tested Every file under src/** is swept; the offender list must be empty
Lowering the threshold The bar moves to wherever the code happens to be The config is read as text and each of the four numbers must literally be 100
Adding an exclude A file that is never measured cannot lower the average The config must contain include: ['src/**'] and must not contain exclude

The guards live in apps/server/test/coverage-honesty.test.ts and its counterparts in packages/core, packages/shared and packages/test-fixtures, plus the coverage honesty block in apps/web/test/honesty.test.tsx. Their premise is stated in the server file’s header and is the same premise as the ground-truth-tokens invariant this whole project is built on: a coverage figure inflated by hiding code is the same category of lie as an inferred token count.

The corollary the guards enforce is that the remedy for a genuinely unreachable branch is to delete it, not to hide it. There is a worked instance of this in the history of apps/server: the branch threshold sat at 99 while server.ts carried an unreachable ?? fallback — Fastify types request.url as string, so the alternative arm was dead code. The arm was deleted and the threshold raised to 100. Suppressing it with a pragma would have produced the same headline number by removing both arms from the denominator, which is the cosmetic version of the same move and the reason the pragma is banned outright rather than merely discouraged.

Three asymmetries in that story, stated rather than smoothed over.

First, apps/web does carry an exclude: src/main.tsx and src/vite-env.d.ts. main.tsx is the DOM entry point — a mount call exercised by the browser, not by jsdom — and is excluded on the same reasoning as a CLI entry point. The exclusion is narrow and named, but it means apps/web is the one package whose 100 is over a set of files chosen by hand rather than over everything under src/**. Worth recording alongside it: seven type-defensive ?? arms in src/views/layout/cost-flow.ts used to be excluded and are now genuinely reached, through the exported toFlowNode / toFlowLink converters and an injectable pathFor seam — the exclusion list shrank by being tested away rather than by being argued away.

Second, and directly downstream of the first, apps/web’s honesty test is weaker than the other four. It sweeps src/ for pragmas and asserts the offender list is empty, but it does not assert the four thresholds and it does not assert the absence of further exclude entries — it cannot, since the package legitimately has one. The practical consequence is that a change widening the web exclude list, or lowering the web thresholds, would not trip a guard. That is a real gap in the mechanism, not a technicality.

(Closed 2026-09-26. The web coverage honesty block in apps/web/test/honesty.test.tsx now also pins the evaluated shape of vitest.config.ts, read as text like the other four: the thresholds object must be exactly lines/branches/functions/statements at 100 with no extra key, include exactly 'src/**', and exclude exactly ['src/main.tsx', 'src/vite-env.d.ts'], each of which must exist on disk. It cannot assert the absence of an exclude, so it pins the list instead — the asymmetry in form stays, the gap in effect is gone. Four mutations were each killed: branches: 99, a third exclude entry, an extra perFile threshold key, and the exclude moved into a variable. Importing the config was tried first and rejected: vitest/config does not load under jsdom, and the file is outside the web tsconfig project.)*

Third, 100% means 100% of src/** — not of the repository. Every package’s coverage include is src/**, so two areas of live code never enter a denominator at all: hooks/install.mjs, which is exercised in earnest by apps/server/test/hooks-installer.test.ts against a throwaway temp directory but is measured by nothing; and the two CI gate scripts scripts/check-no-spawner.mjs and scripts/check-licenses.mjs, which have no unit tests whatsoever and are exercised only by being executed in CI. (Superseded in part, noted 2026-09-26: since 4b3cd2d each gate — and the Node-major guard scripts/check-node-version.mjs — exposes a pure core (scanTree / evaluateLicenses / evaluateNodeVersion plus a formatter) that apps/server/test/scripts-gates.test.ts drives against throwaway fixture trees, so “catches what it claims to catch” is now tested, not only “exits 0 on this repo”. What still holds: the scripts sit outside every coverage denominator, and their CLI wrappers are deliberately left unexecuted by tests, because running them would need another subprocess exemption.) Both gates do run on every CI invocation and both currently pass — the spawner gate reports OK (235 files scanned across 4 roots + repo-root config; 1 allowlisted) and the license gate OK (412 installed packages, all licenses allowlisted), measured 2026-08-15; re-measured 2026-09-19 as OK (272 files scanned across 4 roots + repo-root config; 1 allowlisted; 6 package.json manifests checked for forbidden direct dependencies) and OK (412 packages / 429 installed versions; 411 allowlisted, 1 under a documented exception) — but “the gate script works” is established by its output, not by a test of the script. (Re-measured again 2026-09-23: the spawner gate now scans 282 files across the same 4 roots + repo-root config, still 1 allowlisted and 6 manifests checked, and the license gate is unchanged at 412 packages / 429 installed versions, 411 allowlisted, 1 under the documented exception. The file count tracks the sources; it is a dated measurement, not a constant, and neither gate’s verdict has moved.)

(Amended 2026-09-23 (J-12): the OK lines quoted above also predate the spawner gate’s current output shape. Its real final line now reads check-no-spawner: OK (282 files scanned across 4 roots + repo-root config; 1 allowlisted; 5 line(s) inline-exempt; 6 package.json manifests checked for forbidden direct dependencies), and the run prints one inline opt-out in force line per exempt line plus two inline opt-out suppressed nothing (dead marker, not fatal) lines above it. The five exempt lines carry the three sanctioned exceptions - two exceptions need the marker on the import as well as on the call. Both extra clauses are reporting, not failure: exit code 0.)

(Amended 2026-09-24: the gate now also checks every package.json scripts string for a bind wider than loopback, flags imports of more subprocess-wrapper packages, and flags the vm module. Its final line now ends 6 package.json manifests checked for forbidden direct dependencies and wide-bind scripts); measured at 308 files scanned, still exit 0.)

(Amended again 2026-09-24 (II4): the license gate counted a package once per license it appeared under, so a name installed at two versions under two licenses counted twice. It now counts unique names, and the line reads OK (407 packages / 429 installed versions; 406 allowlisted, 1 under a documented exception). The dependency tree did not change; the earlier 412 / 411 figures were the double count. All three gate scripts also now recognise themselves when invoked through a symlinked path, where they previously exited 0 without checking.)

(Re-measured 2026-09-26: check-no-spawner: OK (313 files scanned across 4 roots + repo-root config; 1 allowlisted; 5 line(s) inline-exempt; 6 package.json manifests checked for forbidden direct dependencies and wide-bind scripts) and check-licenses: OK (406 packages / 428 installed versions; 405 allowlisted, 1 under a documented exception). Both exit 0; the package count moved with the dependency tree, not with the gate. Later the same day the spawner line gained the no-SSRF clause ; 93 server-process files checked for outbound network calls).)

And the standing caveat that no coverage number escapes: 100% line and branch coverage records that every line and branch executed, not that every behaviour was asserted. It is a floor under the test suite, not a statement about its depth. What gives this suite its actual weight is the material in §3.1, §4 and §5 — an independent reader in the token proof, byte-identical replay snapshots, and twelve enumerated ways the system is expected to fail well.

7. Where this lands on the roadmap

Phase Wave(s) Testing-relevant exit gate
0 — Feasibility spike 1–4 WP-S1’s paired corpus + Ivan-labeled trees captured; WP-S7 reads GO/CONDITIONAL-GO on the evidence they produced.
1 — Foundation, security spine, storage, ports 6–8 Coverage harness + CI gate green & blocking (WP-F3, WP-F4, WP-X5); golden fixture corpus promoted and labeled (WP-X1, WP-X2); WP-F7’s security contract tests exist, deliberately red.
3 — Projection, the DAG moat, reconciliation, cost 14, 16 Three P0 tests green & merge-blocking; hierarchy ≥95% vs. the labeled corpus; 12-scenario negative catalogue green (WP-X3, WP-X4, WP-IN13).
4 — Read API + SPA — Web package counts toward the >90% gate.
6 — Alert-track release hardening (post-1.0 per best-path §6.1; the operator alerts UI WP-A8/WP-A9 was cut per §6.2) 17 Alerts modules >90% covered; RELEASE.md enumerates every CD-7 gate with a verification step (WP-A10, WP-X9).

(development-plan §3, §4.) The hard structural point: none of this is a procedural checklist an agent could skip under time pressure — WP-F1 (the monorepo scaffold itself) has a real dependency edge on WP-S7, and WP-IN13/WP-X3 are wired as blocking CI checks, not advisory ones (development-plan §1, §5). Two corrections to that table from the as-built note: the coverage figure it calls “>90%” is 100 in every package that shipped, and “blocking” described the intent rather than the mechanism until 2026-08-25 — the workflow ran on every push and pull request, but main was unprotected, so nothing physically withheld a merge. The mechanism now exists: main requires the ci check, so a red run withholds the merge button from a contributor — but not from the repository owner, because enforce_admins is deliberately off in a repository with one maintainer whose normal working mode is a direct push to main, and enabling it would lock that sole maintainer out (see the standing correction).

What’s undecided

See also