This page is the QA reference for agenthropic: how the golden real-session fixture
corpus is captured, promoted, and labeled with ground truth; the three P0
release-blocker tests that must be green before Phase 3 is considered done; the
12-scenario negative-test catalogue; and the coverage gate — specified at >90%,
shipped at 100 — that is live from Phase 1, not a hardening pass bolted on at the
end. The key takeaway up front: this project treats test infrastructure as a first-class,
phase-spanning engineering problem, not overhead — “you cannot unit-test ‘the
dashboard’; the four real units are ingest correctness, tree correctness, cost
correctness, live-flow correctness, each with a distinct harness. The golden
real-session fixture corpus is the #1 QA investment — without it, >90% coverage is
high-coverage tests of synthetic happy-paths (false safety)” (concept-analysis-v2
§4.3). Everything below traces to
development-plan.md Track X (WP-X1…WP-X5)
and the quantified acceptance criteria in
concept-analysis-v2.md §6.
Update — 2026-08 (as built). This page was written before a single test existed, so the sections below still speak in the future tense about gates that have since been built, superseded by something stricter, or blocked on a human act that has not happened. This note is the verified state of the suite as of 2026-08-15, measured by running
pnpm -r --workspace-concurrency=1 run teston a clean tree and reading the per-packagecoverage/coverage-summary.jsonit writes — with the merge-gating bullet re-verified on 2026-08-25, the daymainbecame branch-protected. Where a section below disagrees with this note, this note is the current truth and the section is the historical intent.
- The three P0 release-blocker tests are green. They live in
apps/server/test/p0/—p0-token-reconciliation.test.ts,p0-double-replay.test.ts,p0-dag-rebuild.test.ts, sharing aharness.ts(a fourth file beside them,p0-five-daily-questions.test.ts, is the CD-10 five-daily-questions proof over real HTTP, not a reconciliation test). One detail matters more than the pass/fail: the token-reconciliation proof does not compare the parser against itself. It reads the JSONL with an independent minimal reader written inside the test against the normative rules ofparser-spec.md§5.1–5.3, so a parser bug cannot make its own proof pass. The double-replay proof compares twoVACUUM INTOsnapshots withBuffer.equalsunder a fixed clock — byte-identical, not “equivalent” — and then re-asserts the same claim a second, independent way against an ordered logical dump of every table. The DAG-rebuild proof additionally demonstrates the hooks-are-liveness-only rule: appending hook events leaves the DAG dump unchanged.- All twelve negative-catalogue scenarios now have executable bodies, and they are green. They are split by nature rather than kept in one file: the pure parser and cost facets (#3, #4, #8, #11, #12) live in
packages/core/test/negative/negative-catalogue.core.test.ts, and the HTTP, persistence and security facets (#1, #2, #3-db, #4-http, #5, #6, #7, #8-halt, #9, #10) live inapps/server/test/negative/. Three scenarios appear on both sides on purpose — a malformed JSONL line and a malformed HTTP body are different failures of the same catalogue entry. The security entries are asserted at the level where an oracle would hide: byte-identical401bodies across four different wrong-token shapes (no token echo, no length oracle) and a403on a foreignOriginboth with and without a valid token, so a probe cannot learn whether its token was good either. One half of scenario #2, “normalized anomaly flagged”, remains untestable as the catalogue words it — there is no hook normalizer, because hooks are a secondary signal — and the test file says so in place of quietly dropping the clause; the anomaly surfaces instead through theWP-IN12watchdog as a visibleunknownstatus.- Totals as of 2026-09-18: 131 test files / 2428 tests, green, across five packages —
apps/server83/1255,apps/web22/697,packages/core14/253,packages/test-fixtures5/139,packages/shared7/84. These counts move with every commit; treat them as a dated measurement, not a constant. (Re-measured 2026-09-23: 140 test files / 2621 tests, still green —apps/server90/1358,apps/web22/726,packages/core16/314,packages/test-fixtures5/139,packages/shared7/84. Coverage still 100% on all four axes in every package. Re-measured 2026-09-26: 167 test files / 3060 tests, still green —apps/server115/1629,apps/web23/860,packages/core17/338,packages/test-fixtures5/139,packages/shared7/94; 100% on all four axes in every package.)- The coverage gate is no longer “>90%” anywhere in the repo. It is 100. All five packages run
vitest run --coverageand all five pinlines/branches/functions/statementsat100, and all five are currently at 100 on every one of those four metrics.packages/test-fixturesis no longer excluded from the gate — that carve-out was reversed on the reasoning that a defect in a fixture builder does not fail loudly, it silently weakens every downstream test that consumes the fixture. §6.1 below is the as-built account of why the number is 100 rather than 90 and what stops it from being bought cheaply. Honest note, kept from the previous revision: until 2026-07-30apps/webranvitest runwithout--coverage, so its configured thresholds silently never executed. That was found and fixed.- §2’s golden real-session corpus is not what shipped. The three-tier raw/redacted/manifested promotion of ≥3 captured real sessions was not built. What exists is
packages/test-fixtureswith eight typed fixtures —flat-tool-use,nested-workflow,queue-operation,task-notification-recovery,depth-2-sync,usage-dedup,legacy-bare-explore,agent-outcome-errors— plus per-suite corpora written into temp directories. Tests never touch the real~/.claude/projects: every corpus is built undermkdtempSyncwith explicitly injected env. The package location that §”What’s undecided” called “a named leaning” is settled — it ispackages/test-fixtures.- §3’s labeled ground truth is half-built, and the missing half cannot be built by an agent. The format, loader, scorer, report and gate runner all exist (§3.1). The labels do not.
packages/test-fixtures/annotations/human/is empty, and the fivespike/corpus/sessions/<short>/LABEL-ME.mdtrees are still unfilled — labeling is Ivan’s act, not an agent’s. Consequence, stated plainly: the ≥95% hierarchy correctness gate in §2/§4 has never been scored. The gate run reportsSUBSTRATE UNAVAILABLE,Phase-3 exit clause: NOT MEASURED - no hand-labeled sessionsand a NOT CERTIFIED verdict, and it passes as a test in that state, because an unlabeled corpus is an honest state rather than a broken build. Every spike-derived accuracy number in this corpus stays PROVISIONAL until that labeling happens. No test result on this page should be read as satisfying that bar.- Merge-blocking for a contributor, not for the repository owner — since 2026-08-25.
.github/workflows/ci.ymlruns the spawner gate, typecheck, lint, format check, the web production build, the full suite with its coverage thresholds, and the license gate — in that order, security first so a broken invariant fails in seconds. Until 2026-08-25 this bullet read “nothing here is physically merge-blocking yet”, and it was true:mainwas unprotected, so a red run withheld nothing.mainis now branch-protected — thecicheck is required (lowercaseci, the job id in that workflow, not itsCIdisplay name), and force-pushes tomainand deletion ofmainare refused for everyone. One exemption is deliberate and is stated wherever it applies:enforce_adminsis off, because this is a single-maintainer repository whose normal working mode is a direct push tomain, and turning it on would lock the sole maintainer out of their own repository. How to read the sections below: where they say “merge-blocking”, read “merge-blocking for anyone who is not the repository owner; CI-failing and loud, but bypassable, for the owner.” Verify withgh api repos/IvanBBaev/agenthropic/branches/main/protection --jq '{contexts: .required_status_checks.contexts, enforce_admins: .enforce_admins.enabled}'→{"contexts":["ci"],"enforce_admins":false}; the full write-up is the standing correction.
The QA lens’s starting move is refusing to treat “test the dashboard” as one problem. It decomposes into four real units, each needing its own harness because each fails in a structurally different way (concept-analysis-v2 §4.3):
| Unit | What it proves | Fails when |
|---|---|---|
| Ingest correctness | A hook event and a JSONL line for the same fact collapse into one events_raw row; nothing is lost or double-counted. |
Dual-write dedup is wrong, or a crash mid-ingest drops data. |
| Tree correctness | The reconstructed subagent hierarchy in orchestration_edges matches reality. |
A parent→child edge is missing, mis-attributed, or the tree is only “almost correct” — which the QA lens judges worse than no graph because it manufactures false trust (concept-analysis-v2 §4.3). |
| Cost correctness | Every displayed dollar traces to ground-truth tokens × a dated, versioned price, including across a compaction. | A stale price, a silent zero-cost default, or a PreCompact repricing bug. |
| Live-flow correctness | The realtime status board reflects reality without a permanent false “working” state. | A missing SubagentStop never resolves to a watchdog “unknown”. |
On top of these four units, the project’s consolidated test model layers a shape and a priority scheme (concept-analysis-v2 §4.3):
PreCompact re-pricing) — see §5.This is scoped as a project-wide, always-on workstream, not a phase: the implementation
plan names it WS-Test — “golden real-session fixtures corpus (happy + pathological:
deep nesting, missing Stop, mid-session compaction, two concurrent instances);
reconciliation/idempotency/compaction/security test suites; CI coverage gate at >90%,
blocking merges” — one of four cross-cutting workstreams that “run through every phase”
(implementation-plan.md §B.1), alongside security, docs, and ops.
Synthetic fixtures cannot stand in for this corpus: a hand-written happy-path JSONL transcript cannot prove the tree survives a crash it was never built to have, so the corpus has to be captured from real Claude Code sessions, not authored. Capture and permanent promotion are two different work packages in two different phases:
Phase 0 (throwaway spike) Phase 1 (permanent corpus)
────────────────────────── ──────────────────────────
WP-S1 paired-capture harness ──────► WP-X1 golden fixture corpus
· ≥3 real sessions · promotes the Phase-0 capture
· paired JSONL + hook log per · three artifacts: raw, redacted,
session manifested
· Ivan-labeled expected tree · ≥3 real sessions; all four
per session pathologies each represented
· throwaway hook block · a "manifest self-test" — CI
reverted after capture asserts pathology coverage from
the manifest, not a comment
(development-plan §5, WP-S1, WP-X1.) WP-X1 depends on WP-S2, WP-S4, and WP-F1
— the two Phase-0 probes that already exercise the corpus, plus the monorepo scaffold —
and is scheduled at wave 6, labeled there as “corpus promotion” (development-plan
§4, wave 6). It is not stood up in isolation: WP-X1 runs alongside WP-X5 (the
coverage gate config) and WP-X7 (the docs-site build) in the same wave, all gated on
the Phase-0 GO/CONDITIONAL-GO verdict (WP-S7) that unblocks WP-F1 in the first
place (development-plan §1, §4).
WP-X1’s own Done-when names three artifacts, not one (development-plan §5, WP-X1):
WP-S1’s paired harness: the JSONL
transcript and the hook log side by side, unmodified.events 90 days, backup files 30 days behind a floor of 7,
token_usage never), closing concept-analysis-v2 §7, open question 6 — see backup & restore.WP-X1).Both WP-S1’s capture harness and WP-X1’s promoted corpus are required to represent
the same four pathological session shapes (development-plan §5, WP-S1;
concept-analysis-v2 §6):
| # | Pathology | Why it’s in the corpus |
|---|---|---|
| 1 | Crashed, no Stop |
Proves the missing-SubagentStop → watchdog “unknown” rule, and that a crash doesn’t leave a permanent false “working” state. |
| 2 | Deep nesting | Stresses the DAG-rebuild-from-JSONL path against multi-level subagent-of-subagent chains, not just a flat parent→child pair. |
| 3 | Mid-session PreCompact |
The one EXPANDED’s own negative-test catalogue omitted — proves the tree and the cost baseline both survive a context compaction (concept-analysis-v2 §4.3, “its catalogue dropping the compaction case”). |
| 4 | Two concurrent instances | Proves instance/host_id correctly partitions two simultaneously running Claude Code sessions instead of merging their trees. |
The corpus size floor is ≥3 real sessions (development-plan §5, WP-S1/WP-X1;
concept-analysis-v2 §6), and it doubles as the population the ≥95% hierarchy
correctness gate is measured against in Phase 3 (concept-analysis-v2 §6) — see
ingest & reconciliation §10 for that gate in
full.
expected/*.json and the typed loaderA corpus of real sessions is only useful to CI if there is a machine-checkable answer
key. WP-X2 — labeled ground-truth annotations + typed fixture loader — is that
answer key: “every session has an expected/*.json; a test fails if any lacks one”
(development-plan §5, WP-X2). Two things make this more than a convention:
expected/*.json, so a new corpus session cannot silently ship unlabeled and
invisible to the P0 and negative-catalogue suites that consume it.WP-X2 is scheduled directly after WP-X1 and depends on
WP-D1 (the shared storage-port row types), at wave 7, where development-plan
labels it “labeled fixtures” (development-plan §4, wave 7; §5, WP-X2) — the expected
trees are consumed as typed fixtures against the same row-shape contracts the
production projection code uses, not as untyped JSON blobs a test has to hand-parse.Everything downstream — the three P0 tests (§4), the negative catalogue (§5), and the
≥95% hierarchy gate (§2) — is scored against these expected/*.json files, which is why
WP-X2 sits on the critical dependency edge into both WP-X3 and WP-X4
(development-plan §5).
What shipped is not expected/*.json but something with the same job and a stricter
posture about its own authority. The ground truth lives in
packages/test-fixtures/annotations/, in a hand-writable markdown format that is
diffable and needs no tooling to author: a ## meta block declaring the session,
provenance, substrate, labeled-by and labeled-on, then a ## edges block of one
line per subagent — <child hex> <- ROOT | ORPHAN | UNKNOWN | <parent hex>, with an
optional trailing comment. Everything outside those two blocks is prose the loader
ignores, so a labeler can leave notes to themselves in the file. The loader, validator,
scorer, report renderer and read-only filesystem adapters are in
packages/test-fixtures/src/annotations/; the runner that wires the real parser to them
is packages/core/test/hierarchy-gate.test.ts.
Three design choices in that tooling are the point of it, and each exists to stop a number from meaning less than it appears to.
The score is a Wilson lower confidence bound, not a ratio. A naive percentage is
silent about sample size — 3/3 and 300/300 both read “100%”, and only one of them is
evidence. wilsonLowerBound() computes the one-sided 95% lower bound instead, which
returns 0 when there are no observations at all: no data, no confidence. The direct
consequence is a hard floor on the sample. Solving n / (n + z²) ≥ 0.95 gives
n ≥ 0.95 · 1.6449² / 0.05 = 51.4, so 52 labeled agents is the minimum at which even a
flawless run can clear the bar, and roughly 90 are needed to survive a single error.
minimumClaimsForThreshold() computes that floor and certifyExitGate() refuses to
certify below it no matter how good the raw percentage looks. The two prepared templates
— b24be30c (42 agents, dual on-disk layout, deepest observed nesting) and f28af3fd
(18 agents, an independent depth-2 population, 5 compactions) — total 60, chosen to clear
52 with a little headroom while covering structurally distinct ground. Labeling only one
of them leaves the sample below the floor, and the gate says so.
Provenance is enforced structurally, not by convention. annotations/synthetic/
holds eight annotations, one per fixture, that state the hierarchy each fixture was
built to have. They are genuinely useful — they prove the loader, the scorer, the
report and every join path work end to end, and they regression-guard the depth-2 case —
but they were written by the same side as the parser, so agreement with them proves
internal consistency and nothing else. Every annotation must therefore declare
provenance: human or provenance: synthetic-by-construction; scoreCorpus() throws if
a corpus mixes the two, so a blended figure cannot be produced by accident;
certifyExitGate() hard-refuses any non-human corpus and prints the reason; and the
report prints an ADMISSIBILITY banner above the numbers so no reader can quote the
figure without also reading what it is made of.
Abstention is a first-class answer. UNKNOWN is not scored as a miss and not scored
as agreement — it is excluded from the accuracy fraction and reported in its own bucket,
because a guess that turns out wrong is strictly worse than an abstention when the exit
gate would be signed against it. ORPHAN, by contrast, is a positive claim (“there is no
parent to find here, and a parser that invents one is wrong”) and is scored. To stop
abstention from becoming a way to launder a number, the report prints label coverage and
a worst-case figure — every abstention assumed wrong — next to the headline, so an
under-labeled corpus cannot masquerade as a passing one.
What the gate reports today. annotations/human/ is empty. Running
pnpm --filter @agenthropic/core exec vitest run test/hierarchy-gate.test.ts therefore
prints SUBSTRATE UNAVAILABLE - not measured for each missing substrate and
Phase-3 exit clause: NOT MEASURED - no hand-labeled sessions, points at
packages/test-fixtures/annotations/README.md, and returns NOT CERTIFIED. The test
itself passes in that state, and the code says why in a comment at the branch: nothing
to certify on this machine — do not manufacture a verdict. This is the distinction the
whole subsystem is built around. A build that fails because a human has not done a manual
task teaches a team to route around the check; a build that green-washes an unmeasured
gate is worse. Passing while loudly reporting n = 0 and refusing to certify is the only
option that is both honest about the state and honest about the number.
Until those templates come back filled in, every hierarchy-accuracy figure anywhere in this corpus is PROVISIONAL and the Phase-3 exit clause is unmet — not failed, unmeasured.
Three tests are named release-blockers: they must be green and merge-blocking in CI before Phase 3 — projection, the DAG moat, reconciliation, cost — can be considered done, and no other feature work substitutes for them (development-plan §3, Phase 3 exit gate; concept-analysis-v2 §4.3). (As built: green, and merge-blocking since 2026-08-25 for anyone who is not the repository owner — read “merge-blocking” here per the note at the top of this page.) The QA lens calls the third one “the make-or-break test both externals omit” (concept-analysis-v2 §4.3):
| # | Test | Proves |
|---|---|---|
| 1 | Σ token_usage == JSONL exact, per session |
The ground-truth-tokens invariant holds in the projected data, not just the raw log — zero drift, no silent rounding, no double count. |
| 2 | Double-replay → byte-identical DB state | Replay-on-startup is deterministic and safe to run on every process start, not just theoretically pure. |
| 3 | DAG-rebuild from JSONL alone, after a simulated outage | The persisted orchestration_edges tree survives the exact failure mode it exists to survive. |
(concept-analysis-v2 §4.3, §6; development-plan §5, WP-X3: “Three P0 reconciliation
release-blocker tests. Σtoken_usage==JSONL exact; double-replay byte-identical;
DAG-rebuild-from-JSONL-alone.”) Ownership is split across two work packages by design:
WP-X3 (QA/fixture track, deps X2, IN10, IN7, D1) owns the test bodies run against the
golden corpus from §2–3; WP-IN13 (deps IN10, IN9, X1) wires them into CI as
“Reconciliation / idempotency / DAG-rebuild suite (P0 blockers). All three P0 tests green
in CI and blocking” (development-plan §5, WP-IN13). Both land at wave 16, alongside
the operator alerts endpoints — the last thing on the moat’s own critical sub-chain
before release (development-plan §4, wave 16 and the “moat spine” note).
Two adjacent, non-P0 bars sharpen what “passing” is allowed to mean, and both are scored against the same corpus:
SubagentStop must resolve to an explicit “unknown” state within the
watchdog window, never a permanent “working” (concept-analysis-v2 §6; WP-IN12).The full mechanics of why these three tests are structured this way — the
events_raw substrate, idempotency keys, replay-on-startup, and the contingent outbox —
are the architecture-level deep dive on
ingest & reconciliation §10; this page is
the QA-process view: who owns the test body, what corpus it runs against, and where it
sits in the release gate.
WP-X4 — expanded negative-test catalogue (12 scenarios) — requires that “each maps
to a CD/acceptance criterion” (development-plan §5, WP-X4). Its composition is
explicit in the consolidated test model (concept-analysis-v2 §4.3):
PreCompact re-pricing — “all ten EXPANDED §7.1 negative scenarios pass, plus
compaction-mid-session and PreCompact re-pricing” (concept-analysis-v2 §6), correcting
what the QA lens itself flags as the external catalogue “dropping the
compaction case” outright (concept-analysis-v2 §4.3).WP-X4 depends on WP-X2 (the labeled corpus + typed loader, §3), WP-IN6 (the pure
Normalizer), and WP-C4 (compaction-baseline repricing), and is scheduled at
wave 14, alongside the reconciliation/backfill and session-tree endpoints
(development-plan §4, wave 14; §5, WP-X4). That dependency set is itself informative
about the catalogue’s shape: it has to exercise the Normalizer’s handling of malformed
or unrecognized input (via IN6) and the cost engine’s repricing path (via C4), not
only the reconciliation logic the three P0 tests already cover.
The base 10 of the catalogue are now recovered from the source in §5.1 (LOST-7); the
table below is the complementary CD-anchored view — the acceptance criteria elsewhere
in concept-analysis-v2 §6 and the CD register (§3) name the concrete, individually
testable conditions the catalogue has to trace to, each the kind of scenario “mapped to a
CD/acceptance criterion” that WP-X4’s Done-when requires:
| Traceable condition | CD / WP | Category |
|---|---|---|
An unknown event_type is stored, not crashed |
CD-2; development-plan §3, Phase 2 exit gate | Ingest robustness |
| Kill+restart resumes JSONL tail-follow at the persisted offset, zero loss/dup | development-plan §3, Phase 2 exit gate | Ingest robustness |
Missing SubagentStop → explicit “unknown” within the watchdog window |
concept-analysis-v2 §6; WP-IN12 |
Live-flow honesty |
A PreCompact session reprices correctly against its preserved baseline |
concept-analysis-v2 §6; WP-C4 |
Cost correctness |
| A fixture model with no price row FAILS CI | concept-analysis-v2 §6; WP-C6 |
Cost correctness |
events_raw exposes no UPDATE/DELETE path |
CD-4, CD-7; concept-analysis-v2 §6 | Data integrity |
Server fails startup when DASHBOARD_TOKEN is unset |
CD-7 | Security |
No-spawner grep/static gate fails the build on a child_process import |
CD-7; WP-F5 |
Security |
| An SSRF test proves no outbound dial to a payload-supplied URL | CD-7; WP-F5 |
Security — static half live since 2026-09-26 (gate:spawner refuses outbound primitives in server-process code, tested in apps/server/test/scripts-gates.test.ts); the dynamic test waits for a dispatcher to exist |
| SSE rejects a cross-origin connection | CD-5, CD-7 | Security |
This table is illustrative of the categories and traceable anchors the 12-scenario catalogue draws from in this project’s own decision set — it is not a claim that these are the literal, final 12 entries. The literal base 10 originate in the external report’s §7.1; they are now recovered verbatim in §5.1 below (finding LOST-7).
The ten literal scenarios below are the external EXPANDED report’s own §7.1 negative-test
catalogue, recovered from the EXPANDED external report (internal source material, not
published in this repo) per
corpus-audit finding LOST-7 (corpus-audit-2026-07-06.md
§4.3). They are the base 10 of WP-X4’s 12-scenario catalogue; the remaining two
(compaction-mid-session and PreCompact re-pricing) are the project-specific
additions described in §5 above (concept-analysis-v2 §6). The expected-behavior column
is quoted from the source; the maps to column is this project’s decision anchor. Two
scenarios were hardened as they crossed into the plan of record — flagged inline.
| # | Scenario (EXPANDED §7.1) | Expected behavior (source) | Maps to (CD / NFR / WP) |
|---|---|---|---|
| 1 | Duplicate hook event | No duplicate normalized event or double token total. | CD-2 idempotency-keyed events_raw; NFR-DATA-02; P0 double-replay test (§4). |
| 2 | SubagentStop arrives before SubagentStart |
Raw event stored; normalized anomaly flagged; UI shows uncertain edge. | CD-2/CD-3; anomaly + watchdog state (WP-IN12); relates to OPEN-2 ('unknown' is in the agents.status CHECK since migration 4, which closed it). |
| 3 | Missing parent id | Agent marked orphan / pending reparent; no fake root unless explicitly synthesized. | CD-4 self-ref parent_agent_id, orphan-safe self-FK (WP-D6); NFR-DATA-01. |
| 4 | Malformed JSON | 400 error; no DB mutation except optional audit log. | CD-2 ingest schema validation (FR-01). Distinct from the §5 anchor “an unknown event_type is stored, not crashed” — this one rejects, that one accepts-and-stores. |
| 5 | Unauthenticated POST | 401/403; no raw event stored. | CD-7 mandatory DASHBOARD_TOKEN; NFR-SEC-02. |
| 6 | Invalid-token timing-attack attempt | Timing-safe compare; no differentiated error leak. | CD-7 timing-safe DASHBOARD_TOKEN compare; NFR-SEC-02. |
| 7 | Connection from a foreign origin (hardened: WS → SSE) | Rejected. | CD-5 same-origin SSE (transport corrected from the source’s “WebSocket” per CD-5/ADR-0007), CD-7. Overlaps §5’s “SSE rejects a cross-origin connection”. |
| 8 | Unknown model pricing (hardened) | Token counts shown; cost marked unknown/estimated, not silently zero. | CD-3/CD-4; the “estimated, never silent zero” rule governs runtime display. Project hardening: at fixture/build time a model with no price row FAILS CI (WP-C6) — a stricter gate the source did not have; both hold, at different times. |
| 9 | Huge payload | Rejected or truncated per policy; no UI lockup. | Payload-size + redaction policy (CD-10); NFR-PERF-01 (no UI lockup — render budget). |
| 10 | Server restart mid-session | Session resumes / marks stale via watchdog; raw events intact. | CD-1 replay-on-startup; NFR-OPS-01; P0 double-replay test (§4). Refinement: the project’s watchdog separates stale (idle session) from “unknown” (an agent whose SubagentStop never arrived, WP-IN12); the source folded both into “stale”. |
Reconciliation with WP-X4’s 12 and §5’s illustrative anchors. WP-X4’s catalogue =
these 10 + compaction-mid-session + PreCompact re-pricing (concept-analysis-v2 §6), so
the recovered 10 are a strict subset — no conflict, no duplication. Where the recovered
scenarios overlap the §5 illustrative anchor table, they are the same control seen from
the negative-test angle: #5/#6 ↔ “Server fails startup when DASHBOARD_TOKEN is unset” +
the timing-safe compare (one CD-7 control, three distinct assertions — keep all three);
#7 ↔ “SSE rejects a cross-origin connection”; #8 ↔ “a fixture model with no price row
FAILS CI” (WP-C6, the hardened build-time facet); #10 ↔ “kill+restart resumes JSONL
tail-follow at the persisted offset” + P0 test #2.
Where the recovered scenarios fill a gap — EXPANDED entries with no row in §5’s
illustrative anchor table, implicit in acceptance criteria but never enumerated as negative
scenarios — the value of the recovery concentrates: #1 duplicate-hook dedup, #3
missing-parent-id orphaning, #4 malformed-JSON rejection (as distinct from the
accept-and-store event_type case), and #9 huge-payload handling. These four are the
scenarios WP-X4 most needs written as explicit test bodies.
The coverage bar is a canonical decision, not a style preference: CD-7 states plainly that “the coverage gate [is a] boundary condition from commit one, CI-blocking… >90% coverage blocks merges” (concept-analysis-v2 §3, CD-7), and the build-sequencing principle underneath it is that “the security invariants and the >90% coverage gate are load-bearing from Phase 1 — not a hardening pass at the end” (implementation-plan.md §B.0). Three work packages implement it end to end:
| WP | Owner | What it does |
|---|---|---|
WP-F3 |
qa | Vitest coverage harness + >90% gate config (scope defined). Produces lcov + json-summary output. |
WP-F4 |
devops | CI pipeline skeleton (GitHub Actions) with the coverage gate blocking merges. |
WP-X5 |
devops | CI coverage gate (>90%, blocking) live from Phase 1. A PR dropping below the threshold is blocked, demonstrated as such. |
(development-plan §5, WP-F3/WP-F4/WP-X5.) All three land in wave 6–7, ahead of
any ingest feature code (development-plan §4, waves 6–7; §7, “security + coverage go
live at Phase 1, never deferred”), and CD-7’s coverage-gate obligation is itself
implemented by WP-X5 in the CD-coverage matrix (development-plan §6, CD-7 row).
(As built: the “>90%” in CD-7’s clause and in those three Done-whens is 100 in every package
that shipped (§6.1 below), and their “blocks merges” half was intent rather than mechanism until
2026-08-25, when main became branch-protected on the ci check — lowercase ci, the job id in
.github/workflows/ci.yml, not its CI display name. A coverage regression now withholds the
merge button from a contributor, but not from the repository owner, whose enforce_admins
exemption is deliberate in a single-maintainer repository — see the note at the top of this page
and the standing correction.)
Scope is resolved explicitly, not left ambiguous. The open scope question the gap
analysis raised — whether the shipped UI would be quietly exempted from the bar — is
resolved by not exempting it: the web package “counts toward the >90% gate” once
Phase 4 ships it (development-plan §3, Phase 4 exit gate), and alerts modules are held
to the same “>90% covered” bar once the alert track ships them (development-plan §3,
Phase 6 exit gate — post-1.0 per best-path §6.1; WP-A10). The one deliberate
carve-out — the labeled-experimental vector-DB stub — no longer exists: WP-X11 was
deleted per best-path §6.3 (applied 2026-07-06), so no coverage exemption remains
in scope.
A concrete illustration of “coverage from commit one” in practice: WP-F7’s
security-invariant contract tests (loopback bind, token compare, SSE origin) are
committed intentionally red at wave 8, before the server bootstrap (WP-U0) exists
to make them pass at wave 9 — “do not merge WP-F7 as ‘passing’; its DoD is jointly
owned with WP-U0” (development-plan §7). The same red-then-green discipline the
coverage gate enforces for ordinary unit tests is applied deliberately to the security
suite, not relaxed for it — see security model for the control
catalogue those contract tests protect.
The gate that shipped is stricter than the gate that was specified. All five packages —
apps/server, apps/web, packages/core, packages/shared, packages/test-fixtures —
run vitest run --coverage and pin lines, branches, functions and statements at
100, and as of 2026-08-15 all five sit at 100 on all four. The reasoning is written into
the configs themselves, in a comment repeated in each: a 90% bar on a package sitting at
100% licenses a ten-point regression to pass in silence, which is the opposite of a gate.
A threshold is only load-bearing when it is set at the level the code actually holds. Set
lower, it does not measure the code; it measures how far the code is allowed to fall
before anyone is told.
That raises the obvious objection: a 100% figure is exactly the kind of number that gets manufactured. There are three ways to buy one, and each has a test that reads the source as text and never imports it, so a mock or a stub cannot satisfy it.
| The cheat | What it actually does | The guard |
|---|---|---|
An ignore pragma — /* v8 ignore */, /* c8 ignore */, /* istanbul ignore */ |
Removes both arms of the operator from the denominator, so the uncovered arm stops existing rather than starts being tested | Every file under src/** is swept; the offender list must be empty |
| Lowering the threshold | The bar moves to wherever the code happens to be | The config is read as text and each of the four numbers must literally be 100 |
Adding an exclude |
A file that is never measured cannot lower the average | The config must contain include: ['src/**'] and must not contain exclude |
The guards live in apps/server/test/coverage-honesty.test.ts and its counterparts in
packages/core, packages/shared and packages/test-fixtures, plus the
coverage honesty block in apps/web/test/honesty.test.tsx. Their premise is stated in
the server file’s header and is the same premise as the ground-truth-tokens invariant
this whole project is built on: a coverage figure inflated by hiding code is the same
category of lie as an inferred token count.
The corollary the guards enforce is that the remedy for a genuinely unreachable branch
is to delete it, not to hide it. There is a worked instance of this in the history of
apps/server: the branch threshold sat at 99 while server.ts carried an unreachable
?? fallback — Fastify types request.url as string, so the alternative arm was dead
code. The arm was deleted and the threshold raised to 100. Suppressing it with a pragma
would have produced the same headline number by removing both arms from the
denominator, which is the cosmetic version of the same move and the reason the pragma is
banned outright rather than merely discouraged.
Three asymmetries in that story, stated rather than smoothed over.
First, apps/web does carry an exclude: src/main.tsx and src/vite-env.d.ts.
main.tsx is the DOM entry point — a mount call exercised by the browser, not by jsdom —
and is excluded on the same reasoning as a CLI entry point. The exclusion is narrow and
named, but it means apps/web is the one package whose 100 is over a set of files chosen
by hand rather than over everything under src/**. Worth recording alongside it: seven
type-defensive ?? arms in src/views/layout/cost-flow.ts used to be excluded and are
now genuinely reached, through the exported toFlowNode / toFlowLink converters and an
injectable pathFor seam — the exclusion list shrank by being tested away rather than by
being argued away.
Second, and directly downstream of the first, apps/web’s honesty test is weaker than the
other four. It sweeps src/ for pragmas and asserts the offender list is empty, but it
does not assert the four thresholds and it does not assert the absence of further
exclude entries — it cannot, since the package legitimately has one. The practical
consequence is that a change widening the web exclude list, or lowering the web
thresholds, would not trip a guard. That is a real gap in the mechanism, not a
technicality.
(Closed 2026-09-26. The web coverage honesty block in apps/web/test/honesty.test.tsx
now also pins the evaluated shape of vitest.config.ts, read as text like the other four:
the thresholds object must be exactly lines/branches/functions/statements at
100 with no extra key, include exactly 'src/**', and exclude exactly
['src/main.tsx', 'src/vite-env.d.ts'], each of which must exist on disk. It cannot assert
the absence of an exclude, so it pins the list instead — the asymmetry in form stays,
the gap in effect is gone. Four mutations were each killed: branches: 99, a third
exclude entry, an extra perFile threshold key, and the exclude moved into a variable.
Importing the config was tried first and rejected: vitest/config does not load under
jsdom, and the file is outside the web tsconfig project.)*
Third, 100% means 100% of src/** — not of the repository. Every package’s coverage
include is src/**, so two areas of live code never enter a denominator at all:
hooks/install.mjs, which is exercised in earnest by apps/server/test/hooks-installer.test.ts
against a throwaway temp directory but is measured by nothing; and the two CI gate scripts
scripts/check-no-spawner.mjs and scripts/check-licenses.mjs, which have no unit tests
whatsoever and are exercised only by being executed in CI. (Superseded in part, noted
2026-09-26: since 4b3cd2d each gate — and the Node-major guard
scripts/check-node-version.mjs — exposes a pure core (scanTree / evaluateLicenses /
evaluateNodeVersion plus a formatter) that apps/server/test/scripts-gates.test.ts drives
against throwaway fixture trees, so “catches what it claims to catch” is now tested, not only
“exits 0 on this repo”. What still holds: the scripts sit outside every coverage
denominator, and their CLI wrappers are deliberately left unexecuted by tests, because
running them would need another subprocess exemption.) Both gates do run on every CI
invocation and both currently pass — the spawner gate reports OK (235 files scanned
across 4 roots + repo-root config; 1 allowlisted) and the license gate OK (412 installed
packages, all licenses allowlisted), measured 2026-08-15; re-measured 2026-09-19 as
OK (272 files scanned across 4 roots + repo-root config; 1 allowlisted; 6 package.json
manifests checked for forbidden direct dependencies) and OK (412 packages / 429
installed versions; 411 allowlisted, 1 under a documented exception) — but “the gate
script works” is established by its output, not by a test of the script.
(Re-measured again 2026-09-23: the spawner gate now scans 282 files across the same
4 roots + repo-root config, still 1 allowlisted and 6 manifests checked, and the license
gate is unchanged at 412 packages / 429 installed versions, 411 allowlisted, 1 under the
documented exception. The file count tracks the sources; it is a dated measurement, not a
constant, and neither gate’s verdict has moved.)
(Amended 2026-09-23 (J-12): the OK lines quoted above also predate the spawner gate’s
current output shape. Its real final line now reads check-no-spawner: OK (282 files scanned
across 4 roots + repo-root config; 1 allowlisted; 5 line(s) inline-exempt; 6 package.json
manifests checked for forbidden direct dependencies), and the run prints one inline opt-out
in force line per exempt line plus two inline opt-out suppressed nothing (dead marker, not
fatal) lines above it. The five exempt lines carry the three sanctioned exceptions - two
exceptions need the marker on the import as well as on the call. Both extra clauses are
reporting, not failure: exit code 0.)
(Amended 2026-09-24: the gate now also checks every package.json scripts string for a
bind wider than loopback, flags imports of more subprocess-wrapper packages, and flags the
vm module. Its final line now ends 6 package.json manifests checked for forbidden direct
dependencies and wide-bind scripts); measured at 308 files scanned, still exit 0.)
(Amended again 2026-09-24 (II4): the license gate counted a package once per license it
appeared under, so a name installed at two versions under two licenses counted twice. It now
counts unique names, and the line reads OK (407 packages / 429 installed versions; 406
allowlisted, 1 under a documented exception). The dependency tree did not change; the earlier
412 / 411 figures were the double count. All three gate scripts also now recognise themselves
when invoked through a symlinked path, where they previously exited 0 without checking.)
(Re-measured 2026-09-26: check-no-spawner: OK (313 files scanned across 4 roots + repo-root
config; 1 allowlisted; 5 line(s) inline-exempt; 6 package.json manifests checked for forbidden
direct dependencies and wide-bind scripts) and check-licenses: OK (406 packages / 428 installed
versions; 405 allowlisted, 1 under a documented exception). Both exit 0; the package count moved
with the dependency tree, not with the gate. Later the same day the spawner line gained the
no-SSRF clause ; 93 server-process files checked for outbound network calls).)
And the standing caveat that no coverage number escapes: 100% line and branch coverage records that every line and branch executed, not that every behaviour was asserted. It is a floor under the test suite, not a statement about its depth. What gives this suite its actual weight is the material in §3.1, §4 and §5 — an independent reader in the token proof, byte-identical replay snapshots, and twelve enumerated ways the system is expected to fail well.
| Phase | Wave(s) | Testing-relevant exit gate |
|---|---|---|
| 0 — Feasibility spike | 1–4 | WP-S1’s paired corpus + Ivan-labeled trees captured; WP-S7 reads GO/CONDITIONAL-GO on the evidence they produced. |
| 1 — Foundation, security spine, storage, ports | 6–8 | Coverage harness + CI gate green & blocking (WP-F3, WP-F4, WP-X5); golden fixture corpus promoted and labeled (WP-X1, WP-X2); WP-F7’s security contract tests exist, deliberately red. |
| 3 — Projection, the DAG moat, reconciliation, cost | 14, 16 | Three P0 tests green & merge-blocking; hierarchy ≥95% vs. the labeled corpus; 12-scenario negative catalogue green (WP-X3, WP-X4, WP-IN13). |
| 4 — Read API + SPA | — | Web package counts toward the >90% gate. |
6 — Alert-track release hardening (post-1.0 per best-path §6.1; the operator alerts UI WP-A8/WP-A9 was cut per §6.2) |
17 | Alerts modules >90% covered; RELEASE.md enumerates every CD-7 gate with a verification step (WP-A10, WP-X9). |
(development-plan §3, §4.) The hard structural point: none of this is a procedural
checklist an agent could skip under time pressure — WP-F1 (the monorepo scaffold
itself) has a real dependency edge on WP-S7, and WP-IN13/WP-X3 are wired as
blocking CI checks, not advisory ones (development-plan §1, §5). Two corrections to
that table from the as-built note: the coverage figure it calls “>90%” is 100 in every
package that shipped, and “blocking” described the intent rather than the mechanism until
2026-08-25 — the workflow ran on every push and pull request, but main was unprotected, so
nothing physically withheld a merge. The mechanism now exists: main requires the ci
check, so a red run withholds the merge button from a contributor — but not from the
repository owner, because enforce_admins is deliberately off in a repository with one
maintainer whose normal working mode is a direct push to main, and enabling it would lock
that sole maintainer out (see
the standing correction).
packages/core/test/negative/ (parser and cost facets) and apps/server/test/negative/
(HTTP, persistence and security facets). Two source-level readings the recovery
surfaced are settled in the tests themselves: scenario #7 is asserted against SSE
(CD-5 supersedes the source’s “WebSocket”, and the test proves it by checking the
text/event-stream content type), and scenario #8 holds at both times — build-time
no-price-row-FAILS-CI and a runtime halt that never degrades to a silent $0. What
is genuinely still open is the clause of scenario #2 that asks for a “normalized
anomaly flagged”: there is no hook normalizer to flag it in, so the anomaly is observed
through the watchdog instead and the gap is recorded in the test file.packages/test-fixtures,
which is now a real workspace package carrying eight typed fixtures, the annotation
corpus, and its own coverage gate.apps/server/src/retention/, and
the owner signed the v1.0 numbers on 2026-09-08 (D3) — events 90 days, token_usage
never, backup files 30 days behind a floor of 7 — wired into the composition root on
2026-09-10 and run after each successful daily backup. The library default is still a
byte-identical no-op; the server no longer uses it. See
backup & restore.token_usage.agent_id (WP-S3, G0.1b) turned out not to need
the confidence-scored heuristic the open question contemplated. Attribution is a hard
structural key: extractUsageRows() in packages/core/src/parser/parse-session.ts
stamps each usage row with the owner of the transcript file the line was read from, and
usage in the main transcript is stamped null rather than attributed to a guess. That
is why the P0 token-reconciliation test can assert integer equality per session, per
model, per bucket instead of a tolerance. What has not been ratified is the accuracy
of the surrounding hierarchy attribution — that is the §3.1 gate, and it is unmeasured.events_raw
substrate, idempotency keys, and replay-on-startup mechanics the three P0 tests prove
correct.orchestration_edges
derivation the DAG-rebuild test and the deep-nesting/two-instances pathologies exist
to validate.PreCompact re-pricing case and the
no-priceless-model-fails-CI gate (WP-C6).development-plan.md — the full Track X
work-package catalog (WP-X1…WP-X10; WP-X11 deleted per best-path §6.3), waves,
and Global Definition of Done (§8).concept-analysis-v2.md §4.3, §6 — the QA
lens verdict and the quantified acceptance criteria in full.