agenthropic

Parser specification — implementation-ready, derived from the Phase-0 spike (WP-S1…S7)

Status. This is the normative parser contract, distilled from the completed Phase-0 feasibility spike (WP-S1…S7, verdict phase0-verdict.md, 2026-07-10). It turns the spike’s empirical findings into the specification the production parser and its golden-fixture tests are written against — one place, so the parser is not reconstructed from six scattered spike/*/README.md files when the build starts.

IMPLEMENTED as of 2026-08-15. This is no longer a proposal on paper. The contract is realised in packages/core/src/parser/, persisted by the eighteen migrations in apps/server/src/db/migrations.ts, and driven over the real corpus by apps/server/src/corpus/ingest-corpus.ts; all fourteen gate items in §3 have code behind them. Implemented is not measured, certified or ratified, and this file keeps those axes apart on purpose (§3). Every number below remains PROVISIONAL / self-check — scored against machine inventories, not human ground truth — until Ivan’s LABEL-ME.md sign-off lands, and the WP-S7 verdict it rests on is still CONDITIONAL GO.

Historical note (2026-08-15). This block used to read “this is a design document, not code … CD-8 still binds”. CD-8 was overridden by the owner on 2026-07-11 — for dispatching only, recorded in §8 item 1. The override released the scaffold, not the evidence: the security invariants, the LABEL-ME ratification requirement and the never-commit-without-an-explicit-ask rule survive it untouched.

Authority. Where this document and the pre-spike architecture pages disagree, the spike evidence wins for the parser mechanics it measured — but this file does not silently rewrite the site docs; §7 lists the exact edits to fold into ingest-reconciliation.md, hooks.md, cost-model.md, dag-moat.md, and data-model.md. Those edits were held behind the scaffold gate; that gate opened on 2026-07-11 and they are now due (§7).


1. Scope

This spec covers the read side only — the pure function on-disk JSONL substrate → (agents, orchestration_edges, token_usage). It does not cover the substrate contract (events_raw, idempotency keys, replay-on-startup, the Normalizer/Projection split) — that is fixed by CD-2/CD-3 and documented in ../site/architecture/ingest-reconciliation.md, unchanged by the spike. The spike confirmed CD-2/CD-3 hold; it sharpened what the reconstruction step inside them must do.

The corpus it was proven against: 5 sessions, 224 agents, 3 projects, deliberately loaded with the four hostile pathologies (post-compaction block eviction, live-crash / no-Stop snapshot, same-slug concurrency, dual-layout depth-2).


2. Input substrate — what the parser reads

Per session, from ~/.claude/projects/<slug>/ (read-only, always):

Artifact Role
top-level <session-uuid>.jsonl the main-agent transcript; carries parent-side Agent/Workflow tool_use spawn blocks, usage rows for the main agent’s own turns, and compact_boundary records
subagents/.../agent-<hex>.jsonl one per subagent; carries that agent’s usage rows (inline agentId == <hex>) and, for nested layouts, its child spawns
agent-<hex>.meta.json per-agent sidecar, a single-line JSON object, one beside each subagent transcript. Flat subagents: {agentType, description, toolUseId, spawnDepth} — its toolUseId is the primary child→parent join anchor (§4). Workflow subagents: {agentType, spawnDepth} (and, in worktree-isolation mode, worktreePath) — no toolUseId, so they join by directory instead. Measured on real data (1855 sidecars): spawnDepth ∈ {1: 1842, 2: 11}; toolUseId present on 282, absent on 1573; worktreePath on 19. (The earlier “promptId, layout, model” description was wrong — those keys do not exist; retracted.)
journal.jsonl (subagents/workflows/wf_* dirs) per-workflow dispatch journal; started line-order + content-hash key

Two on-disk layouts exist and the parser branches on directory shape, not version string (gate item #2, #10):

A single session may be mixed (b24be30c: 20 flat / 22 nested). Branch per-agent-file, not per-session.


3. The parser-requirements gate — final status (14 items)

The original probe defined an 11-item gate. The spike exercised it and discovered three additional load-bearing requirements (N1–N3) the original list did not have. All 14 are normative MUSTs for the production parser.

# Requirement Status Proved by Affects
1 Spawn edges key on Agent/Workflow tools, never Task (0 Task blocks corpus-wide) ✅ green S2 WP-IN8
2 Walk both layouts (flat + nested wf_*); branch on directory shape ✅ green S2 WP-IN8
3 Support the structural join schemas (see §4) ✅ green — sidecar-anchored, branched on layout; real-data 1855/1855 = 100% (§4.2, PROVISIONAL) S2, S3 + real-data WP-IN8
4 Index subagents as self-referential candidate parents (depth-2 parents live inside depth-1 agent transcripts, not ROOT) ✅ green S2, S5 WP-IN8, data-model
5 Match block ids by structural equality, never substring (a toolu_/hex id also appears in prose — substring join forges false edges) ✅ green S3 WP-IN8
6 Sum tokens from child transcripts; parent rollup is ≈0% (measured 0.00%, disjoint message.id sets) ✅ green S6 WP-IN7, WP-IN9
7 Legacy 2.1.70 bare-Explore fallback ✅ implemented (defensive, 2026-08) — narrowest reading: bare {agentType:'Explore'} sidecar (no toolUseId/spawnDepth keys) joined via a raw top-level agentId on a foreign progress record, only when every modern anchor misses; edge carries the DISTINCT legacy_explore provenance (persisted verbatim). Exercised by synthetic fixture legacy-bare-explore (packages/test-fixtures); still absent from corpus, so the shape stays PROVISIONAL until a real pre-2.1.71 transcript ratifies it; pointer session site/08871133-82a3-4ae2-8303-781a8761e92a S1 WP-IN8 (defensive)
8 Handle compaction resets (compact_boundary + compactMetadata, JSONL-native) ✅ green S2, S4 WP-IN8, cost
9 Concurrency-safe: key on session-uuid, never slug (two same-slug concurrent sessions must stay two roots) ✅ green — hardened 2026-08 for the mirror case, one uuid under several slugs: enumeration keeps one deterministic ref and records every shadow copy as a duplicate-session skip (§4.3); the mirror-mirror case (a renamed copy, whose file name no longer collides but whose records still name the original session) is refused pre-write as session-id-mismatch (§4.3) S2, S5 WP-IN8
10 Version-detect / branch-on-shape (not on a version string) ✅ green — the version-detect half added 2026-09-26: until then only branch-on-shape existed and no record’s version was read. parseSession now returns claudeCodeVersions (sorted distinct per-record version strings, provenance only; parsing still never branches on it). Not persisted — no reader needs it yet S2 WP-IN8
11 Intra-workflow sibling ordering (EMP-1) ✅ amended — see §6.3; original “total order via journal+promptId” is false, order is wave-partial only S5 dag-moat, UX
N1 <task-notification> as a flat join path — a legacy child-side re-anchor when the parent-side tool_use block was evicted ✅ implemented (defensive) — extractTaskNotificationToolUseId in packages/core/src/parser/parse-session.ts, persisted as the task_notification provenance; absent from the real corpus (0/1855 edges), exercised only by the synthetic fixture task-notification-recovery (its spike “load-bearing” role was fixture-only) S2; real-data WP-IN8 (defensive)
N2 queue-operation as a hard join schema for run_in_background (queued) Agent spawns whose parent tool_use block is never materialized ✅ green — but marginal: 3/1855 on real data S3; real-data WP-IN8
N3 Dedup usage rows by message.id before summing, and price per-bucket/per-model ✅ green (reconciled, PROVISIONAL) — the §5.2 byte-identical premise was false (lines sharing a message.id are streaming partials whose output grows and whose first row may carry a transient fast-mode model label), so dedup collapses to the per-bucket maximum and settles the model to the greatest-output row; UsageConflictError is reserved for a genuine collision — two distinct models tied at the same max output, or any agentId clash (§5.2, §8) S6; real-data WP-IN7, cost

Real-data note (PROVISIONAL — LABEL-ME). The “green” marks above were originally self-scored against 5 synthetic spike fixtures. Re-measured read-only over all ~/.claude/projects/* (141 sessions), edge reconstruction is 1855/1855 = 100% (§4.2), but two premises did not survive contact with real data: N1 <task-notification> fires 0 times, and N3’s identical-usage-per-message.id assumption is false on two axes — output grows and the first partial can carry a transient fast-mode model label (§8) — N3 is now reconciled to per-bucket-max collapse plus greatest-output model settle (§5.2). With both reconciled, a full read-only re-measurement with usage intact parses all 141 sessions end-to-end (0 UsageConflictError, 0 SubstrateError, 0 orphan subagents).

Two axes, deliberately not collapsed into one number (2026-08-15). Gate coverage is reported on two independent axes, because a single figure would let “we built it” pass itself off as “the corpus proved it”:

So the honest one-liner is “14 implemented, 11 measured, 2 defensive-and-unwitnessed, 1 amended”. Do not report “14/14 green.” Two of those items are fallbacks whose on-disk shape no real transcript has ever confirmed; their shape stays PROVISIONAL until a genuine pre-2.1.71 transcript ratifies it, and if one ever arrives and disagrees, the fixture is what was wrong. The rest of the caveats stand unchanged: the numbers are self-check, and the LABEL-ME ratification has not happened.


4. Join schemas — sidecar-anchored resolution, branched on layout

RETRACTED — “224/224 (100%)”. The prior version of this section claimed spawn-edge resolution reached 224/224 across four co-equal paths, with <task-notification> (N1) as the load-bearing compaction-recovery path and a fictional meta.toolUseId↔queue handshake. That number was self-scored against 5 synthetic spike fixtures that encoded anchors which do not match real on-disk shapes; it is withdrawn. Measured against real data (§4.2) the distribution is entirely different: the flat anchor is the sidecar toolUseId, <task-notification> never fires (0/1855), and queue-operation is marginal (3/1855). The real numbers below are PROVISIONAL — pending owner (Ivan) ratification (LABEL-ME).

4.1 The resolution model (as implemented)

Resolution branches on layout (gate #2), keying edges on structural id position, never a substring (item #5):

The provenance set is closed at the storage layer. Migration 13 rebuilds orchestration_edges with CHECK (source IN ('tool_use','directory','task_notification', 'queue_operation','legacy_explore')) (apps/server/src/db/migrations.ts). A sixth join path cannot be introduced by a parser edit alone: it needs a migration, which is a deliberate speed bump on exactly the change that would otherwise dilute provenance quietly.

4.2 Real-data measurement — census of record (PROVISIONAL — LABEL-ME)

Census of record (note added 2026-08-15). This block supersedes the 2026-07-04 probe census (17 projects / 117 sessions / 148 flat + ~849 nested agent files) still quoted in phase0-probe.md §1, README.md, corpus-audit-2026-07-06.md §9.5 and ux0-design.md §3. Cite this table, not those. Note the unit carefully: 1855 counts subagent TRANSCRIPTS, one per agent-<hex>.jsonl file — the session count is 141, of which 54 have any subagent at all. Conflating the two turns a per-file rate into a per-session claim roughly thirteen times larger than reality.

Measured read-only over all ~/.claude/projects/* (20 slugs, 141 sessions, 54 with subagents, 1855 subagent transcripts) — not the 5-session spike corpus. At the time of this measurement the end-to-end parser still threw UsageConflictError on every subagent-bearing session (§8, a real-data violation of the since-retired §5.2 identical-usage assumption), so the edge pipeline was exercised with usage blocks neutralized, isolating reconstruction from the unrelated usage gate; edge resolution reads no usage field, so this did not perturb it. That conflict has since been reconciled (§5.2: per-bucket-max collapse plus greatest-output model settle) and a re-measurement with usage intact now parses all 141 sessions with the same edge numbers — the neutralized run is retained here because it is what the table below was actually taken from.

Metric Value
Edge reconstruction rate 1855 / 1855 = 100.00%, 0 orphans
directory (workflow subagents) 1571 (84.7%)
tool_use (flat, via sidecar toolUseId) 281 (15.1%)
queue_operation (run_in_background flat) 3 (0.16%)
task_notification 0 (0.00%)
parent = main / parent = another agent (depth-2) 1844 / 11

Flat total 284 = 281 tool_use + 3 queue_operation + 0 task_notification + 0 orphan; nested total 1571 = all directory. The join key is intentionally cross-file (a tool_use.id / sidecar toolUseId matching an inline hex in a child file) — which is exactly why matching must be on structural id position, never a substring grep.

The legacy_explore path is not in this table because it contributed 0 edges: nothing in the corpus is old enough to reach it (§3).

4.3 One session uuid, several slugs — the duplicate-session rule (hardens gate #9)

Gate #9 says: key on the session uuid, never the slug. The corpus supplies a case that rule alone does not settle — the same uuid appearing under more than one slug directory, which is what a copied or backed-up project dir produces (~/…/myproj and ~/…/myproj-backup both holding <uuid>.jsonl). Keying on the uuid is still right; the question is which of the copies the uuid refers to.

Enumeration answers it before any parsing happens (dedupeSessionRefs in apps/server/src/corpus/disk-substrate.ts): exactly one ref survives per uuid — the one whose projectSlug sorts first lexicographically — and every losing copy is recorded as a duplicate-session skip. Two properties matter more than the choice itself:

duplicate-session therefore joins the other skip reasons (oversize, symlink, not-regular-file, unreadable, empty-agent, empty-main, non-artifact) under one posture: a file the ingest declined to read is reported, not forgotten. A skip counter that quietly reads zero while the corpus is being half-ingested is the failure this design exists to prevent.

AMENDED 2026-09-23 (J-6). That enumeration is one short, and because this section is normative for ingest work the omission is a gap in the contract rather than a prose slip. The union has a ninth member, too-deep, and an implementation that does not emit it does not satisfy §4.3.

The mirror case — one file name, a different uuid inside (session-id-mismatch)

dedupeSessionRefs settles collisions between file names. It cannot see the other half of the same hazard, because a renamed copy no longer collides: cp <uuid>.jsonl <other>.jsonl leaves a file whose name is new but whose records still carry the original sessionId on every line. Nothing on disk forces a transcript’s two identities to agree.

That divergence is not cosmetic. The enumerator keys the ref (and the tail-follow fingerprint) on the file name; parseSession derives the id the writer keys every row on from the records. Ingest the copy and its agents, edges and usage upsert into the original session from a second file, its dispatch events re-fire, and — because the copy is a distinct fingerprint that never reaches a settled state — it does so again on every tick, indefinitely.

Rule: before any write, the id the substrate’s records declare must equal the uuid the file name declares. A mismatch is refused; the session is not ingested. Three details are load-bearing:

Measured against the corpus available on this machine at the time of writing (21 slugs, 51 main transcripts, read-only), the rule rejects nothing: every main transcript’s declared sessionId equals its file-name stem, and no transcript carries two different sessionId values. The guard is therefore a containment rule for a hazard that user action creates (copy, rename, restore-from-backup), not a correction of anything Claude Code writes.


5. Token → cost

5.1 Attribution is a hard field-read, not a join

token_usage row → agent_id is 6654/6654 = 100% hard key, 0 heuristic: the agentId field on the row is byte-equal to the hex in the enclosing agent-<hex>.jsonl filename. This makes CD-3’s backfill a hard join, not a confidence-scored inference — no UI-surfaced uncertainty is required for attribution.

5.2 Dedup by message.id is a correctness gate (N3)

Claude Code writes one JSONL line per content block, and lines sharing a message.id are streamed partials of one message: the input and cache buckets are constant across them while output_tokens grows toward the final total (verified raw, e.g. output_tokens 7→7→309 on one message). Dedup therefore collapses each message.id to its per-bucket maximum — equivalently the final streamed state — never a sum. Naive row-summation over-counts by ~2.4–2.7× (this corpus: 8,540 raw usage rows → 3,339 deduped messages); summing without dedup produces ~$900 of phantom spend where the true figure is ~$346. The parser MUST collapse usage to the per-message.id per-bucket maximum before summing. This is correctness, not optimization.

PROVISIONAL (LABEL-ME). This supersedes the original §5.2/N3 premise that the partials were byte-identical and that ANY usage, model, or agentId divergence was a loud UsageConflictError. Real transcripts (the full ~/.claude/projects corpus, 141 sessions) disprove the premise on two axes, both reconciled by the same rule — the final streamed row is ground truth:

  1. Usage — only output grows across the partials, so usage collapses to the per-bucket maximum (the final total) instead of asserting equality.
  2. Model — the first partial can carry a transient fast-mode routing label (observed: claude-fable-5 on the first row, then claude-opus-4-8 on every later row and the final row, sharing one message.id and one requestId; output 2→…→6747 / 9→…→360). Fast mode runs Opus (it does not downgrade), so the settled label is the true model. The parser therefore settles the model to the greatest-output row. This is pricing-relevant, not cosmetic: naively taking the first partial’s fable-5 would price the message at the fable rate ($10/$50 per MTok) instead of the true opus rate ($5/$25) — a ~2× over-charge. Measured impact: 2 of 141 sessions (both in servicenow-preflight) previously threw here; with the settle, 0/141 throw and the edge rate holds at 100%.

What stays loud (genuine corruption streaming cannot produce): two distinct models tied at the same maximum output, or any agentId disagreement on a shared message.id (§5.3 makes message ids per-agent-disjoint). Selecting the final-row model/usage is choosing ground truth, not inferring it, so it is defensible under the ground-truth invariant — but both loud-failure relaxations (usage-max and model-settle) are flagged for Ivan’s ratification and the corpus numbers stay PROVISIONAL. The §5.4 “halt loudly on an unknown model id” pricing MUST is unchanged — the settle only chooses between two known models, and a genuinely unknown settled id still halts.

5.3 Parent-rollup is 0.00% (gate #6)

Parent usage rows are the main agent’s own turns (agentId=null, isSidechain=false); child usage lives only in agent-<hex>.jsonl. A subagent’s result returns to the parent as a tool_result (user row, no usage block), so it is never re-counted. Parent and child message.id sets are disjoint → provably no double-count. Partition identity sum(per-agent) + ROOT == session total holds to <1e-6 USD in all 5 sessions.

5.4 Pricing must be bucket-and-model-aware (N3)

88% of corpus tokens are cheap cache reads — a flat per-token rate would be wildly wrong. Price each of the five buckets (fresh input, output, cache-write-5m, cache-write-1h, cache-read) at the model’s rate. That set of five is single-sourced as the TokenBucket union in packages/shared/src/types/enums.ts, and the token_usage.bucket / model_pricing.bucket CHECK constraints are generated from the same five-member list — so “how many buckets are there” has exactly one answer in the codebase. (The earlier “four” in this sentence was a miscount against its own parenthetical; corrected 2026-08-15.)

The spike’s approximate table (list prices, for a mechanism proof — not a billing source): opus-4-8 $5/$25, sonnet-5 $3/$15 (list; the intro $2/$10 through 2026-08-31 would cut Sonnet ~⅓), fable-5 $10/$50, haiku-4-5 $1/$5 per MTok; cache-read ×0.1 of input, cache-write ×1.25 (5m) / ×2.0 (1h). <synthetic> = $0. The parser MUST halt loudly on an unknown model id — never silently price it at $0.

Shipped as the model_pricing seed (note added 2026-08-15). This table is no longer only a spec proposal: migrations 7 and 11 in apps/server/src/db/migrations.ts INSERT exactly these rates into model_pricing (model, bucket, usd_per_mtok, effective_from), with the derived multipliers applied per bucket; migration 18 (2026-09-10) then added ten explicit per-bucket rows for the two model ids the real corpus exposed as unpriced, claude-opus-5 and claude-fable-5-1 (Fable 5.1’s cache-read rate is 0.025× input, so those rows are not derived by the seed’s multipliers); migration 19 (2026-09-26) added five more for claude-opus-5-5 (cache-read 0.05× input, a third ratio), the id a boot at schema 18 that day found refusing 27 of 61 sessions; migration 20 (the same day) rewrote the five claude-sonnet-5 rows to the official 2 / 10, because the pricing page then said the intro price above had become the standard one and the 3 / 15 increase would not occur (decision D10). Three details a reader should not have to reverse-engineer:

The rates themselves remain PROVISIONAL and unratified; they are frozen against in-place editing by the per-migration sha256 checksum, so changing a price must ship as a new migration carrying its own data rather than as a quiet constant edit (that exact edit is how migration 7 once diverged from an already-migrated operator database).

Corpus total (approximate): ≈ $345.91 over 206,001,429 tokens / 3,339 messages.


6. Tree construction

6.1 Self-referential parent index (gate #4)

Depth-2 parents live inside depth-1 agent transcripts, so the parent index must admit agents as candidate parents, not only ROOT. Witnessed by two independent sessions (b24be30c 6/6, f28af3fd 5/5). Depth is capped at 2 corpus-wide — metas {1: 1339, 2: 11, none: 2}, no depth-3 anywhere. Deeper trees are therefore unmeasured, not proven absent: the parser must handle depth-N by construction (recursive index), but the ≥95% hierarchy bar has only been demonstrated to depth 2.

The ≥95% bar currently reports NOT CERTIFIED (note added 2026-08-15). The gate itself is built — packages/test-fixtures/src/annotations/ scores parsed parents against labelled ones and prints a verdict — and it is deliberately hard to satisfy. It scores on the Wilson one-sided 95% lower bound, not the naive ratio, because the naive ratio is silent about n: 3/3 and 300/300 both read “100%” and only one of them is evidence. Solving for a flawless run gives a minimum of 52 labelled claims. And it refuses fixture data outright — a synthetic-by-construction corpus is marked NOT ADMISSIBLE, since machine-authored truth scored against the same machine’s parser proves only self-consistency. The hand-labelled LABEL-ME corpus does not exist, so the gate has never had an admissible sample and returns NOT CERTIFIED. It cannot be closed by an agent: producing the ground truth is Ivan’s act.

6.2 Concurrency independence — key on session-uuid (gate #9)

69ac12d0 (103 agents) and a362e15d (57) share slug -Users-ivanbaev-Development- agenthropic and overlap ~2.5 days of wall-clock, yet are two independent roots: empty hex intersection, distinct roots, zero cross-session edges. Keying on slug would fuse 160 agents into one phantom tree; keying on session-uuid keeps them apart. This is the one tree fact Ivan most needs to confirm by hand (LABEL-ME.md).

6.3 Sibling ordering — EMP-1, amended (gate #11)

The original gate assumed journal.jsonl + promptId yields a total sibling order. The spike disproves that:

Normative rule: order siblings by first-record timestamp; collapse same-wave siblings (Δt below ~10 ms) into an explicitly UNORDERED concurrent group (“N started together”); use journal started order only as corroboration, never for flat layout (no journal.jsonl exists there); do not manufacture a total order the substrate does not carry. Ship agent → workflow → ROOT with time-ordered waves. (This is “honest uncertainty over confident nonsense” made concrete — see ux0-design.md.)


7. Edits to fold into the site docs — the gate has opened

Status (2026-08-15). This section used to end with “do not apply these now”. The gate it was waiting for — CD-8 — opened on 2026-07-11 by the owner’s override, the code that implements these results has since shipped, and the site-docs lane is applying them. The edits below are due, not held. The list itself is kept verbatim so the audit trail survives: this file remains the source each edit must be checked against, and the underlying WP-S7 verdict is still CONDITIONAL pending human sign-off, so any page that adopts these results inherits the PROVISIONAL label with them.

Fold these spike results into the published architecture pages (this file is the source for each edit):

Two additions the original list predates, both now shipped and both belonging with the same sweep: the fifth edge provenance legacy_explore and the migration-13 five-value source CHECK (§4.1), and the duplicate-session enumeration rule with its /api/health.ingestSkips surface (§4.3).


8. What still gates production code

  1. CD-8 — overridden, not passed. The original rule was: no package.json/src/ until Gate A is signed and WP-S7 GO is ratified. On 2026-07-11 Ivan explicitly overrode it and authorised implementation; code has shipped since. The override is narrow — it released dispatching only. It did not sign Gate A, did not ratify WP-S7, and did not turn a single PROVISIONAL number below into a fact. This spec stays the contract, and the LABEL-ME sign-off in item 2 is still outstanding.
  2. The “224/224 (100%)” self-score is WITHDRAWN. It was scored against 5 synthetic spike fixtures encoding anchors that do not match real on-disk shapes; it never measured real data. It is replaced by a real, read-only measurement over all ~/.claude/projects/* (141 sessions, 54 with subagents, 1855 subagent transcripts): edge reconstruction = 1855/1855 = 100%, 0 orphans — breakdown directory 1571 / tool_use 281 / queue_operation 3 / task_notification 0 (§4.2). This number is PROVISIONAL — pending owner (Ivan) ratification (LABEL-ME). If Ivan re-parents any edge by hand, it moves.
  3. Two real-data caveats the edge number rests on (do not overstate).
    • The end-to-end parser previously threw UsageConflictError on real subagent-bearing sessions on two streaming axes, both now RECONCILED by “the final streamed row is ground truth” (§5.2): (a) usage — lines sharing one message.id carry growing output (e.g. 7→309), which the old identical-usage assertion rejected; dedup now collapses each message.id to its per-bucket maximum. (b) model — the first partial can carry a transient fast-mode label (claude-fable-5) that settles to the real model (claude-opus-4-8) on the same message.id/requestId; dedup now settles the model to the greatest-output row (2 of 141 sessions, both servicenow-preflight, threw here). A genuine collision still throws (two distinct models tied at the same max output, or an agentId clash). A read-only re-measurement with usage intact now parses all 141 sessions (0 UsageConflictError, 0 SubstrateError, edge rate 100%, 0 orphans). Both reconciliations and the corpus cost numbers stay PROVISIONAL pending Ivan’s ratification (LABEL-ME).
    • The pure parser correctly throws SubstrateError on any non-JSONL file, so the WP-IN5 disk adapter MUST feed only the four §2 artifact types; handing it the whole session directory (tool-results/*.txt, workflows/scripts/*.js, *.pdf) throws on 32 sessions.
  4. KC-0 — Gate A’s two physical Step-0 acts (friction log; rival-dashboard trial) are Ivan’s, not machine-closable, and are unrelated to the technical evidence above. The spike de-risks the build; it does not satisfy the governance checkpoint.

See also