agenthropic

Implementation Review — 2026-08-09 (owner-ordered)

Ordered explicitly by the owner in chat. This document reviews the shipped implementation at commit 2f8d103 and is not design-analysis #10 — the roadmap §8 design-analysis freeze (9 of 9) remains intact. All spike/bench numbers quoted anywhere in this report are PROVISIONAL pending LABEL-ME ratification.

Review date: 2026-08-09 (between KC-1, passed under recorded owner override, and KC-2 due 2026-09-14). Ten dimensions were reviewed; findings below carry one of three verdict labels: CONFIRMED (adversarially verified in code), PLAUSIBLE (credible from code, not settled by a verification pass), and UNVERIFIED-LOW (low severity, not verified). REFUTED findings were removed before synthesis.


Executive summary

The implementation is solid and unusually honest — the verdict is positive. The security invariants hold as coded, not just as documented; the cost engine enforces “no silent $0” at every boundary with the halt gate ordered before any DB write; and the honesty discipline (coverage-honesty guards, as-built doc boxes, a UI that renders uncertainty instead of coercing it) is real and rare. The three most important strengths: (1) defense-in-depth on the security invariants — loopback bind enforced twice, auth on the routed path, TOCTOU-hard read-only corpus access; (2) cost integrity end-to-end — token ground truth, no stored dollars, idempotent replay proven byte-identical; (3) test quality — mutation-killing negative suites on real infrastructure. The three most important weaknesses: (1) CONFIRMED: migration 7 was edited in place after being applied — the live DB at data/agenthropic.db still holds the old seed, so under HEAD whole-corpus ingest sinks on that DB while a fresh DB works; (2) HIGH/PLAUSIBLE: per-file skip diagnostics (onWarning) are never wired in production — an oversize/EACCES transcript silently freezes a session’s dollars; (3) the composed system drops signals its parts emit (ingest-failed SSE frames no client hears) and the synchronous watcher/read paths scale with corpus size, not delta size. Single highest-leverage improvement: corrective migration 11 repairing the model_pricing seed plus a migration checksum — it un-sinks the operator’s actual database today and prevents the divergence class.


Disposition since this review (amendment, 2026-08-15)

This report is a dated record of what the tree looked like on 2026-08-09, and the body below is left exactly as written — a review that quietly edits its own findings once they are fixed stops being evidence of anything. What follows is the delta: the tree has since moved, and a reader who acts on the findings without this section will chase work that is already done.

The method here is narrow on purpose. A finding is marked closed only where the code that implements the fix names the finding in a comment and the mechanism was read and matched against the Fix: line recorded below — fourteen of the twenty-seven weaknesses carry such a reference. Where a fix landed partly, or where the shipped code addresses one half of a two-part finding, it says partial and names which half. Everything else is marked not re-verified: that is not a claim it is still broken, it is a statement that nobody looked again. The Low block was not re-reviewed at all.

Both High findings are closed:

# Status (2026-08-15) What changed
H-1 closed Migration 11 + per-migration content checksum verified on start
H-2 closed SkipReporter wired to onWarning; /api/health.ingestSkips
M-1 closed Gate #7 defensive fallback with distinct legacy_explore provenance, CHECK-constrained by migration 13. Scope stays PROVISIONAL and the shape is unwitnessed in the real corpus — see parser-spec §3
M-2 closed Watcher resolves pricing per tick and resets attempts when the pricing content changes
M-4 closed Restore removes a pre-existing -wal/-shm before opening
M-5 closed usage_by_agent restricted to the selected agent id-list; migration 12 adds the edge endpoint indexes
M-6 closed SERVER_EVENT_TYPES moved into packages/shared and imported by both sides
M-7 closed Root-scope setErrorHandler with 5xx message suppression
M-8 closed Top-burners table shipped (apps/web/src/views/top-burners.ts)
M-9 partial Today/this-week windows shipped (cost-windows.ts); an aggregate delegation-saved figure was not verified as existing
M-10 open, acknowledged in code CostView.tsx names M-10 and records that the fix is a shared clock tick, not a per-view workaround
M-11 closed Argv-free curl delivery (--variable / --expand-header), curl ≥ 8.3.0 floor, fail-closed below it
M-12 closed A first-ingested-session ownership rule, with skipped messages counted rather than silently dropped
M-13 closed Persisted-slug hint lets a late SubagentStop reconcile against the agent row
M-14 closed Duplicate session uuid across slugs recorded as a duplicate-session skip — parser-spec §4.3
M-15 partial A tail-read path and a lastTickDurationMs health field shipped; whether the synchronous full-fingerprint pass is gone was not verified
M-16 closed but incomplete Boot ingest moved after listen and /api/health.ingest reports replaying / idle — but the replay is synchronous, so the endpoint cannot answer for most of that window. Measured below.
M-18 partial crossSessionUsageCollisions exposed on health under this item’s number; the endpoint’s re-enumeration cost was not re-measured
M-20 closed Daily backups wired in the composition root, not only as a manual drill
M-22 partial CI now runs a web production build; the production run path was not verified
M-23 superseded All five packages now pin 100% and all five carry an anti-pragma guard (four named coverage-honesty.test.ts; apps/web’s is the coverage honesty block of test/honesty.test.tsx)
M-3, M-17, M-19, M-21 not re-verified No code in the tree names them
M-24 still open, owner-only The hierarchy-accuracy gate remains unmeasured: the LABEL-ME hand-labelled corpus does not exist, so the ≥95% bar reports NOT CERTIFIED and every Phase-0 number stays PROVISIONAL. No agent can close this — producing ground truth is Ivan’s act
M-25 open, owner-only on 2026-08-15 — closed 2026-08-25 As read on 2026-08-15: branch protection on main was not enabled, so no gate was merge-blocking. The owner act has since happened — on 2026-08-25 main was branch-protected with ci (the job id in .github/workflows/ci.yml, not the workflow’s CI display name) as the required status check, force-pushes and branch deletion refused. A red run now withholds the merge button from a contributor and, deliberately, not from the repository owner: enforce_admins is off because agenthropic has one maintainer whose normal working mode is a direct push to main, and admin enforcement would lock the sole maintainer out of their own repository — see the standing correction. The KC calendar’s other owner-only acts (LABEL-ME, the <30 s stopwatch) are unchanged
Low (all) not re-verified The block was not re-reviewed

The two items at the bottom of that table are the ones worth re-reading. Everything above them was work an agent could do and did; M-24 and M-25 are the findings that cannot be closed by writing code, and they are precisely the ones that gate the project’s honesty claims — an uncertified accuracy number and an unenforced quality bar. Fourteen fixes had not moved them by one inch. One of the two has moved since, and only in the way it always could: M-25 was closed on 2026-08-25 by the owner act itself, not by a commit. M-24 stands exactly as written — the hand-labelled corpus does not exist, the ≥95% gate still reports NOT CERTIFIED at n = 0, and every Phase-0 number is still PROVISIONAL.

Second amendment (2026-08-22)

Same method as above, and for the same reason the table above is left untouched rather than rewritten in place. This section records only what moved after 2026-08-15. (One later exception, dated where it happened: on 2026-08-25 the M-25 row of that table and the paragraph under it were amended, because the owner act they were waiting on finally happened; both keep their 2026-08-15 reading and label it as such.)

Read this before trusting either table. The body of this report is a dated 2026-08-09 snapshot, and it has now caused two misreadings in one session: the body’s M-1 text (line 338) describes gate #7 as unimplemented, when it had in fact already shipped — LEGACY_EXPLORE_EDGE_SOURCE (packages/core/src/parser/types.ts), the branch-5 fallback (packages/core/src/parser/parse-session.ts), migration 13’s five-value source CHECK, and parser-spec §3’s ✅ implemented (defensive, 2026-08) row. The 2026-08-15 table had it right. Where the body and an amendment disagree, the amendment wins, and where an amendment says “not re-verified” the only authority is the tree itself.

# Status (2026-08-22) What changed
M-10 closed apps/web/src/clock.ts — one module-level useSyncExternalStore clock, one reference-counted setInterval shared by every consumer, cached reading refreshed on the 0→1 subscribe transition so a remount after an idle gap cannot read a stopped value. LiveView and CostView both read useNowMs(); the per-view timer and the duplicated CLOCK_INTERVAL_MS are gone. CLOCK_INTERVAL_MS = 30_000 is PROVISIONAL against the review’s 30–60 s guidance
M-18 closed The residue named in the first table is wired: apps/server/src/index.ts hoists a single createTailCachingFs(nodeCorpusFs()) above the ingest branch and hands the same decorator to both the watcher and createSubstrateProvider. apps/server/test/api-substrate-shared-fs.test.ts asserts exactly one decorator exists process-wide, that boot replay warmed it, and that a cost-analysis request reads through it as a tail read with no full .jsonl read
M-21 partial — pricing half closed Migration 14 (model-pricing-canonical-effective-from) rewrites every model_pricing.effective_from to the canonical YYYY-MM-DDTHH:mm:ss.sssZ instant and installs 4 triggers (2 BEFORE guards that RAISE(ABORT), 2 AFTER rewrites) so the column’s lexicographic order is its chronological order — the ordering the SQL resolver’s effective_from <= occurred_at + ORDER BY … DESC LIMIT 1 had been silently assuming. Three of four pinned divergences in apps/server/test/rate-resolver-parity.test.ts graduated to PARITY_CASES, including the one where core priced $30 and the API reported $10. One survivor stays pinned: offset-form-occurred-at — that divergence is on token_usage.occurred_at, written verbatim by ingest, so it needs the write path
M-19 documented and guarded — NOT closed; now the largest measured cost in the system getCostSummary still scans all of token_usage on every request. The L-26 run below measures one GET /api/cost/summary at what it called real corpus scale — in fact ~11x real scale, 1590 sessions / 10.97 GiB, see the scale-label correction below — as 15.62 s of blocked event loop, per click — roughly 100× one fingerprint sweep tick. What shipped is honesty, not a fix: an explicit “read this before recording the finding as closed” block on getCostSummary, an addendum to the costSummaryStateKey docstring stating that false invalidation is the norm under ingest, and apps/server/test/api-cost-summary-equivalence.test.ts, which mutates the ledger five ways and after each one asserts the served summary still equals a cache-cold direct scan. Every sound narrowing of the cache key turned out to be a write-side seam; the one in-process option (PRAGMA data_version) moves only on other-connection commits, and ingest shares the handle, so it would have served stale dollars
M-9 still partial The half the first table could not verify — an aggregate delegation-saved figure — is still absent. The rest of M-9 is in fact done: cost-windows.ts supplies the today/this-week KPIs and every SessionsView row carries an analyse button, so the “only top-5 sessions are analysable” clause no longer holds
M-1 closed (correcting the body, not the first table) See the paragraph above; the first table was already right

One design error in this report’s own improvement plan is worth recording, since acting on it would have shipped wrong dollars. Bucket 2 item 4 proposes a (session, model, day) rollup for M-19. That grain is unsound: when a model_pricing.effective_from falls mid-day, two rows sharing that key resolve different rates, and 1000 tok @ $1 + 1000 tok @ $9 is not 2000 tok at either rate. Any rollup must pin the resolved rate and store tokens — never a dollar figure that outlives the rate that produced it.

L-26 measurement, and the M-17 verdict it settles (2026-08-22)

Bucket 2 item 5 asked for event-loop-delay and concurrent-inject contention phases in apps/server/bench/corpus-scale.ts, precisely so item 6 (M-17) could be decided “from measurement, not speculation”. Both phases now exist and were run. The benchmark was run twice at what this section called real corpus scale — 1590 sessions discovered, 10.97 GiB, 7.16 M token_usage rows — rather than answering M-17 from a linear projection.

Scale label corrected — 2026-09-01. “Real corpus scale” is wrong, and the figures below should be read at the scale they were actually taken at. The census of record (parser-spec.md §4.2) is 141 sessions; the 1590 here is what you get when the bench is sized by 1855 — the count of subagent transcripts — mistaken for a session count (1855 clones, one fixture in seven planting no root transcript, gives 1590 discovered sessions). So this run covers roughly 11x the real session count and 11x the corpus bytes measured today (996.4 MiB). The numbers themselves are not withdrawn — they were measured, and they are the reason M-17’s verdict is what it is — but they describe a synthetic corpus about an order of magnitude larger than the real one, and every “at real scale” phrase in this section means “at ~11x real scale”. The constant is fixed in code: REAL_CORPUS_SESSIONS = 141 in apps/server/bench/corpus-scale.ts, with the retraction written out beside it. M-17’s conclusion survives the correction and gets stronger — the shortlist was already judged not worth building below ~1100–1250 sessions, and the real corpus is 141.

Two defects in the new instrumentation were found and fixed before any number was trusted, and both matter to anyone reading a loop-delay figure here again:

Measured at ~11x real scale (1590 sessions / 10.97 GiB — see the scale-label correction above; this label said “at real scale” until 2026-09-01), the blocking costs order like this. The ordering, not any single figure, is the finding — and the ordering survives the correction intact, because every row was inflated by the same factor. The individual figures did not survive: cold replay is the one quantity since measured at census scale, and it comes out at 34.87-39.92 s, an order of magnitude below the row below — measured over a synthetic corpus whose per-session size is 3.7x below the real one, so read it as a lower bound (BENCH-SHAPE). Read the table as a ranking, not as magnitudes.

What Blocked event loop Trigger
GET /api/cost/summary 15.62 s every click (M-19)
cold replay 341.7–370.1 s every boot (M-16)
GET /api/dag/global 3.62 s every click
GET /api/sessions 489.7 ms every click
one fingerprint sweep tick 148–178 ms every 3 s (M-17)

M-17 verdict: do not build the fingerprint shortlist now. The sweep is linear and cheap — 0.08 ms/session at 172 sessions, 0.09 ms/session at 1590, a constant per-session cost across a 9.2× range. At that ~11x scale it is a 4.9–5.2% duty cycle, and its effect on in-flight reads is ~0.8 ms at p99 (6.7 → 7.5 ms, worst observed 24.0 ms). At small corpus size the effect is below the run-to-run noise: across four runs the p99 delta was +1.4, +0.1, −4.2, −1.1 ms — the sign is not stable, and in two runs the ticking window was faster than the quiescent one. On a 100 ms-blocked-loop criterion the shortlist becomes worth building at roughly 1100–1250 sessions. One CostView load blocks the loop about 100× longer than one tick.

One incidental result confirms M-17 named the right cost centre: at that ~11x scale an incremental tick (135.5 ms, one session grew by one record) is cheaper than a warm tick (148.0 ms), so accepting the changed session vanishes into the noise and essentially the whole tick is the sweep.

When M-17 does come up, two risk-free levers come before a shortlist: the poll interval is a PROVISIONAL constant (3 s → 10 s drops the duty cycle from ~5% to ~1.5% for a one-constant change — a product decision about liveness latency, so it is Ivan’s to make, not an agent’s), and the stat sweep can move to a worker thread (lstat/readdir need no better-sqlite3), which removes the stall without touching discovery correctness.

The per-directory mtime cut is unsound and must not be built. A directory’s mtime changes when an entry is created, deleted or renamed — not when an existing file inside it is appended to. Appending to <uuid>/subagents/<child>.jsonl moves that file’s mtime and leaves subagents/ untouched. That is exactly the live case this product exists to show: the subagent writes while the parent transcript sits still. The cut would silently miss live subagent growth — the worst possible failure mode for a dashboard whose claim is that the subagent tree is a data fact. The saving it buys is 148–156 ms per poll. If a shortlist is ever built it must be fs.watch as a hint layer over a periodic full sweep (Node’s recursive watch drops events silently under load, especially on macOS and network filesystems), never a replacement, and never a directory mtime cut.

Honest limits of these numbers, stated by the run itself: the OS page cache was warm (both corpora had just been written), so the sweep figures are an underestimate; there is no seam separating the sweep’s cost from the rest of a tick, so warm tick is a proxy; app.inject cannot measure a request that arrives mid-stall, so for that case the honest number is the event-loop max (156–178 ms), not the request p99; and the fixtures pin one subagent-files-per-session shape, so a corpus with deeper subagent trees costs more at the same session count.

M-16 is closed but incomplete: the boot health probe cannot answer

The cold-replay row above was checked against the server rather than assumed to be a property of the benchmark harness, because a 341.7–370.1 s figure only matters if the real boot path has the same shape. It does:

That yield does what it claims and no more: it drains what was already accepted at the instant of bind. A probe that connects one second into the replay is accepted by the kernel backlog and then waits — for the better part of six minutes at what this review called real corpus scale (in fact ~11x it), and even in the smaller run in the same series for 22.5 s. At census scale the same wait is 34.87-39.92 s (measured 2026-09-01, over a synthetic corpus 3.7x lighter per session than the real one — a lower bound; see BENCH-SHAPE). The behaviour is unchanged by the correction; only its duration is. So the observable boot behaviour is a socket that accepts and never answers, which for a health probe is worse than connection-refused: refused fails fast and is unambiguous, accepted-and-silent hangs until the client’s own timeout and is indistinguishable from a wedged server.

The ingestPhase state (:553, :676) is correct and the endpoint is wired; the phase simply is not reachable during almost all of the window it exists to describe. Whatever closes this is a change to how the replay runs — chunking it across macrotasks, or moving it off the main loop — not a change to the health endpoint, which is already right. Nothing here is a regression: M-16 improved on replay-before-listen, and this is the next layer of the same problem, now measured.

The validation corpus is a rolling window, so corpus-measured figures expire

Chased down while checking an unexplained discrepancy in a lane report, and it turns out to matter well beyond the number that surfaced it.

docs/analysis/parser-spec.md cites 141 sessions (20 slugs, 54 with subagents, 1855 subagent transcripts) in eight places, including its headline validation claims — 1855/1855 edge reconstruction, “parses all 141 sessions end-to-end, 0 UsageConflictError, 0 SubstrateError”, and the 2-of-141 → 0-of-141 pricing-settle result. Measured on this machine on 2026-08-22:

  parser-spec today
slugs 20 21
main transcripts 141 51
all .jsonl incl. subagents ~1996 2551

The oldest main transcript in ~/.claude/projects today is dated 2026-07-24 — about four weeks back. Main transcripts fell while subagent transcripts rose, which is what a rolling retention window over an increasingly subagent-heavy workload looks like. Whether the older ones aged out or were cleaned up by hand is not something this repo can determine, and it does not change either consequence:

  1. Every corpus-measured figure in the doc corpus is perishable and must carry its measurement date. The parser-spec numbers above cannot be reproduced today — not because anything regressed, but because the corpus they describe is gone. They are historical results, and re-running the same validation now answers a different question. No figure quoted from ~/.claude/projects should ever again be written without its date and its slug/session counts beside it.
  2. The JSONL corpus is not an archive, so it is not a universal recovery path. The design rests on JSONL being ground truth (CD-1), and rebuilding the DB from JSONL is the obvious answer to almost any ingest bug. That answer only reaches back as far as the window. Past its edge, the dashboard’s own SQLite is the only record of that spend — which raises the stakes on backups and makes the retention values (OPEN-1/2/3) a question about what history is permanently lost, not merely about disk use.

Nothing here is a code defect and nothing needs fixing today. It is a standing correction to how this project’s measured claims should be read and written.


Strengths

Merged and deduplicated across the ten dimensions. Every claim below was cited against code by the dimension reviews.

1. Security invariants are enforced in code, not by convention

2. Cost integrity: “Token counts are ground truth … never inferred” holds end-to-end

3. Parser: 13 of the 14 normative gate items conform, with structural (not textual) joins

The heading’s “13 of the 14” was true on the review date and is kept for the record; the fourteenth (gate #7) landed afterwards — see M-1 in the disposition table. The count that replaced it is not “14 of 14” either: parser-spec §3 now separates implemented (14) from exercised by the real corpus (11), and gate #7 is one of the two shapes that exist only against fixtures.

4. Ingest is fail-safe in the correct direction: extra work, never wrong data

5. Data layer: invariants live in the storage engine

6. API/HTTP honesty and measured performance work

7. Frontend renders uncertainty instead of coercing it

8. Architecture and docs discipline

9. Test quality: mutations killed on real infrastructure


Weaknesses

Grouped by severity; cross-dimension duplicates merged (merge noted in the entry). Format: title — dimension(s) · verdict · location; failure scenario; suggestion. (known, parked for owner) marks items already on an owner list.

High

H-1. Migration 7 was edited in place after being applied; no checksum, no repair migration — data-layer · CONFIRMED · apps/server/src/db/migrations.ts:40 Migration id 7 changed between commits (effective_from ‘2026-07-11’ → ‘2026-01-01’ and all four real model keys bare → claude--prefixed), but runMigrations skips by recorded id only. The live DB at data/agenthropic.db has ids 1–7 applied with the old seed: under HEAD every real-model message throws PricingError, fails the halt gate, and is quarantined — whole-corpus ingest sinks on that DB while a fresh DB works, with nothing detecting the divergence. Fix: corrective migration id 11 rewriting the seed rows, plus a content checksum recorded in schema_version so an edited applied migration fails loudly.

H-2. Per-file skip diagnostics are dead in production; an oversize/EACCES main transcript silently freezes a session’s dollars — ingest-watcher + performance-roadmap (merged) · PLAUSIBLE · apps/server/src/index.ts:356 The production watcher wiring never passes onWarning (the per-file skip sink), and logReplaySummary omits filesSkipped. A main transcript crossing the 64 MiB cap (PROVISIONAL value, itself parked) or chmod’d to EACCES is skipped; the session ingests partially from subagent files, counts as ok, is checkpointed — its persisted token totals and dollars freeze while the JSONL ground truth grows, with zero signal in logs, SSE, health or UI. The entire NoSubstrate/SkippedFile design is unreachable in the composed server. Fix: wire onWarning to a rate-limited structured log and/or SSE diagnostic, add filesSkipped to replay/tick log lines and a cumulative skip counter to /api/health; consider a per-session partial-substrate flag rendered like unpricedTokens.

Medium — CONFIRMED

M-1. Parser gate #7 (legacy 2.1.70 bare-Explore fallback) is the one unimplemented gate item — parser-core · packages/core/src/parser/parse-session.ts:327 A pre-2.1.71 transcript’s legacy spawn shape never enters toolUseOwner, so a flat child with no other anchor silently orphans — quiet edge loss in the moat DAG when ingesting an older machine’s corpus (the product’s core use case); meanwhile TODO.md:325 claims the 14-item gate is satisfied and parser-spec §3 records no waiver. Fix: add the defensive fallback with distinct provenance, or record an owner-signed waiver in the spec’s §3 table so 14/14 stops overstating.

M-2. Ingest prices against a boot-time pricing snapshot; seeding a new model row does not unblock ingest — cost-engine · apps/server/src/index.ts:358 loadPricing(db) is captured once at start; the operator follows the PricingError’s own advice, seeds the row, and the watcher still burns its 3 attempts on the stale snapshot and parks the session — while the cost-analysis route (fresh loadPricing per request) works, a confusing split-brain. The watcher’s own comment says the retry budget exists for “a pricing row that arrives”, contradicting the snapshot. Fix: reload pricing per tick (µs against a 3 s poll) and reset attempt counters when the pricing table changes.

M-3. Retention residue semantics inverted now that replay checkpoints are wired — data-layer (merges the ingest-watcher stale-doc finding) · apps/server/src/retention/prune.ts:41 The documented “restart replay puts pruned rows back” is false: checkpoints are honored while the sessions row exists (retention never deletes sessions), so idle sessions’ pruned token_usage stays gone; a later append resurrects it, and a subsequent prune journals the same dollars again — reconciliation by summing receipts over-counts. Opt-in policy (defaults NO_RETENTION), so harm is wrong operator documentation, not wrong live dollars. Fix: at OPEN-1 ratification decide the interplay explicitly — invalidate affected checkpoints inside the prune transaction, or exclude pruned windows on re-ingest — and fix the prune.ts comment + journal note either way.

M-4. restoreDatabase can replay a stale WAL from the pre-restore database into the restored file — data-layer · apps/server/src/db/backup.ts:25 copyFileSync then openDatabase with nothing removing a pre-existing destPath-wal/-shm: after an unclean shutdown plus in-place restore, SQLite recovers the OLD database’s WAL frames into the restored copy — a refused legitimate restore or a silent mix of two states, in exactly the disaster path backups exist for. Fix: delete (or refuse on) ${destPath}-wal/-shm before the copy; document that in-place restore requires the server stopped.

M-5. getGlobalDag prices and groups the entire token_usage table on every request — data-layer · apps/server/src/api/queries.ts:563 usage_by_agent builds from the unfiltered priced CTE regardless of nodeLimit (measured 432 ms over 752k rows; ~5 s projected at what was then called real corpus scale — the inflated 1855-as-sessions target, so this projection is now unverified), and the edge query full-scans orchestration_edges (no parent/child index). Page cost grows with corpus size, not response size. Fix: restrict usage_by_agent to the selected agent set (id-list injection, as sessionSummarySelect already does); add edge indexes on (parent_agent_id) / (child_agent_id).

M-6. ingest-failed SSE frames are published but no client ever listens — the dashboard stays silent on quarantine — api-realtime + tests-quality (merged: the event-type list is hand-duplicated per package and has already drifted) · apps/server/src/realtime/bridge.ts:52, apps/web/src/sse.ts:65 EventSource drops named events with no registered listener; SERVER_EVENT_TYPES lists only two of the three emitted types, so a quarantined session produces nothing in the UI — the exact silent-failure mode WP-IN5 claims fixed (“reaches BOTH the operator … and the dashboard”), and api.md’s “Two typed frames only” is stale. Fix: move the event-type list into packages/shared, import it on both sides, add a contract test that every published type is a member, subscribe and render ingest-failed (e.g. LiveView banner); update api.md.

M-7. Hook receiver escapes the uniform error contract: 5xx leaks raw error.message via Fastify’s default handler — api-realtime · apps/server/src/hooks/routes.ts:115 registerHookRoutes sits on the root app, outside the apiRoutes plugin’s scoped setErrorHandler; a throwing append/applyStatus (SQLITE_BUSY/FULL, I/O) returns the raw SQLite message and a shape matching neither ApiErrorSchema nor the declared 202 — empirically reproduced. Bounded by loopback+auth. Fix: root-scope setErrorHandler with the uniform { error } shape and 5xx message suppression; declare 400/500 responses on the route.

M-8. Q2 (biggest agent/subagent burner) is answerable only by hover-hunting — web-frontend · apps/web/src/views/CostView.tsx:265 No view lists agents by token count — per-agent tokens exist only in SVG <title> hover tooltips, unreachable for keyboard/AT users; the ux0 TOP BURNERS leaderboard is absent and TODO.md’s “all 5 questions answerable” is overstated for Q2. A material KC-4 exit-gate gap. Fix: per-agent top-burners table (per-agent usage is already persisted; even a client-side sort of served DAG nodes would be honest).

M-9. Q4 today/this-week KPIs and any aggregate delegation-saved figure are absent; only top-5 sessions are analysable — web-frontend · apps/web/src/views/CostView.tsx:91 KPIs are all-time only; delegation savings is reachable solely by clicking one of the top-5-by-cost sessions — no route or link makes any other session analysable — so half of Q4 has no on-screen answer. Same overstated-gate pattern as M-8; RELEASE.md’s unchecked [HUMAN] daily-questions box would catch it pre-tag. Fix: today/this-week KPIs from the existing perDay rows; link SessionsView rows into cost analysis; an aggregate savings figure needs a server endpoint (cross-scope — flag it).

M-10. Relative-time labels freeze: “just now” can persist for hours — web-frontend · apps/web/src/views/LiveView.tsx:131 Date.now() is captured per render and no timer exists anywhere in the web app; on a quiet stream (heartbeats are comment frames that never reach EventSource) the recency view presents a stale claim as current until the next event or navigation. Fix: a 30–60 s interval bumping a clock state; formatRelativeTime already takes injected nowMs.

Medium — PLAUSIBLE

M-11. Installed hook command exposes the dashboard token in the process table — security · hooks/install.mjs:119 The generated shell command expands ${DASHBOARD_TOKEN} into curl’s argv; any other OS account on the machine can harvest it from ps during the up-to-3 s curl window — exactly the “local multi-user” attacker the threat model claims the token defends against. Fix: pass headers via a 0600 file (curl --config/--header @file) or a wrapper reading the env var itself; at minimum document the residual exposure.

M-12. Cross-session message.id collision (resume/fork) would silently rewrite agent attribution across sessions — cost-engine · verified in code, trigger unproven on this machine’s corpus · apps/server/src/db/token-usage.ts:69 UNIQUE(message_id, bucket) is global and the upsert rewrites agent_id without touching session_id; if a CLI resume ever replays history into a new file (the behavior that forces ccusage-class tools to dedupe globally), the same agent shows different dollars in /api/dag/global vs the session tree and attribution flip-flops with ingest order. A full scan of this machine’s corpus found zero colliding ids today. Fix: decide the ownership rule explicitly (per-session uniqueness + parse-time dedupe of replayed history, or refuse cross-session agent_id rewrites and count the collision); add a resumed-session fixture either way.

M-13. SubagentStop verdict lost for subagents shorter than one poll interval — ingest-watcher · apps/server/src/db/agents.ts:150 A hook arriving before the agent row exists is stored as raw liveness and never replayed; every fast subagent (< 3 s, common for quick explores) ends unknown instead of completed, degrading terminal-state accuracy for the most numerous agent class. Fix: on ingest of a newly inserted agent id, reconcile pending SubagentStop rows from events_raw through applyAgentStatus (sticky-terminal CASE makes this idempotent).

M-14. Duplicate session uuid across two slug dirs causes perpetual re-ingest and project_slug flapping — ingest-watcher · apps/server/src/ingest/corpus-watcher.ts:348 The fingerprint map keys on sessionId alone (last ref wins) with no duplicate-stem guard at enumeration; a copied project dir yields a permanent 3-second full re-read loop for that session plus UI-visible flapping project attribution. Fix: detect duplicates during enumeration, keep one ref, record the rest as skipped with a duplicate-session reason through the (to-be-wired) onWarning channel.

M-15. Tail-follow never tails: any change triggers a synchronous full re-read + re-parse on the event loop — ingest-watcher + performance-roadmap (merged) · apps/server/src/ingest/corpus-watcher.ts:333 An active session with a tens-of-MiB transcript costs a full re-read/re-parse every 3 s tick — per-tick cost is O(session size), not O(delta); cumulative I/O is quadratic in session size, and each tick blocks every API request and SSE heartbeat. This degrades the dashboard exactly while the user watches a live session — its core use case. Highest-leverage pre-KC-2 performance fix. Fix: use the fingerprinted size as a byte offset and read only the appended tail (full re-read on shrink/identity change), and/or parse off-thread; add a tick-duration metric to health.

M-16. Startup replay blocks before listen — minutes of no server at real corpus scale — performance-roadmap · apps/server/src/index.ts:407 watcher.tick() runs before app.listen; a first boot (or any checkpoint-degrade) against the real corpus reads and parses everything with no HTTP surface, no health endpoint and no progress output — a health-checked supervisor may kill it into a loop. (Unit corrected 2026-09-01: this sentence said “the real ~1855-session corpus”. 1855 counts subagent transcripts; the census of record is 141 sessions — parser-spec §4.2. The finding does not depend on the number: cold replay at census scale is measured at 34.87–39.92 s of one synchronous stall, and “minutes” in this heading was a projection at the inflated scale, now unverified.) Fix: listen first with a “replaying” status (replay is idempotent, partial visibility is safe), emit periodic progress, or chunk replay across ticks.

M-17. Steady-state fingerprint pass lstats every file of every session every 3 seconds — performance-roadmap · apps/server/src/corpus/fingerprint.ts:86 No shortlist: at 10x corpus scale ~60–90k synchronous lstats per tick, plausibly 0.3–1 s of blocked event loop with zero changes — permanent API/SSE jitter. (Numbers PROVISIONAL, from the bench’s floor-labeled projection.) Fix: shortlist via fs.watch or a project-dir mtime cut; revisit the 3 s PROVISIONAL interval with warm-tick numbers before v1.0.

M-18. Cost-analysis endpoint re-enumerates the entire corpus and re-parses the full session per request, synchronously — performance-roadmap · apps/server/src/api/substrate-provider.ts:106 Every loadSession readdir+lstats every project dir to find one ref, then re-reads up to 64 MiB — several hundred ms to ~1 s of frozen server per click. Fix: resolve the ref directly from the DB’s project_slug (keeping the same containment vetting) or cache enumeration briefly; share the tail/worker fix from M-15.

M-19. getCostSummary scans all of token_usage on every request; no cache or materialization — performance-roadmap · apps/server/src/api/queries.ts:473 The one-scan rewrite fixed the 4x-scan shape, but each CostView load still prices every row on the event loop; linear with corpus age forever (retention of token_usage deliberately refused). Fix: maintain the (session, model, day) rollup incrementally at ingest, or cache the summary invalidated on session-ingested.

M-20. Backups exist as capability + manual drill, but nothing in production ever runs one — and events_raw is the only non-re-derivable table — performance-roadmap + data-layer (merged) · apps/server/src/db/backup.ts:14 backupDatabase is test-proven and never scheduled; a disk failure loses all hook liveness history permanently. The invariant reads “SQLite in WAL mode with backups” — the letter (WAL + tested restore) is met; the spirit (backups actually happening) is not. Fix: schedule backupDatabase in start() (e.g. daily) into data/backups/ with the retention module’s keepMinimum-floored expiry; log each run. Small, no owner decision needed.

M-21. Dated-price resolution is implemented twice (core TS epoch-ms vs SQL lexicographic) with no reconciliation test — architecture-dx + cost-engine (merged; the lexicographic-timestamp sub-aspect is (known, parked for owner)) · apps/server/src/api/queries.ts:36, compute-cost.ts:70-80 A pricing edge (e.g. 'Z'-suffixed effective_from vs millisecond occurred_at — ordered differently as strings, identically as ms) lets the API present dollars the ingest halt gate never approved, undetected. Fix: one property-style test asserting the two resolvers agree to the cent across boundary timestamps; normalize effective_from at write time; longer-term make SQL the single resolver.

M-22. CI never builds the web app and there is no production run path for the SPA — architecture-dx · .github/workflows/ci.yml A change that breaks the Vite production build merges green; the first vite build may happen at v1.0 release time under the KC-4 hard date. The server never serves dist/; the supported deployment topology is undeclared. Fix: add pnpm --filter @agenthropic/web build to CI; decide and document the production topology (server serves dist/, or two-process is declared supported in RELEASE.md).

M-23. The zero-pragma coverage guard exists only in apps/server and apps/web — packages/core, shared and test-fixtures are unguarded — tests-quality · packages/core/vitest.config.ts:16 A /* v8 ignore */ above the unknown-model branch in core would keep 100% green while untesting the code that owns the no-silent-$0 invariant — the exact failure mode the server guard’s own header documents. Fix: extract the sweep into a shared helper and instantiate the guard test (incl. threshold-pin and no-exclude assertions) in all three packages.

M-24. Hierarchy-accuracy exit gate is unmeasured: the human annotation corpus ships empty — tests-quality · (known, parked for owner) · packages/core/test/hierarchy-gate.test.ts:244 All parser ground truth is machine-authored (6 synthetic fixtures); the ≥95% claim has never been measured against reality — the tooling honestly prints NOT CERTIFIED, but honesty does not close the gap. Fix: hand-label the minimum 52 claims from the local spike corpus before the Phase-3 exit; treat NOT MEASURED as release-blocking.

M-25. v1.0 critical path is pinched by owner-only acts: KC-3’s merge-blocking clause is unmeetable by agents — performance-roadmap · (known, parked for owner) · TODO.md:109 KC-3 (2026-10-12) requires the three P0 proofs to be merge-blocking, but main is not branch-protected — “a passing test that nothing gates on is a test, not a gate” (RELEASE.md); LABEL-ME is unstarted, the <30 s measurement deferred. The stay-alive condition can fail on a 10-minute Settings click. Fix: put the two 10-minute owner acts (branch-protection click, LABEL-ME start) in front of Ivan now with the KC-3 date attached; run the <30 s stopwatch measurement in September, not November. Resolved 2026-08-25 — branch-protection half only. main is now branch-protected: the required status check is ci (the job id in .github/workflows/ci.yml, not the workflow’s CI display name), and force-pushes to main and deletion of main are refused. The three P0 proofs are therefore merge-blocking for a contributor and, by deliberate design, not for the repository owner — enforce_admins is off because agenthropic has one maintainer whose normal working mode is a direct push to main, and admin enforcement would lock the sole maintainer out of their own repository. KC-3’s stay-alive clause no longer fails on this act. LABEL-ME (M-24) and the <30 s measurement are untouched.

Low (all UNVERIFIED-LOW)


Prioritized improvement plan

Bucket 1 — quick wins (each under a day)

# What Why now Size
1 Migration 11: rewrite model_pricing seed (keys + effective_from) and add a migration content checksum to schema_version (H-1) CONFIRMED; the operator’s live DB is sunk under HEAD today S–M
2 Wire onWarning → structured log/SSE; filesSkipped in replay/tick summaries; skip counter on /api/health (H-2, enables M-14’s reporting) Restores the entire skip-diagnostics design the composed server currently discards S
3 Reload pricing per watcher tick + reset attempts on pricing-table change (M-2) CONFIRMED split-brain; the fix is microseconds per tick S
4 Shared SSE event-type list + contract test + subscribe/render ingest-failed; update api.md (M-6) CONFIRMED silent-failure mode; drift already shipped once S
5 Root-scope error handler for hook routes (M-7) CONFIRMED contract escape; a few lines S
6 restoreDatabase: remove/refuse stale -wal/-shm before copy (M-4) CONFIRMED disaster-path defect S
7 Restrict getGlobalDag usage rollup to selected agents + edge indexes (M-5) CONFIRMED; measured 432 ms → the only unfixed endpoint from that bench run S
8 Schedule daily backupDatabase in start() (M-20) Makes “SQLite in WAL mode with backups” an operating fact; no owner decision needed S
9 LiveView clock interval (M-10); Retry buttons + refetch-on-ingest (L-16); health-chip re-probe (L-18) CONFIRMED honesty defect on the recency view + cheap UX polish S
10 SQL-vs-TS rate-resolver reconciliation test (M-21) Guards the two authoritative dollar paths against divergence S
11 CI web build step (M-22, CI half) Catches a whole defect class before the KC-4 hard date S
12 Coverage-pragma guard in core/shared/test-fixtures (M-23) Protects the no-silent-$0 code’s coverage denominator S
13 Doc/consistency touches: prune.ts residue note (M-3 doc half), retention-queries header (L-11), PROJECT-STATE counts (L-23), README quickstart (L-20), drop d3 (L-17), meta-classification tightening (L-5), sessionId cross-check (L-4) All small, all reduce the honest-docs debt this project trades on S

Bucket 2 — pre-v1.0, sequenced against KC-2 (2026-09-14) and the 2026-12-01 hard date

Before KC-2 (performance and measurement, so numbers exist with runway):

# What Why now Size
1 Byte-offset tail-read for changed sessions (+ optional worker-thread parse); tick-duration metric on health (M-15) Highest-leverage perf fix; the live-watch path is the core use case M–L
2 Listen-before-replay with a “replaying” status and progress output (M-16) Removes the boot blackout — measured at 34.87-39.92 s of one synchronous stall at census scale (2026-09-01). The “minutes-long” that justified this row was a projection at ~11x real scale and is now unverified; the fix may well still be worth its slot at 35-40 s, but that argument has to be made rather than inherited. The 35-40 s is itself a lower bound: it was measured over a synthetic corpus 3.7x lighter per session than the real one (BENCH-SHAPE) M
3 Direct-ref lookup for cost-analysis (skip corpus enumeration) (M-18) Removes a per-click whole-server freeze; pairs with #1 S–M
4 Incremental (session, model, day) rollup or cached summary for getCostSummary (M-19) The read grows with corpus age forever otherwise M
5 Bench: event-loop-delay + concurrent-inject contention phases (L-26); then run the <30 s stopwatch measurement in September (M-25c) Findings need two months of fix runway before 2026-12-01 S–M
6 Fingerprint shortlist (fs.watch or mtime cut) if the contention numbers demand it (M-17) Decide from measurement, not speculation; 3 s interval is PROVISIONAL M

Owner-only acts — put in front of Ivan immediately (KC-3 is 2026-10-12):

# What Why now Size
7 Branch-protection click making the three P0 proofs merge-blocking (M-25a) KC-3’s stay-alive clause fails on this alone; agents cannot do it 10 min (owner)
8 Start LABEL-ME: hand-label ≥52 claims from the spike corpus (M-24, M-25b) The ≥95% hierarchy bar is unsignable and every number stays PROVISIONAL until this exists M (owner)

Item 7 was carried out on 2026-08-25: main is branch-protected with ci as the required status check, and force-pushes and branch deletion are refused, so the three P0 proofs withhold the merge button from a contributor — and, deliberately, not from the repository owner, enforce_admins being off because agenthropic has one maintainer whose normal working mode is a direct push to main. Item 8 has not been carried out: the human annotation corpus is still empty, so the ≥95% bar still reports NOT CERTIFIED and every hierarchy number stays PROVISIONAL.

Product/correctness before the KC-4 exit gate:

# What Why now Size
9 Q2 top-burners table + Q4 today/this-week KPIs + cost-analysis entry from SessionsView; aggregate-savings endpoint (M-8, M-9) Two of the five daily questions are currently unanswerable as specified; the RELEASE.md [HUMAN] box will fail M
10 Gate #7: defensive legacy fallback or owner-signed waiver in parser-spec §3 (M-1) The 14/14 claim currently overstates; old-corpus ingest silently loses edges S–M
11 Cross-session message.id ownership rule + resumed-session fixture (M-12) Settle the rule while it is cheap; the trigger is external CLI behavior that may return M
12 SubagentStop reconciliation on agent-row creation (M-13); duplicate-session detection at enumeration (M-14) Systematic status inaccuracy for the most numerous agent class; a footgun loop S–M each
13 Hook token out of argv (curl config file) or an honest threat-model amendment (M-11) The threat model currently claims a mitigation the process table undermines S
14 Retention/checkpoint interplay decision at OPEN-1 ratification (M-3) The journal’s reconciliation semantics are wrong in both directions until decided S (decision) + S (code)
15 Production topology decision for the SPA: server serves dist/ or two-process declared supported (M-22, topology half) RELEASE.md cannot honestly describe a deployable artifact without it S–M

Bucket 3 — later / v2 territory (alerts land only via KC-5)


Review coverage

Dimensions run (10): security, parser-core, cost-engine, ingest-watcher, data-layer, api-realtime, web-frontend, architecture-dx, tests-quality, performance-roadmap.

Findings kept per dimension: security 4 · parser-core 4 · cost-engine 5 · ingest-watcher 7 · data-layer 7 · api-realtime 5 · web-frontend 7 · architecture-dx 7 · tests-quality 5 · performance-roadmap 10 — 61 raw, merged to 53 entries in this report (8 cross-dimension duplicates folded, noted inline). After dedup: 2 high, 25 medium, 26 low.

Verification stats: the adversarial verification pass was budget-capped at 14 findings. Of those 14: 11 CONFIRMED, 1 remained PLAUSIBLE after inspection (M-12 — mechanism fully real in code; the triggering substrate behavior was absent from this machine’s entire corpus, so code alone cannot settle it), and 2 were REFUTED and removed before synthesis. 18 further findings were queued past the cap and ship as PLAUSIBLE — credible from code reading, not adversarially verified. All low-severity findings are UNVERIFIED-LOW by policy.

What this review did NOT cover: