agenthropic

Concept Analysis v2 — Consolidated

Authoritative re-analysis of the agenthropic conceptual brief. This v2 folds four inputs into one decision-useful view:

  1. the original internal analysis — concept-analysis.md + implementation-plan.md (v1);
  2. two externally-produced parallel reports — BASE (~3.7k words) and EXPANDED (~7.7k words);
  3. the adversarial cross-check of those two — external-docs-review.md;
  4. the design of record — docs/ai/DESIGN.md and the invariants in CLAUDE.md.

It was produced by a six-lens senior workflow (Architect · Developer · QA · Business Analyst · brutal Gap · Holistic), Opus / high effort, ~468k subagent tokens, run against all of the above. Where v1 and the externals agree, that is recorded as confirmed; where they diverge, this document decides. The consolidated decisions here are the input to development-plan.md, which decomposes them into agent-distributable work packages.

Naming: decisions are tagged by originating lens — AD* Architect, SD* Developer, QA-D* QA, BA-D* Business, G-D* Gap, LB*/H-* Holistic. The canonical set in §3 dedupes them into ten load-bearing decisions the build actually turns on.

Empirical update — 2026-07-04 desktop probe. The full-corpus read-only probe (phase0-probe.md, 17 projects · 117 sessions · 33 subagent dirs) pre-answers CD-1 as CONDITIONAL-GO → build (confidence 85) and corrects the mechanism assumptions below. It de-risks but does not replace the formal Phase-0 spike — WP-S1/WP-S5 still need the paired-capture corpus + Ivan’s tree sign-off, and the WP-S7 GO gate still stands (no production code before it).


1. Bottom line

Build. The concept is unusually coherent because it descends from a source-level audit of six real rivals, not a blank page. One dimension — the security posture (loopback-only, mandatory-token-or-fail-startup, no-spawner, same-origin, no-SSRF) — is genuinely best-in-class and is the system’s spine, not a bolt-on. The external docs converge with v1 on the two decision-critical tripwires (neither repeats the refuted “simple10 has no DAG”; both diagnose hoangsonww’s RCE as a bypassPermissions spawner, not a concurrency cap), which raises confidence in the whole rather than adding new content.

The gap between this A-grade vision and a shippable v1 is a small number of load-bearing decisions that are cheap on paper and ruinous in code. Two of them govern everything else:

Everything else resolves cleanly around these two.

Synthesis rule for v2: BASE’s sequencing + accuracy (Phase-0 spike, security in Phase 1) + EXPANDED’s formal apparatus (FR/NFR/ADR/traceability, negative-test catalogue, quantified metrics, the events_raw+events schema split) + the eleven items the internal analysis carries that both externals miss. EXPANDED’s own generation defects — seven byte-identical §10 holistic tables, security deferred to Phase 6, Phase-0 reduced to paperwork, SSRF dropped — are explicitly quarantined and not inherited.


2. The two load-bearing decisions

Everything hangs off these; if either is wrong the rest is wasted motion.

LB1 — Ingest primacy (the data-foundation seam)

Decision: ingest is JSONL-primary + replay-on-startup, with hooks providing sub-second liveness only, and every write is an idempotent upsert on a stable event id — contingent on the Phase-0 spike proving JSONL carries the subagent parent→child linkage. If it does not, fall back to hooks-primary + a durable local outbox/spool (at-least-once, idempotent upsert), and only then.

Why it is load-bearing: this single choice determines whether the persistent-DAG moat is trustworthy (survives an outage), whether the ground-truth-tokens invariant is satisfied naturally (tokens read from the durable log, never inferred), whether history is crash-tolerant, and whether >90% coverage is even achievable (replay-from-fixtures needs a deterministic durable source). The primacy decision is an output of Phase-0, not an assumption baked in ahead of it — which is precisely EXPANDED’s error. (The 2026-07-04 desktop probe has since pre-answered it CONDITIONAL-GO (confidence 85); the formal paired-capture spike confirms rather than decides it — see the empirical-update note above and phase0-probe.md.)

LB2 — Identity: personal-first / commercial-clean

Decision: build the single-user Mac Mini cockpit; take only the cheap commercial hedges now (MIT-clean code only; instance/host_id on every row from the first migration; a schema that does not block tenancy); explicitly defer fleet and multi-tenancy. Resolve the “OPCⁿ” commercial-line token — define it or drop it — before it drives tenancy/schema/license-strictness investment.

Why it is load-bearing: it resolves two cross-dimension tensions at once — scope discipline for a solo owner and legal reality. The constraints happen to align: the copyable repos (simple10, hoangsonww) carry the large patterns (tree-building, webhook/alert schema), while the uncopyable ones (cast/disler/nirdiamant) carry only small ideas cheap to reimplement clean-room.


3. Canonical decision register (v2)

The ten decisions the build turns on, each consolidating the per-lens decisions that agree on it. These are the durable output of this analysis.

# Decision Consolidates Satisfies (FR/NFR — recovered EXPANDED §3, LOST-5) The rule
CD-1 Ingest primacy is JSONL-primary + replay-on-startup, contingent on Phase-0; else hooks-primary + durable outbox. LB1, AD3, SD3, G-D1 FR-01, FR-05; NFR-OPS-01 Decided by the Phase-0 diff (tree-from-JSONL vs tree-from-hooks), never assumed.
CD-2 Single immutable substrate + deterministic projection. Both sources write into append-only, idempotency-keyed events_raw; sessions/agents/orchestration_edges/token_usage are a pure replayable projection over it. AD1, SD2 FR-04; NFR-DATA-01/02 Reconciliation is per-field precedence at projection time, not a two-store merge at query time.
CD-3 Reconciliation precedence. Tokens are JSONL-authoritative (never inferred); interim liveness/state from hooks; final session/agent state + cost from JSONL. token_usage.agent_id is nullable at first write, deterministically backfilled once the agent is known. AD2, SD4 FR-03, FR-05; NFR-DATA-02 Cross-source idempotent upsert: a fact seen by both a hook and JSONL lands once.
CD-4 Schema: events_raw(immutable) + events(normalized); persisted orchestration_edges (self-ref parent_agent_id, instance/host_id, derived_from_event_id, idempotent); fine-grained token_usage (service_tier / speed / inference_geo + compaction baseline); versioned model_pricing (effective_from, verified_on). AD4, G-D6, SD5 FR-02/03/04/06/08; NFR-MAINT-01 events_raw is provably append-only (no UPDATE/DELETE path, enforced by test).
CD-5 Transport is SSE with same-origin enforcement from Phase 1. AD5 FR-06, FR-07; NFR-SEC-02 Server→browser-only feed; revisit WebSocket only if bidirectional control is ever needed (it is not).
CD-6 Ports & adapters: named set — HookSource, TokenReader/TokenSource, StoragePort, RealtimeHub, AlertSink, PricingProvider, CostEngine + an EventStore port + a pure Normalizer/Projection; simple10’s strategy-pattern agent classes are the per-runtime adapter (Claude Code now, Codex later). AD6, SD1 FR-01/05/08/09 (via named ports); NFR-MAINT-01 Keeps the normalizer testable/replayable and the system multi-runtime-ready without a core rewrite.
CD-7 Security + the coverage gate are boundary conditions from commit one, CI-blocking: loopback-or-fail bind; mandatory DASHBOARD_TOKEN-or-fail-startup (timing-safe compare); SSE same-origin; no-spawner grep/static gate; no-SSRF (webhook targets operator-configured, never dialed from a payload); WAL + tested restore; >90% coverage blocks merges. AD7, SD8, QA-D3/D4, G-D3, H-SEQ NFR-SEC-01/02/03, NFR-MAINT-01 Rejects EXPANDED’s security→Phase 6 / backup→Phase 8. Slice 8 is polish only.
CD-8 Phase 0 is a throwaway GO/NO-GO feasibility spike with a hard ❌ stop — G0.1 ingest-primacy probe · G0.2 hook-catalog enumeration (don’t assume “the twelve”; confirm/deny SubagentStart) · G0.3 tree smoke gate · G0.4 token-reconciliation probe. No production code until green. AD-Phase0, G-D2, H-SEQ — (process gate; de-risks FR-01/03/05/06) Rejects EXPANDED’s paperwork “decision lock” that validates linkage only at the Phase-4 UI.
CD-9 Per-artifact licensing. COPY simple10 tree/ports + hoangsonww Telegram/webhook schema + dual-sqlite driver (dropped — single better-sqlite3 driver per best-path §6.3, applied 2026-07-06) with attribution; CLEAN-ROOM reimplement cast controlGate + delegation-savings and nirdiamant checkpoint (never view their source while writing). Enforced by a CI provenance/license scan. SD6, BA-D4, G-D4, LB2 — (licensing/provenance; underpins FR-09) cast/disler/nirdiamant are all-rights-reserved by Berne default — not “ambiguous”.
CD-10 Scope + secrets + retention. MVP = the 5 daily questions; Phase-3 vector-DB “observability-becomes-memory” feed on a labeled experimental track; fleet deferred until a second host exists; Telegram token via token_ref → launchd env / chmod-600 (never in SQLite, never to the browser); retention TTL + payload redaction from Phase 1. BA-D1/D3, G-D7, AD8, SD7, LB2 FR-09; NFR-PRIV-01 ANTHROPIC_API_KEY stays out of the dashboard env entirely.

Cost-trust chain (CD-3 + CD-4 crosscut, H-COST): every displayed dollar traces to (ground-truth tokens × a dated, priced model); a model observed in fixtures with no price row FAILS CI — silent staleness is a red build, not a runtime “estimated” label. This extends the byte-exact-tokens guarantee to the priced output.

How the register stands as built (note added 2026-08-15). The ten decisions were implemented, not revised: the register is still the right description of the system, and the rows are left exactly as they were resolved. Four details have to be corrected against the code, and one row records an event rather than a design.

CD-4 names a verified_on column that was never built. The shipped model_pricing table is (model, bucket, usd_per_mtok, effective_from) keyed on PRIMARY KEY (model, bucket, effective_from); there is no provenance column, and effective_from alone carries the versioning. Whether provenance is worth adding is the surviving half of OPEN-6 in open-decisions.md — a genuinely open question, not an oversight to be quietly patched into the table above. CD-7’s closing clause, “>90% coverage blocks merges”, was half true and half false in a way worth naming: the threshold shipped at 100% in all five packages, above what was asked, while blocks merges was enforced by nothing, because main was unprotected — a coverage run could go red and a merge still happen. As built, 2026-08-25: main is branch-protected, the required status check is ci (the job id in .github/workflows/ci.yml, not the workflow’s CI display name), and force-pushes and branch deletion are refused. So the clause now holds with one deliberate exemption that must travel with it: a red run withholds the merge button from a contributor, not from the repository owner, because enforce_admins is left off on purpose — agenthropic has exactly one maintainer whose normal working mode is a direct push to main, and admin enforcement would lock the sole maintainer out of their own repository. See the standing correction. CD-9’s per-artifact licensing now has the LICENSE file it presupposed. CD-1 was decided the way it was meant to be — by the Phase-0 evidence, which returned CONDITIONAL GO for the JSONL-primary branch.

CD-8 was not satisfied; it was overridden. Its hard stop said no production code until the spike is green and Gate A is signed. The spike returned CONDITIONAL GO, Gate A is only partially signed, and on 2026-07-11 the owner instructed implementation to begin anyway. That is a legitimate decision by the only person entitled to make it, and it is recorded here as what it was: an override of a gate, not a passage through one. It released dispatching only — every security condition in CD-7 and CD-10 still binds, and the Phase-0 numbers it did not ratify remain PROVISIONAL.


4. The six lenses (consolidated)

4.1 Senior Architect

Verdict: the topology (hooks + JSONL → ingest → SQLite/WAL → SSE → SPA), the ports-&-adapters backbone, and the “hierarchy is persisted data, not UI reconstruction” invariant are all correct and confirmed — build proceeds. The one load-bearing gap v1 named and both externals still miss is the reconciliation contract; v2 resolves it structurally via CD-2/CD-3 (single events_raw substrate + deterministic projection + per-field precedence), which collapses replay-on-startup, idempotency, and outage-backfill into one mechanism.

4.2 Senior Developer (buildability by a solo owner)

Verdict: feasible for one owner; the stack lean (Fastify + better-sqlite3 + React/Vite/D3, pnpm monorepo, SSE) is correct. Developer-critical value sits in three places the externals only half-cover: (1) the normalizer is a two-path problem — forward-link on SubagentStart if it exists, else post-hoc reconstruction on SubagentStop, with the JSONL Agent/Workflow spawn-tool chain as the primary durable linkage source; (2) the cost engine must combine EXPANDED’s versioned model_pricing with v1’s tier/speed/geo buckets + PreCompact baseline as an append-only recomputable layer; (3) licensing is a per-artifact engineering rule, not a vibe.

4.3 Senior QA

Verdict: the QA posture is sound and now well-instrumented. You cannot unit-test “the dashboard”; the four real units are ingest correctness, tree correctness, cost correctness, live-flow correctness, each with a distinct harness. The golden real-session fixture corpus is the #1 QA investment — without it, >90% coverage is high-coverage tests of synthetic happy-paths (false safety).

4.4 Senior Business Analyst

Verdict: the business case is sound and sharper than v1 — a real, narrow gap (no persistent-DAG cockpit with dollar cost + owned persistence + Telegram alerts + security-by-default that is also popular/maintained) against a named incumbent (davila7/claude-code-templates, 28.4k★). The identity split is resolved to personal-first / commercial-clean (LB2).

4.5 Brutal Gap Analysis

The eleven gaps between vision and a shippable v1, ranked by danger. None is fatal; each has a smallest provable slice that closes it, and none is closed by “more documentation”.

# Gap Danger Smallest provable slice
1 Reconciliation/durability contract make-or-break for the moat Phase-0 script: rebuild tree from JSONL alone for one real session, diff vs hook-derived tree
2 Phase 0 must be a real GO/NO-GO spike builds normalizer on unproven premise Gate G0 with a hard ❌ stop; throwaway code only
3 Security sequencing (Phase 1, not 6) cross-origin-vulnerable socket for two phases security-invariant tests in CI from commit one
4 Licensing legal blocker infringement under commercial intent CI provenance/license scan; copy only simple10/hoangsonww
5 Hook-catalog uncertainty both normalizers rest on unverified “twelve” G0.2 enumerates actual hooks + fields
6 Coarse token costing misprices after PreCompact compaction re-pricing test + priceless-model-fails-CI
7 Missing instance/host_id future forced migration one column on every row in the first migration
8 Missing coverage/badges/Pages gate coverage theatre CI gate blocking <90% live from Phase 1
9 Solo-owner scope creep stalls in Phase 2 cut vector-DB to experimental; defer fleet; gate MVP to daily questions
10 SSRF + secret home + retention dial-out, leaked token, unbounded growth SSRF guard test; token via launchd env; retention/redaction in Phase 1
11 External docs’ own defects inherited as if analysis one reconciled KPI; events_raw+events; orchestration_edges first-class

Decisions: G-D1–D7 → the canonical set.

4.6 Holistic / Systems

Verdict: an unusually coherent concept. Coherence is not automatic — it holds only once LB1 and LB2 are locked. Five commitments then reinforce rather than fight each other: the security spine keeps the surface small (helps the solo owner and maintainability); JSONL-primary ingest makes the ground-truth-tokens invariant free; the >90%-coverage CI gate is the enforcement layer that turns the security spine from prose into fact.

Four seams where coherence is stressed:

  1. the persistent-DAG moat’s durability promise depends entirely on LB1;
  2. the cost moat inherits exact tokens but a churn-prone hand-maintained pricing table — precision on the input can be silently betrayed downstream (→ H-COST / CD-4);
  3. ambition vs ownership — a solo owner out-building a 28.4k★ incumbent on five axes + a coverage gate + a docs site is the exact hoangsonww “enterprise-cosplay-over-solo-project” trap (→ scope discipline, CD-10);
  4. “steal these patterns” collides with the flagship grafts being all-rights-reserved (→ CD-9).

The reassuring finding: the constraints are well-aligned — the copyable repos carry the large patterns, the uncopyable ones only small cheap ideas. EXPANDED is itself the negative holistic exemplar (seven identical tables, security to Phase 6) — the precise failure this lens exists to prevent. v2 must be the counter-example.

Decisions: LB1, LB2, H-SEQ, H-APPARATUS, H-COST → CD-1/7/8 + the synthesis rule.


5. Strengths & weaknesses (v2)

Strengths (confirmed and reinforced)

Weaknesses (open until closed by the plan)


6. Acceptance criteria the MVP must meet (quantified)

Merged across lenses; these become the gates in development-plan.md.

Data foundation & reconciliation

Tree correctness

Cost

Security (build-failing, from Phase 1)

Product / business

Delivery bar (Ivan’s, missed by both externals)

Which of these have actually been met (note added 2026-08-15). The criteria are left as written; what follows is the honest split between implemented, measured and neither, because the difference is the whole value of a quantified gate.

The data-foundation and security blocks are implemented and covered by tests, including the append-only substrate, the loopback-or-fail bind, the token-or-fail startup and the exercised restore. The cost block is implemented with one correction: model_pricing is versioned by effective_from only — there is no verified_on — and the seeded rates are self-described in migrations.ts as approximate list prices, a mechanism proof rather than a source of record.

Two criteria are not met, and cannot be met by writing code. Hierarchy correctness ≥95% against a labeled golden corpus is NOT CERTIFIED: the LABEL-ME trees have never been hand-filled, so there is no ground truth to measure against, and every hierarchy number in this corpus remains PROVISIONAL self-check output. Time to understand a session < 30 s has never been measured — not passed, not failed, untested; nobody has sat with a stopwatch. Neither of those two has moved since.

Two of the delivery bar’s own three clauses have moved. “Blocks merges below 90%” was unenforced when this note was written; as built, 2026-08-25 main is branch-protected with ci as the required check, so it now blocks a contributor’s merge and — deliberately, enforce_admins being off for a single-maintainer repository whose normal working mode is a direct push to main — not the owner’s; the reasoning is under the decision register above and in the standing correction. And “the GitHub Pages docs site builds” is no longer a build that publishes nowhere: Pages was enabled on 2026-08-25 and the site serves at https://ivanbbaev.github.io/agenthropic/.


7. Open questions → Phase-0 inputs

The empirical unknowns that must be answered before architecture is poured. These are the Phase-0 gate’s job (CD-8); every one feeds a G0.* probe in the development plan.

  1. G0.1 — THE make-or-break: does ~/.claude/projects/*.jsonl carry the subagent parent→child linkage (Agent/Workflow spawn-tool → child sessionId → parent ref) so the DAG rebuilds from JSONL alone after a full outage? Governs CD-1 (JSONL-primary vs hooks-primary+outbox).
  2. G0.1b — join key: what is the exact key from a JSONL token row to a specific agent_id? Does a subagent get its own transcript with a recorded parent, or must attribution fall back to a timestamp/model heuristic? Determines whether CD-3’s backfill is a hard join or a confidence-scored inference (surface uncertainty in the UI if the latter).
  3. G0.2 — hook catalog: which of the assumed “twelve” actually fire — specifically does SubagentStart exist? If not, edge derivation keys off the Agent/Workflow PostToolUse + SubagentStop.
  4. G0.2b — PreCompact mechanism: does the log carry pre/post-compaction markers that let token_usage preserve a repriceable baseline, or must it be snapshotted at hook time?
  5. State reconciliation: once a late SubagentStop or the JSONL final arrives, does a watchdog-set “unknown/stale” revert to completed or stay flagged? Needs an explicit state-transition rule.
  6. Policy numbers: retention window + payload-redaction rule (which fields, at ingest or at query), “huge payload” reject-vs-truncate threshold, and the coverage-gate scope (line vs branch, per-package vs global, does the web package count).
  7. Pricing source: the authoritative dated source for model_pricing and the refresh cadence that keeps the staleness-fails-CI test honest as the model lineup churns.
  8. Hook-POST auth: is the loopback hook endpoint itself authenticated, and how does the hook script obtain the token without leaking it into ~/.claude scripts?
  9. “OPCⁿ”: define the commercial line concretely or drop the token (BA-D6).

8. What changed vs v1

Area v1 v2 (this document)
Reconciliation named as make-or-break, unresolved resolved — single events_raw substrate + deterministic projection + per-field precedence (CD-2/3)
Schema implicit single events events_raw(immutable) + events(normalized) adopted from EXPANDED (CD-4)
Ports prose “ports & adapters” named adapter set + EventStore + pure Normalizer/Projection (CD-6)
Requirements prose decisions D1–D7 FR/NFR register + ADR set + traceability matrix + negative-test catalogue folded in
Metrics qualitative “daily questions” quantified — hierarchy ≥95%, time-to-understand <30s, 0 lost raw events (CD + §6)
Normalizer tree-building, SubagentStart hedged dual-path design; plan for SubagentStart absence as base case (SD3)
Cost compaction sleeper flagged versioned pricing + buckets + baseline + cost-trust chain (CD-4, H-COST)
Licensing copy-vs-reimplement line hardened to a CI-enforced per-artifact rule (CD-9)
KPI — reconciled the external 100%-vs-95% contradiction to a ≥95% gate
Phase 0 go/no-go gate reaffirmed as an empirical spike, explicitly against EXPANDED’s paperwork degradation (CD-8)

9. Verdict

Build — personal-first, security-spine-first, moat-led. Freeze LB1 (ingest primacy, answered by the Phase-0 spike) and LB2 (personal-first / commercial-clean). Adopt BASE’s sequencing (empirical Phase-0, security in Phase 1) + EXPANDED’s formal apparatus (FR/NFR/ADR/traceability, events_raw+events) + the eleven internal-only items. Quarantine EXPANDED’s generation defects. The consolidated decisions CD-1 … CD-10 are the input to the work-package decomposition in development-plan.md; open work is tracked in ../../TODO.md, completed milestones in ../../DONE.md.


Six-lens re-analysis (Architect · Developer · QA · BA · Gap · Holistic), Opus / high effort, ~468k subagent tokens, run against v1 + BASE + EXPANDED + the adversarial review. Reviewed against the design of record in docs/ai/DESIGN.md.