Authoritative re-analysis of the agenthropic conceptual brief. This v2 folds four inputs into one decision-useful view:
concept-analysis.md + implementation-plan.md (v1);external-docs-review.md;docs/ai/DESIGN.md and the invariants in CLAUDE.md.It was produced by a six-lens senior workflow (Architect · Developer · QA · Business
Analyst · brutal Gap · Holistic), Opus / high effort, ~468k subagent tokens, run against
all of the above. Where v1 and the externals agree, that is recorded as confirmed;
where they diverge, this document decides. The consolidated decisions here are the
input to development-plan.md, which decomposes them into
agent-distributable work packages.
Naming: decisions are tagged by originating lens —
AD*Architect,SD*Developer,QA-D*QA,BA-D*Business,G-D*Gap,LB*/H-*Holistic. The canonical set in §3 dedupes them into ten load-bearing decisions the build actually turns on.
Empirical update — 2026-07-04 desktop probe. The full-corpus read-only probe (
phase0-probe.md, 17 projects · 117 sessions · 33 subagent dirs) pre-answers CD-1 asCONDITIONAL-GO→ build (confidence 85) and corrects the mechanism assumptions below. It de-risks but does not replace the formal Phase-0 spike — WP-S1/WP-S5 still need the paired-capture corpus + Ivan’s tree sign-off, and the WP-S7 GO gate still stands (no production code before it).
- Spawn tool is
Agent/Workflow, neverTask— 0Taskblocks in the real corpus (Agent= 142,Workflow= 29); aTask-keyed parser rebuilds an empty DAG.- The two on-disk layouts are spawn-mechanism-driven, not version-driven — a general-purpose
Agentwrites flatsubagents/agent-<hex>.jsonl, aWorkflowwrites nestedsubagents/workflows/wf_<id>/; both coexist within the same CC version, so the parser branches on directory shape, notversion.- The proven load-bearing hedges are dual-layout parsing (85% nested) + child-transcript token summation (parent rollup ≈ 0%) — these, not the outbox, are what protects the moat.
- The durable outbox (CD-1 fallback / WP-IN11) is YAGNI-leaning and pulled off the v1 critical path — JSONL self-reconciles by backfill (≈ 0 historical crashes); add it only on a real trigger (a sub-second-liveness need, or a hooks-only data source).
Build. The concept is unusually coherent because it descends from a source-level audit
of six real rivals, not a blank page. One dimension — the security posture (loopback-only,
mandatory-token-or-fail-startup, no-spawner, same-origin, no-SSRF) — is genuinely
best-in-class and is the system’s spine, not a bolt-on. The external docs converge with
v1 on the two decision-critical tripwires (neither repeats the refuted “simple10 has no
DAG”; both diagnose hoangsonww’s RCE as a bypassPermissions spawner, not a concurrency
cap), which raises confidence in the whole rather than adding new content.
The gap between this A-grade vision and a shippable v1 is a small number of load-bearing decisions that are cheap on paper and ruinous in code. Two of them govern everything else:
~/.claude/projects/*.jsonl
carry the subagent parent→child linkage well enough to rebuild the DAG from the durable
log alone after an outage? This is make-or-break for the persistent-DAG moat and both
externals silently default to hooks-primary with no durability contract. It cannot be
settled on paper — a Phase-0 empirical spike must answer it before any architecture is
poured.Everything else resolves cleanly around these two.
Synthesis rule for v2: BASE’s sequencing + accuracy (Phase-0 spike, security in
Phase 1) + EXPANDED’s formal apparatus (FR/NFR/ADR/traceability, negative-test
catalogue, quantified metrics, the events_raw+events schema split) + the eleven
items the internal analysis carries that both externals miss. EXPANDED’s own generation
defects — seven byte-identical §10 holistic tables, security deferred to Phase 6, Phase-0
reduced to paperwork, SSRF dropped — are explicitly quarantined and not inherited.
Everything hangs off these; if either is wrong the rest is wasted motion.
Decision: ingest is JSONL-primary + replay-on-startup, with hooks providing sub-second liveness only, and every write is an idempotent upsert on a stable event id — contingent on the Phase-0 spike proving JSONL carries the subagent parent→child linkage. If it does not, fall back to hooks-primary + a durable local outbox/spool (at-least-once, idempotent upsert), and only then.
Why it is load-bearing: this single choice determines whether the persistent-DAG moat is
trustworthy (survives an outage), whether the ground-truth-tokens invariant is satisfied
naturally (tokens read from the durable log, never inferred), whether history is
crash-tolerant, and whether >90% coverage is even achievable (replay-from-fixtures needs a
deterministic durable source). The primacy decision is an output of Phase-0, not an
assumption baked in ahead of it — which is precisely EXPANDED’s error. (The 2026-07-04
desktop probe has since pre-answered it CONDITIONAL-GO (confidence 85); the formal
paired-capture spike confirms rather than decides it — see the empirical-update note above
and phase0-probe.md.)
Decision: build the single-user Mac Mini cockpit; take only the cheap commercial
hedges now (MIT-clean code only; instance/host_id on every row from the first
migration; a schema that does not block tenancy); explicitly defer fleet and
multi-tenancy. Resolve the “OPCⁿ” commercial-line token — define it or drop it — before
it drives tenancy/schema/license-strictness investment.
Why it is load-bearing: it resolves two cross-dimension tensions at once — scope discipline for a solo owner and legal reality. The constraints happen to align: the copyable repos (simple10, hoangsonww) carry the large patterns (tree-building, webhook/alert schema), while the uncopyable ones (cast/disler/nirdiamant) carry only small ideas cheap to reimplement clean-room.
The ten decisions the build turns on, each consolidating the per-lens decisions that agree on it. These are the durable output of this analysis.
| # | Decision | Consolidates | Satisfies (FR/NFR — recovered EXPANDED §3, LOST-5) | The rule |
|---|---|---|---|---|
| CD-1 | Ingest primacy is JSONL-primary + replay-on-startup, contingent on Phase-0; else hooks-primary + durable outbox. | LB1, AD3, SD3, G-D1 | FR-01, FR-05; NFR-OPS-01 | Decided by the Phase-0 diff (tree-from-JSONL vs tree-from-hooks), never assumed. |
| CD-2 | Single immutable substrate + deterministic projection. Both sources write into append-only, idempotency-keyed events_raw; sessions/agents/orchestration_edges/token_usage are a pure replayable projection over it. |
AD1, SD2 | FR-04; NFR-DATA-01/02 | Reconciliation is per-field precedence at projection time, not a two-store merge at query time. |
| CD-3 | Reconciliation precedence. Tokens are JSONL-authoritative (never inferred); interim liveness/state from hooks; final session/agent state + cost from JSONL. token_usage.agent_id is nullable at first write, deterministically backfilled once the agent is known. |
AD2, SD4 | FR-03, FR-05; NFR-DATA-02 | Cross-source idempotent upsert: a fact seen by both a hook and JSONL lands once. |
| CD-4 | Schema: events_raw(immutable) + events(normalized); persisted orchestration_edges (self-ref parent_agent_id, instance/host_id, derived_from_event_id, idempotent); fine-grained token_usage (service_tier / speed / inference_geo + compaction baseline); versioned model_pricing (effective_from, verified_on). |
AD4, G-D6, SD5 | FR-02/03/04/06/08; NFR-MAINT-01 | events_raw is provably append-only (no UPDATE/DELETE path, enforced by test). |
| CD-5 | Transport is SSE with same-origin enforcement from Phase 1. | AD5 | FR-06, FR-07; NFR-SEC-02 | Server→browser-only feed; revisit WebSocket only if bidirectional control is ever needed (it is not). |
| CD-6 | Ports & adapters: named set — HookSource, TokenReader/TokenSource, StoragePort, RealtimeHub, AlertSink, PricingProvider, CostEngine + an EventStore port + a pure Normalizer/Projection; simple10’s strategy-pattern agent classes are the per-runtime adapter (Claude Code now, Codex later). |
AD6, SD1 | FR-01/05/08/09 (via named ports); NFR-MAINT-01 | Keeps the normalizer testable/replayable and the system multi-runtime-ready without a core rewrite. |
| CD-7 | Security + the coverage gate are boundary conditions from commit one, CI-blocking: loopback-or-fail bind; mandatory DASHBOARD_TOKEN-or-fail-startup (timing-safe compare); SSE same-origin; no-spawner grep/static gate; no-SSRF (webhook targets operator-configured, never dialed from a payload); WAL + tested restore; >90% coverage blocks merges. |
AD7, SD8, QA-D3/D4, G-D3, H-SEQ | NFR-SEC-01/02/03, NFR-MAINT-01 | Rejects EXPANDED’s security→Phase 6 / backup→Phase 8. Slice 8 is polish only. |
| CD-8 | Phase 0 is a throwaway GO/NO-GO feasibility spike with a hard ❌ stop — G0.1 ingest-primacy probe · G0.2 hook-catalog enumeration (don’t assume “the twelve”; confirm/deny SubagentStart) · G0.3 tree smoke gate · G0.4 token-reconciliation probe. No production code until green. |
AD-Phase0, G-D2, H-SEQ | — (process gate; de-risks FR-01/03/05/06) | Rejects EXPANDED’s paperwork “decision lock” that validates linkage only at the Phase-4 UI. |
| CD-9 | Per-artifact licensing. COPY simple10 tree/ports + hoangsonww Telegram/webhook schema controlGate + delegation-savings and nirdiamant checkpoint (never view their source while writing). Enforced by a CI provenance/license scan. |
SD6, BA-D4, G-D4, LB2 | — (licensing/provenance; underpins FR-09) | cast/disler/nirdiamant are all-rights-reserved by Berne default — not “ambiguous”. |
| CD-10 | Scope + secrets + retention. MVP = the 5 daily questions; Phase-3 vector-DB “observability-becomes-memory” feed on a labeled experimental track; fleet deferred until a second host exists; Telegram token via token_ref → launchd env / chmod-600 (never in SQLite, never to the browser); retention TTL + payload redaction from Phase 1. |
BA-D1/D3, G-D7, AD8, SD7, LB2 | FR-09; NFR-PRIV-01 | ANTHROPIC_API_KEY stays out of the dashboard env entirely. |
Cost-trust chain (CD-3 + CD-4 crosscut, H-COST): every displayed dollar traces to (ground-truth tokens × a dated, priced model); a model observed in fixtures with no price row FAILS CI — silent staleness is a red build, not a runtime “estimated” label. This extends the byte-exact-tokens guarantee to the priced output.
How the register stands as built (note added 2026-08-15). The ten decisions were implemented, not revised: the register is still the right description of the system, and the rows are left exactly as they were resolved. Four details have to be corrected against the code, and one row records an event rather than a design.
CD-4 names a
verified_oncolumn that was never built. The shippedmodel_pricingtable is(model, bucket, usd_per_mtok, effective_from)keyed onPRIMARY KEY (model, bucket, effective_from); there is no provenance column, andeffective_fromalone carries the versioning. Whether provenance is worth adding is the surviving half of OPEN-6 inopen-decisions.md— a genuinely open question, not an oversight to be quietly patched into the table above. CD-7’s closing clause, “>90% coverage blocks merges”, was half true and half false in a way worth naming: the threshold shipped at 100% in all five packages, above what was asked, while blocks merges was enforced by nothing, becausemainwas unprotected — a coverage run could go red and a merge still happen. As built, 2026-08-25:mainis branch-protected, the required status check isci(the job id in.github/workflows/ci.yml, not the workflow’sCIdisplay name), and force-pushes and branch deletion are refused. So the clause now holds with one deliberate exemption that must travel with it: a red run withholds the merge button from a contributor, not from the repository owner, becauseenforce_adminsis left off on purpose — agenthropic has exactly one maintainer whose normal working mode is a direct push tomain, and admin enforcement would lock the sole maintainer out of their own repository. See the standing correction. CD-9’s per-artifact licensing now has the LICENSE file it presupposed. CD-1 was decided the way it was meant to be — by the Phase-0 evidence, which returned CONDITIONAL GO for the JSONL-primary branch.CD-8 was not satisfied; it was overridden. Its hard stop said no production code until the spike is green and Gate A is signed. The spike returned CONDITIONAL GO, Gate A is only partially signed, and on 2026-07-11 the owner instructed implementation to begin anyway. That is a legitimate decision by the only person entitled to make it, and it is recorded here as what it was: an override of a gate, not a passage through one. It released dispatching only — every security condition in CD-7 and CD-10 still binds, and the Phase-0 numbers it did not ratify remain PROVISIONAL.
Verdict: the topology (hooks + JSONL → ingest → SQLite/WAL → SSE → SPA), the
ports-&-adapters backbone, and the “hierarchy is persisted data, not UI reconstruction”
invariant are all correct and confirmed — build proceeds. The one load-bearing gap v1 named
and both externals still miss is the reconciliation contract; v2 resolves it structurally
via CD-2/CD-3 (single events_raw substrate + deterministic projection + per-field
precedence), which collapses replay-on-startup, idempotency, and outage-backfill into one
mechanism.
orchestration_edges as the moat’s
core artefact (queried from the table, never render-time reconstructed); fine-grained
token_usage + compaction baseline that both externals coarsen away; the instance/host_id
near-zero-cost fleet hedge.events_raw(immutable)+events(normalized) split — the one
genuine schema improvement — elevated to the reconciliation substrate both sources write
into; the named adapter set; the FR/NFR register with proof columns; ADRs as decision-memory.Verdict: feasible for one owner; the stack lean (Fastify + better-sqlite3 + React/Vite/D3,
pnpm monorepo, SSE) is correct. Developer-critical value sits in three places the externals
only half-cover: (1) the normalizer is a two-path problem — forward-link on SubagentStart
if it exists, else post-hoc reconstruction on SubagentStop, with the JSONL Agent/Workflow
spawn-tool chain as the primary durable linkage source; (2) the cost engine must combine
EXPANDED’s versioned model_pricing with v1’s tier/speed/geo buckets + PreCompact baseline as
an append-only recomputable layer; (3) licensing is a per-artifact engineering rule, not a
vibe.
SubagentStart is probably not a real hook — the documented set is
PreToolUse/PostToolUse/UserPromptSubmit/Notification/Stop/SubagentStop/SessionStart/SessionEnd/PreCompact.
Plan for its absence as the base case: JSONL-primary + post-hoc-on-Stop is the likely
real design. G0.2 confirms before the normalizer is committed.apps/server + apps/web (deployables), packages/shared + packages/test-fixtures
(libraries), hooks/ (installable scripts) — reconciles BASE packages/* vs EXPANDED
apps/*; test-fixtures becomes first-class.schema_version.Verdict: the QA posture is sound and now well-instrumented. You cannot unit-test “the dashboard”; the four real units are ingest correctness, tree correctness, cost correctness, live-flow correctness, each with a distinct harness. The golden real-session fixture corpus is the #1 QA investment — without it, >90% coverage is high-coverage tests of synthetic happy-paths (false safety).
token_usage ==
JSONL exact per session; (2) double-replay → byte-identical DB state; (3) DAG-rebuild from
JSONL alone after a simulated outage — the make-or-break test both externals omit.Verdict: the business case is sound and sharper than v1 — a real, narrow gap (no persistent-DAG cockpit with dollar cost + owned persistence + Telegram alerts + security-by-default that is also popular/maintained) against a named incumbent (davila7/claude-code-templates, 28.4k★). The identity split is resolved to personal-first / commercial-clean (LB2).
The eleven gaps between vision and a shippable v1, ranked by danger. None is fatal; each has a smallest provable slice that closes it, and none is closed by “more documentation”.
| # | Gap | Danger | Smallest provable slice |
|---|---|---|---|
| 1 | Reconciliation/durability contract | make-or-break for the moat | Phase-0 script: rebuild tree from JSONL alone for one real session, diff vs hook-derived tree |
| 2 | Phase 0 must be a real GO/NO-GO spike | builds normalizer on unproven premise | Gate G0 with a hard ❌ stop; throwaway code only |
| 3 | Security sequencing (Phase 1, not 6) | cross-origin-vulnerable socket for two phases | security-invariant tests in CI from commit one |
| 4 | Licensing legal blocker | infringement under commercial intent | CI provenance/license scan; copy only simple10/hoangsonww |
| 5 | Hook-catalog uncertainty | both normalizers rest on unverified “twelve” | G0.2 enumerates actual hooks + fields |
| 6 | Coarse token costing | misprices after PreCompact | compaction re-pricing test + priceless-model-fails-CI |
| 7 | Missing instance/host_id |
future forced migration | one column on every row in the first migration |
| 8 | Missing coverage/badges/Pages gate | coverage theatre | CI gate blocking <90% live from Phase 1 |
| 9 | Solo-owner scope creep | stalls in Phase 2 | cut vector-DB to experimental; defer fleet; gate MVP to daily questions |
| 10 | SSRF + secret home + retention | dial-out, leaked token, unbounded growth | SSRF guard test; token via launchd env; retention/redaction in Phase 1 |
| 11 | External docs’ own defects | inherited as if analysis | one reconciled KPI; events_raw+events; orchestration_edges first-class |
Decisions: G-D1–D7 → the canonical set.
Verdict: an unusually coherent concept. Coherence is not automatic — it holds only once LB1 and LB2 are locked. Five commitments then reinforce rather than fight each other: the security spine keeps the surface small (helps the solo owner and maintainability); JSONL-primary ingest makes the ground-truth-tokens invariant free; the >90%-coverage CI gate is the enforcement layer that turns the security spine from prose into fact.
Four seams where coherence is stressed:
The reassuring finding: the constraints are well-aligned — the copyable repos carry the large patterns, the uncopyable ones only small cheap ideas. EXPANDED is itself the negative holistic exemplar (seven identical tables, security to Phase 6) — the precise failure this lens exists to prevent. v2 must be the counter-example.
Decisions: LB1, LB2, H-SEQ, H-APPARATUS, H-COST → CD-1/7/8 + the synthesis rule.
file:line evidence.0.0.0.0 and/or ships no-op auth, and hoangsonww ships an actual RCE spawner.CONDITIONAL-GO (confidence 85; phase0-probe.md),
yet the paired-capture Phase-0 spike still has to confirm it on a labeled corpus; it remains
the single biggest risk carrier.SubagentStart likely absent).Merged across lenses; these become the gates in development-plan.md.
Data foundation & reconciliation
events_raw with a stable idempotency key; re-ingesting the
same log yields byte-identical events_raw and an identical projected DB state.token_usage per session == JSONL ground truth exactly; a static check proves no token
row can originate from inference.orchestration_edges/token_usage; zero data loss.orchestration_edges is persisted; the global/cross-session DAG is served by querying the
table, never render-time reconstruction. Every row carries a non-null instance/host_id.token_usage.agent_id may be NULL at first write, backfilled deterministically; a
reconciliation test asserts no double-count or misattribution after backfill.Tree correctness
Agent/Workflow spawn chain even if SubagentStart never fires.SubagentStop → explicit “unknown” state within the watchdog window, never a
permanent “working”.Cost
model_pricing is versioned
(effective_from/verified_on).Security (build-failing, from Phase 1)
127.0.0.1 only and FAILS startup when DASHBOARD_TOKEN is unset (never
“auth disabled”); token compare is timing-safe; SSE rejects cross-origin; no wildcard CORS.events_raw exposes no UPDATE/DELETE path (enforced by test); SQLite runs WAL; a backup is
taken and a restore is exercised at least once per release candidate.Product / business
instance/host_id hedge present).Delivery bar (Ivan’s, missed by both externals)
127.0.0.1
(docs public, app never exposed).Which of these have actually been met (note added 2026-08-15). The criteria are left as written; what follows is the honest split between implemented, measured and neither, because the difference is the whole value of a quantified gate.
The data-foundation and security blocks are implemented and covered by tests, including the append-only substrate, the loopback-or-fail bind, the token-or-fail startup and the exercised restore. The cost block is implemented with one correction:
model_pricingis versioned byeffective_fromonly — there is noverified_on— and the seeded rates are self-described inmigrations.tsas approximate list prices, a mechanism proof rather than a source of record.Two criteria are not met, and cannot be met by writing code. Hierarchy correctness ≥95% against a labeled golden corpus is NOT CERTIFIED: the
LABEL-MEtrees have never been hand-filled, so there is no ground truth to measure against, and every hierarchy number in this corpus remains PROVISIONAL self-check output. Time to understand a session < 30 s has never been measured — not passed, not failed, untested; nobody has sat with a stopwatch. Neither of those two has moved since.Two of the delivery bar’s own three clauses have moved. “Blocks merges below 90%” was unenforced when this note was written; as built, 2026-08-25
mainis branch-protected withcias the required check, so it now blocks a contributor’s merge and — deliberately,enforce_adminsbeing off for a single-maintainer repository whose normal working mode is a direct push tomain— not the owner’s; the reasoning is under the decision register above and in the standing correction. And “the GitHub Pages docs site builds” is no longer a build that publishes nowhere: Pages was enabled on 2026-08-25 and the site serves at https://ivanbbaev.github.io/agenthropic/.
The empirical unknowns that must be answered before architecture is poured. These are the
Phase-0 gate’s job (CD-8); every one feeds a G0.* probe in the development plan.
~/.claude/projects/*.jsonl carry the subagent
parent→child linkage (Agent/Workflow spawn-tool → child sessionId → parent ref) so the DAG rebuilds
from JSONL alone after a full outage? Governs CD-1 (JSONL-primary vs hooks-primary+outbox).agent_id?
Does a subagent get its own transcript with a recorded parent, or must attribution fall back
to a timestamp/model heuristic? Determines whether CD-3’s backfill is a hard join or a
confidence-scored inference (surface uncertainty in the UI if the latter).SubagentStart exist? If not, edge derivation keys off the Agent/Workflow PostToolUse + SubagentStop.token_usage preserve a repriceable baseline, or must it be snapshotted at hook time?SubagentStop or the JSONL final arrives, does a
watchdog-set “unknown/stale” revert to completed or stay flagged? Needs an explicit
state-transition rule.model_pricing and the refresh
cadence that keeps the staleness-fails-CI test honest as the model lineup churns.~/.claude scripts?| Area | v1 | v2 (this document) |
|---|---|---|
| Reconciliation | named as make-or-break, unresolved | resolved — single events_raw substrate + deterministic projection + per-field precedence (CD-2/3) |
| Schema | implicit single events |
events_raw(immutable) + events(normalized) adopted from EXPANDED (CD-4) |
| Ports | prose “ports & adapters” | named adapter set + EventStore + pure Normalizer/Projection (CD-6) |
| Requirements | prose decisions D1–D7 | FR/NFR register + ADR set + traceability matrix + negative-test catalogue folded in |
| Metrics | qualitative “daily questions” | quantified — hierarchy ≥95%, time-to-understand <30s, 0 lost raw events (CD + §6) |
| Normalizer | tree-building, SubagentStart hedged |
dual-path design; plan for SubagentStart absence as base case (SD3) |
| Cost | compaction sleeper flagged | versioned pricing + buckets + baseline + cost-trust chain (CD-4, H-COST) |
| Licensing | copy-vs-reimplement line | hardened to a CI-enforced per-artifact rule (CD-9) |
| KPI | — | reconciled the external 100%-vs-95% contradiction to a ≥95% gate |
| Phase 0 | go/no-go gate | reaffirmed as an empirical spike, explicitly against EXPANDED’s paperwork degradation (CD-8) |
Build — personal-first, security-spine-first, moat-led. Freeze LB1 (ingest primacy,
answered by the Phase-0 spike) and LB2 (personal-first / commercial-clean). Adopt BASE’s
sequencing (empirical Phase-0, security in Phase 1) + EXPANDED’s formal apparatus
(FR/NFR/ADR/traceability, events_raw+events) + the eleven internal-only items. Quarantine
EXPANDED’s generation defects. The consolidated decisions CD-1 … CD-10 are the input to the
work-package decomposition in development-plan.md; open work is tracked
in ../../TODO.md, completed milestones in ../../DONE.md.
Six-lens re-analysis (Architect · Developer · QA · BA · Gap · Holistic), Opus / high effort,
~468k subagent tokens, run against v1 + BASE + EXPANDED + the adversarial review. Reviewed
against the design of record in docs/ai/DESIGN.md.