agenthropic is a self-hosted, local-first cockpit for observing and visualising
Claude Code agent and subagent activity — real sessions, on your own machine, with
no cloud dependency and no telemetry egress. This page covers what the tool is, the
one-paragraph pitch, the moat that justifies building it instead of adopting an
existing dashboard, who it’s for, and where the project actually stands today. The
key takeaway up front: it observes, it never operates — it ingests Claude Code’s
own lifecycle hooks and the ground-truth ~/.claude/projects/*.jsonl logs into a
SQLite database you own, and renders the resulting subagent tree, token cost, and
(post-1.0) Telegram alerts — and, when this page was written, it was still in the
bootstrap phase: the design settled, the code not yet scaffolded. (That last clause
is no longer true — see the update immediately below.)
Update — 2026-07 (as built). This page was written pre-code. Implementation began 2026-07-11, by explicit owner override of the CD-8 “no production code before the Phase-0 GO” gate. agenthropic is no longer a design; it is a running program, and the “bootstrap phase” framing below is design history.
What runs today, all verified against the repository: the Fastify server bound to
127.0.0.1:4317and gated by a mandatoryDASHBOARD_TOKEN; the SQLite/WAL substrate with a forward-only migration runner (eighteen migrations) and a daily backup timer; JSONL corpus ingest with replay-on-startup and tail-follow polling that re-reads only new bytes; the persisted subagent DAG (orchestration_edges, five structural join provenances since migration 13 —tool_use,directory,task_notification,queue_operationand the legacylegacy_explorefallback); the cost engine including compaction repricing and delegation savings; the hook receiver and its installer; a five-value agent status lifecycle (working,waiting,completed,error,unknown) with a watchdog that ages an unobserved agent rather than guessing at its ending; the SSE realtime hub; the read API; and all four dashboard views — live status, session tree, global DAG, cost/Sankey — plus a per-session cost-analysis panel. It is not released: no tag, no published package — version0.3.0, publishable in principle (publishConfig.access: public, noprivateflag on the root package) but never published, run from a checkout.Test figures, re-measured 2026-09-18: 131 test files / 2428 tests (re-measured again 2026-09-23: 140 test files / 2621 tests; 2026-09-26: 167 test files / 3060 tests), with 100% statements, branches, functions and lines enforced in all five packages —
packages/test-fixtureswas folded into the gate rather than left outside it. Two things that does not mean. It is not unconditionally merge-blocking:mainhas been branch-protected since 2026-08-25 and requires thecicheck, so a red run does withhold the merge button — from a contributor, and not from the repository owner, becauseenforce_adminsis deliberately off (agenthropic has one maintainer whose normal mode is a direct push tomain, and turning admin enforcement on would lock the sole maintainer out of their own repository). The full note is the standing correction. And covering the code is not measuring the output — the hierarchy-accuracy exit gate reports NOT CERTIFIED at n = 0 because no session has been hand-labeled.Three corrections to the prose below. (1) Four hooks, not twelve. The installer registers
UserPromptSubmit,Stop,SubagentStopandPreCompact;SubagentStartdoes not exist, and no hook contributes structure — hooks are liveness only. The subagent tree is built entirely from the JSONL transcripts. (2) There is no separate Normalizer→Projection pipeline. The parser reads a transcript and writessessions/agents/orchestration_edges/token_usagein one transaction per session;events_rawholds hook events only. This is a deliberate, recorded divergence from the design sketch. (3) Telegram alerting is not built and may never be. It is v2.0 work, entered only via checkpoint KC-5 (14 consecutive days of real daily v1.0 use plus ≥3 friction-log entries asking for alerts); the operator-alerts API and UI were cut outright. The server currently makes no outbound network request of any kind.Two caveats this page must not lose. The Phase-0 spike numbers below (including the confidence-85
CONDITIONAL-GO) remain PROVISIONAL — they were scored against machine inventories, not a hand-labeled corpus, and the ratification act is still open. And the v1.0 “<30 seconds to understand a session” claim is unmeasured: nobody has sat in front of the running dashboard with a real corpus and timed it.
A self-hosted, local-first cockpit for Claude Code agent/subagent activity on a Mac Mini M4 — persisted subagent DAG + dollar-cost/delegation-savings + Telegram alerts
- owned persistence — differentiated from the 28.4k★ baseline (
claude-code-templates) by the four things it lacks, and from the whole field by a loopback-only, no-spawner, mandatory-token security posture.
This is the idea-in-one-paragraph from the project’s own concept analysis, reproduced
verbatim because it is the most precise summary available (see
docs/analysis/README.md, “The idea in one paragraph”).
Two things in it matter more than the rest: the phrase “the four things it lacks”
(the market gap the moat is carved from — see below) and “Mac Mini M4” — this is a real, personal-first tool built to
run one operator’s actual subagent-heavy workflow, not a generic multi-tenant SaaS.
At its core, agenthropic closes one loop: Claude Code emits lifecycle hook events as agents and subagents run; those events (plus the JSONL transcripts Claude Code already writes) are ingested into an owned, persisted store; a browser renders the resulting orchestration graph and cost picture live. This was the target architecture — the design of record at the time of writing:
Claude Code (subagents)
│ hooks (lifecycle events)
▼
hook-ingest ──► SQLite (persisted, WAL) ──► SSE ──► browser SPA
▲ │ (DAG + Sankey)
│ └──► webhook sink ──► Telegram relay (@baev_bot_bot)
└── reads ~/.claude/projects/*.jsonl (ground-truth token counts)
As built, the loop is real but the emphasis is inverted, and one branch is missing. JSONL is the load-bearing input, not the sidecar: the corpus watcher reads
~/.claude/projects/*.jsonland the parser writessessions/agents/orchestration_edges/token_usagedirectly, in one transaction per session. Thehook-ingestarrow is thinner than drawn — a hook delivery lands one raw row inevents_rawplus one identifier-only liveness row inevents, and never touches the hierarchy or token tables. The webhook-sink → Telegram branch does not exist: it is v2.0 work behind KC-5, and the server makes no outbound network request at all today. The SQLite → SSE → browser SPA path is built and running.
Two invariants hold across every layer of that loop and are non-negotiable design decisions, not implementation details:
~/.claude/projects/*.jsonl — never
inferred or estimated from tool-call heuristics.parent_agent_id foreign key on
an agents table), not something the browser reconstructs on the fly from a flat
event log.The design assumed ingestion would listen for Claude Code’s full lifecycle hook set
(PreToolUse, PostToolUse, UserPromptSubmit, Notification, Stop, SubagentStart/
SubagentStop, SessionStart/SessionEnd, PreCompact, PermissionRequest,
PostToolUseFailure), with SubagentStart/SubagentStop given dedicated handling
because they were expected to feed the hierarchy tables directly. One hedge applied to
that list even then: it was 12 hooks assumed, 9 documented — SubagentStart,
PermissionRequest, and PostToolUseFailure are not in Claude Code’s documented
nine-event set, and whether the runtime fires them was unconfirmed until the Phase-0 spike
reported (see Hook ingestion, “The SubagentStart hedge”).
As built, the hedge was right and then some. The spike confirmed
SubagentStartis not a real hook, and the installer registers exactly four:UserPromptSubmit,Stop,SubagentStop,PreCompact. More importantly, no hook has dedicated structural handling — not evenSubagentStop. Hooks are liveness only, never structure: no hook creates an agent row, asserts a parent→child edge, or writes a token row. The hierarchy comes entirely from the JSONL transcripts, via the parser’s join paths — four modern ones (tool_use,directory,queue_operation,task_notification) pluslegacy_explore, a last-resort fallback for pre-2.1.71 transcripts added in 2026-08. Each edge stores which path produced it, andlegacy_exploreis deliberately not filed astool_use: one edge was observed in a materialised spawn block, the other was inferred from a degraded legacy shape, and once both are written under the same name no reader can ever tell them apart again. Theevents_raw→ normalizedevents→ projection chain below was likewise never built as separate stages —events_rawholds hook events only, and JSONL is parsed straight into the projections. Both divergences are deliberate and recorded.
The full hook catalog and the ingest design (events_raw, events,
orchestration_edges, token_usage) are covered in depth in the architecture
section — see Architecture overview,
Hook ingestion, and
Ingest & reconciliation.
“Local-first” is not a marketing label here — it is a set of hard constraints the project holds itself to everywhere:
127.0.0.1 loopback only — never 0.0.0.0.timingSafeEqual — not opt-in, and
never a no-op when unset.claude or any subprocess driven by request input. This is a
deliberate, permanent line: one of the six dashboards surveyed during
due diligence exposes exactly this as an RCE (/api/run accepts a
permission-mode from the request body whose allow-list includes
bypassPermissions) — agenthropic does not build this surface, ever.None of this is aspirational polish added at the end; it is called out as non-negotiable in the project’s own instructions precisely because every rival dashboard surveyed during due diligence got at least one of these wrong (see the Threat model page for the per-project breakdown).
Six existing Claude Code observability dashboards were audited before deciding to
build agenthropic at all, plus the 28.4k-star zero-install baseline
(davila7/claude-code-templates), which already nails self-hosted, zero-install,
live token attribution — but ships as a flat leaderboard. Across all seven, five
capabilities are confirmed absent everywhere: a global, persistent, per-instance
orchestration DAG (every existing tool has at most a session-scoped tree with
event-derived, non-persisted edges); live dollar-cost attribution and
delegation-savings (quantifying what routing to Haiku/Sonnet instead of a top-tier
model actually saves); a Telegram alert sink; cross-machine/fleet
aggregation; and persistence you control for historical, time-series analysis.
That gap — not a wish list, but five things independently verified absent across the
whole surveyed field — is the reason to build greenfield rather than fork or adopt an
existing project; the tagline’s “the four things it lacks” counts the four of those
five measured directly against the baseline (fleet aggregation is the fifth, absent
across the entire field, baseline included). The moat proper is narrower still:
two capabilities that are genuinely hard to retrofit — the persistent
cross-session DAG and dollar-cost attribution. Telegram alerting is planned as
a post-1.0 convenience, not a differentiator, and fleet aggregation is deferred until
a second host physically exists. The full breakdown, including which project each
idea is borrowed from and why forking was rejected, is on
The moat — why build.
As built, the moat proper is real; the two conveniences are not. Both hard capabilities ship: the persistent cross-session DAG (
agents+orchestration_edges, the latter keyed byinstance/host_id) and dollar-cost attribution including delegation savings. Three P0 moat proofs (plus a fourth over real HTTP) guard them and, since 2026-08-25, block the merge button for anyone who is not the repository owner —mainrequires thecicheck, whileenforce_adminsstays off by design so the sole maintainer is not locked out of their own repository (standing correction). The proofs: Σtoken_usageequals the JSONL as checked by an independently written reader inside the test; a double replay produces a byte-identical database; the DAG rebuilds from JSONL alone after a simulated outage, with hooks proven liveness-only (appending them leaves the DAG dump unchanged). Telegram alerting is not built — it is v2.0, entered only via KC-5, and its operator-alerts API and UI were cut outright. Fleet aggregation is not built either; only the schema key exists, and no second host does. One caveat on the survey itself: the six rivals and the baseline were assessed by reading their code and docs during due diligence — none was installed and run head-to-head, and the project’s friction log was never opened, so nothing here rests on lived comparative use.
| Built for | A senior engineer running genuinely subagent-heavy Claude Code sessions on their own machine, who wants a queryable, persisted record of what every agent and subagent actually did — including cost — without sending anything to a third-party service. |
| Identity | Personal-first, commercial-clean: a single-user cockpit for one operator’s own hardware, deliberately not a multi-tenant product today. The schema carries an instance/host_id key on every row from the first migration so fleet aggregation is possible later, but multi-tenancy and fleet views are explicitly deferred, not shipped. (As built: the hedge is narrower than this promised — instance and host_id are NOT NULL on orchestration_edges only, not on every table. Fleet views remain unbuilt, as stated.) |
| Not (yet) for | Teams needing a shared, hosted, multi-tenant dashboard, or fleet-wide observability across many machines — that is an explicit later phase, not the MVP. |
| Prerequisite mindset | Comfortable self-hosting a small service, reading SQL, and treating “no telemetry egress” as a feature rather than friction. |
Update — 2026-07 (as built). The bootstrap phase is over. This section describes the project as it stood before 2026-07-11. Each bullet below now carries an
*(As built: … )*note with what it resolved to. The heading is kept because other pages link to it.
As of this writing, agenthropic is in the bootstrap phase:
better-sqlite3 + React/Vite/D3 in a pnpm monorepo (server + web),
aligned with a sibling project’s pattern, but repo structure and the MVP schema
scope are still open. (As built: no longer true. The leaning became the decision and
shipped unchanged — apps/server, apps/web, packages/shared, packages/core,
packages/test-fixtures, hooks/, on Node 22, with 131 test files / 2428 tests
passing and 100% statements/branches/functions/lines enforced in all five
packages, re-measured 2026-09-18 *(re-measured again 2026-09-23: 140 test files /
2621 tests; 2026-09-26: 167 test files / 3060 tests, same 100%). What has not happened is a release: no tag, no
published package — version 0.3.0, publishable (publishConfig.access: public)
but never published.)*CONDITIONAL-GO (confidence 85) by a read-only probe of the real
~/.claude/projects corpus on 2026-07-04: JSONL is a trustworthy, outage-surviving
source, provided the parser walks both on-disk layouts (85% of agent files are
nested) and sums tokens from child transcripts. Those two — dual-layout
parsing and child-transcript token summation — are the proven load-bearing
hedges; a hooks-primary ingest with a durable outbox is a deferrable, contingent
fallback (JSONL self-reconciles by backfill), pulled off the v1 critical path unless
a sub-second-liveness or hooks-only data need actually appears. The probe de-risks
but does not replace the formal spike: the paired-capture corpus and the operator’s
tree sign-off still run, and this go/no-go decision stays empirical, not assumed —
tracked in the repository’s own TODO.md as Gate A (decision approval) and the
Phase 0 feasibility spike’s GO/CONDITIONAL-GO/NO-GO verdict (WP-S7).
(As built: the gate did not hold as written. Implementation began 2026-07-11 by
explicit owner override of CD-8, before the paired-capture corpus and the operator’s
tree sign-off were completed. Both load-bearing hedges — dual-layout parsing and
child-transcript token summation — are implemented and covered by the three P0 moat
proofs, but the spike numbers themselves remain PROVISIONAL until ratified
against the hand-labeled corpus. Treat the confidence figure above as an estimate,
not a measurement.)STYLE-GUIDE.md § “As-built amendments” for how.)The empirical basis for the
CONDITIONAL-GOverdict — the corpus census, the four CD-1 questions answered with real numbers, and the parser acceptance gate — is written up indocs/analysis/phase0-probe.md.
For the phase-by-phase build sequence once Phase 0 goes green, see the Roadmap. For the open design questions this status section intentionally does not resolve, see the FAQ.