agenthropic

The DAG moat

This page explains the single hardest, most defensible piece of agenthropic’s design: the global, persisted, per-instance orchestration DAG, why it is unclaimed ground across every audited rival, how it is actually built (dual-path edge derivation, WP-IN8), and how its correctness is proven even after a total outage (rebuild-from-JSONL-alone, one of Phase 3’s three release-blocking tests). The key takeaway: every other project in the audit renders a subagent tree by walking an event log at render time, scoped to one session; agenthropic instead writes orchestration_edges rows once, at ingest/projection time, keyed per agent instance (not per subagent type), tagged with a host/instance key from day one, and queries that table for both the session tree and the cross-session global DAG — never reconstructing it in the browser (DESIGN.md §2, §4, §6). This is deliberately the hardest thing on the roadmap to get right, which is exactly why it is the moat and not a nice-to-have.

Update — 2026-07 (as built). The moat artifact exists. orchestration_edges is a real table (migration 5, apps/server/src/db/migrations.ts), written at ingest time in the same transaction as sessions/agents/token rows, and served by real endpoints (GET /api/sessions/:id/tree and GET /api/dag/global — a query over persisted rows, never a render-time reconstruction). Three things landed differently than this page sketches:

Why this is “the moat” and not just a feature

DESIGN.md §2 lists five capabilities “confirmed absent across all six audited projects” as the reason to build rather than adopt one of them; item one is this DAG:

“Global, persistent, per-instance orchestration DAG. Everyone has at most a session-scoped tree with event-derived, non-persisted edges. A real, queryable, cross-session per-instance graph is unclaimed ground.” (DESIGN.md §2, item 1)

Three words in that sentence carry the entire design decision, and each maps to a concrete, checkable property of the schema and the build plan below:

Word What it rules out What agenthropic does instead
persisted Deriving edges from the event/JSONL stream at render time, every time the UI opens Writing orchestration_edges rows once, at ingest/projection time (DESIGN.md §4)
per-instance Collapsing many concrete subagent runs into a handful of “type” nodes in the diagram One row/node per actual agent instance — no aggregation by subagent_type
cross-session A tree that only makes sense inside the one session that produced it A graph a query can span across sessions, keyed by a host/instance column, for a global view

The rest of this page is the mechanics behind each of those three words.

Persisted, not event-derived at render time

DESIGN.md §4 states the requirement directly, immediately after grafting the agents self-referential hierarchy from hoangsonww:

“For the moat (§2.1), extend beyond any existing schema: edges must be persisted (not event-derived at render time) and per-instance (not type-aggregated), and carry an instance/host key for future fleet aggregation (§2.4).” (DESIGN.md §4)

Concretely, this means orchestration_edges is a real table with real rows, written during ingestion/projection — not a data structure the frontend builds by walking events (or a JSONL transcript) every time a page loads. The development plan encodes this as its own persisted-data work package:

WP-D7 — data, size M, deps D6, D4: “orchestration_edges persisted table (the moat artifact). Duplicate logical edge → exactly one row (UNIQUE + INSERT OR IGNORE). Non-null instance/host_id.” (development-plan.md, Track D)

Two properties fall directly out of that Done-when clause:

By the time the read side needs to serve a tree or a global graph, the answer is already sitting in a table — it is queried, never rebuilt in the request path. That is exactly the property WP-U3 (session/agent/subagent-tree endpoints) and WP-U8 (the global DAG view) are built to prove:

WP-U3: “GET /sessions/:id/tree built from a query over orchestration_edges (proven, not reconstruction).” (development-plan.md, Track U)

WP-U8: “Global persistent per-instance orchestration DAG view (the moat). Spans multiple sessions, sourced from a query over persisted edges.” (development-plan.md, Track U)

As built: both endpoints exist in apps/server/src/api/routes.ts — GET /api/sessions/:id/tree and GET /api/dag/global — and both answer from a query over the persisted orchestration_edges rows. Idempotent writes shipped exactly as named: INSERT OR IGNORE against UNIQUE (session_id, parent_agent_id, child_agent_id), so whole-session re-ingest (the replay mechanism) collapses duplicate derivations to one row. instance and host_id are NOT NULL from migration 5.

AMENDED 2026-09-23 (J-3). One thing the “query over persisted edges” framing does not settle, and which the global query got wrong until this date: how a node’s tokens are attributed. The graph spans sessions; a node does not. GET /api/dag/global now groups token_usage by (agent_id, session_id) and joins on both columns - the scoping the per-session tree has always used - so a node’s totalTokens / costUsd / unpricedTokens are that agent’s usage in its own session, and the two endpoints report the same figure for the same agent. Previously the global query grouped by agent_id alone, so an agent id appearing in more than one session had every session’s tokens summed onto the one node: a cross-session figure rendered as one agent’s spend, larger than the session tree’s figure for the same node and carrying nothing to say so. Nothing in the schema prevents the collision - no foreign key ties token_usage.agent_id to a session, and a usage row naming that id in another session is that session’s unattributed usage. Spanning sessions is the moat claim; silently summing across them never was.

Per-instance, not type-aggregated

The second word matters because at least one audited rival built something that looks like a DAG but collapses individual agent runs into categories. DESIGN.md §6 is explicit that this is not the same thing, and names it directly:

“hoangsonww’s ‘DAG cockpit’ is oversold — OrchestrationDAG.tsx is a type-aggregated 3–4 layer diagram; its true nesting is a collapsible indented tree reconstructed post-hoc on SubagentStop. Study its D3 Sankey / aggregate polish, but know that the global, persistent, per-instance DAG (§2.1) is the thing we still have to build.” (DESIGN.md §6)

The distinction: a type-aggregated diagram draws one node per subagent_type (say, one node for “code-reviewer” no matter how many times it ran) and shows a handful of layers of category flow. A per-instance graph draws one node per actual agent row — every individual invocation, with its own id, its own parent, its own status — the way agents.parent_agent_id already models the hierarchy (DESIGN.md §4). agenthropic’s DAG is explicitly the second kind; the first kind is useful for aggregate visual polish (hence “study its D3 Sankey”), but it is not a substitute for the per-instance graph this project still has to build.

The host/instance key: built for a fleet that doesn’t exist yet

DESIGN.md §2 names cross-machine/fleet aggregation as a separate absent capability (item 4) — “All six [audited projects] are single-host.” The DAG schema is designed so that adding fleet support later is a query change, not a migration:

So today the key exists and is populated (single Mac Mini, one instance/host_id value), but there is no fleet UI or cross-host rollup yet — the column is a hedge against a future migration, not a shipped feature. Until Phase 5+ lands, treat any “fleet” framing as schema-ready, not built.

Dual-path edge construction (WP-IN8)

As built, the “dual-path” collapsed to one path — deeper than designed. The hook leg was never built: SubagentStart does not exist, and no hook (not even SubagentStop) asserts an edge. What shipped instead is the parser’s five JSONL join paths — tool_use (the Agent/Workflow spawn’s tool_use.id matched to the child’s meta.toolUseId), directory (nested workflows/wf_*/ containment), task_notification, queue_operation, and the defensive legacy_explore fallback for pre-2.1.71 transcripts — each recorded in the row’s source column, so an edge’s provenance stays queryable. The parser decides directory first, by file layout alone; for a flat file it tries tool_use, then queue_operation, then task_notification, and legacy_explore last — the first match wins. An agent none of the five can join gets no edge (visible as an orphan), never a guessed one. The section below is the design record of why the hedge existed.

This is the core mechanism, and the single largest execution-risk item on the moat. The development plan’s work package is explicit that the persisted edge must be derivable two independent ways:

WP-IN8 — backend, size L, deps IN7, S1, S2, S3: “Dual-path edge derivation → persisted orchestration_edges (moat core). Correct parent→child tree via the JSONL Agent/Workflow spawn-chain even if SubagentStart never fires.” (development-plan.md, Track IN)

Read literally, this names two paths that both feed the same table:

  1. The hook path. Live lifecycle events — SubagentStart/SubagentStop — arrive at ingest and (once projected through WP-IN7) can directly assert a parent→child edge as it happens.
  2. The JSONL Agent/Workflow spawn-chain path. Independently of any hook firing, the ~/.claude/projects/*.jsonl transcript records the Agent and Workflow tool invocations that spawn subagents — a general-purpose Agent spawn writes a flat subagents/agent-<hex>.jsonl (+ .meta.json), a Workflow spawn writes a nested subagents/workflows/wf_<id>/ subtree — and the parser branches on that directory shape, not on CC version. Walking that chain in the transcript lets the same parent→child edge be derived even if SubagentStart never fires for that session — the exact contingency WP-IN8’s Done-when calls out by name.

Empirically grounded — 2026-07-04 corpus probe. The full-corpus read-only probe found zero Task tool blocks in the real corpus; every spawn is an Agent (142) or Workflow (29) tool call, so the JSONL path keys on those two tools and branches on directory shape (flat agent-<hex>.jsonl vs nested workflows/wf_<id>/ — 85% of agent files are nested), never on CC version. A Task-keyed parser would rebuild an empty DAG. This also pre-answers CD-1 as CONDITIONAL-GO → build (confidence 85), de-risking but not replacing the formal Phase-0 spike (S1–S3) or the production-code gate below. Full evidence: Phase-0 corpus probe.

 live hook events                  ~/.claude/projects/*.jsonl
 (SubagentStart/Stop,               (Agent/Workflow spawn chain,
  when they fire)                    always present)
        │                                    │
        ▼                                    ▼
   ┌─────────────────────────────────────────────┐
   │        WP-IN8 — dual-path edge derivation     │
   │   (consumes WP-IN7's projected sessions/      │
   │    agents; either path can assert an edge)    │
   └─────────────────────────────────────────────┘
                        │
                        ▼
        orchestration_edges  (UNIQUE + INSERT OR IGNORE,
                              non-null instance/host_id — WP-D7)

Both paths write into the same UNIQUE-constrained table (WP-D7), so if a hook event and a JSONL-derived inference describe the same logical edge, the second write is a no-op (INSERT OR IGNORE) rather than a duplicate row or a conflict. This is what makes the two paths genuinely redundant rather than merely “two features that both sort of build a tree” — either one, alone, is sufficient to populate the table correctly for a given edge.

WP-IN8 depends on WP-IN7 (the projection that turns raw events into sessions/agents/token_usage — see ingest & reconciliation for the projection in full) and on the Phase-0 spike outputs S1–S3, i.e. this mechanism is not designed in a vacuum — it is built against the labeled real-session corpus captured before any production code exists.

Downstream of WP-IN8, two more work packages close the loop that dual-path derivation opens:

As built: WP-IN12 shipped as written — stale non-terminal agents flip to 'unknown' after a bounded window (DASHBOARD_WATCHDOG_MINUTES, default 10, PROVISIONAL), and a later re-ingest lets JSONL evidence win the status back. WP-IN9’s backfill pass was never needed: token_usage.agent_id is attributed inside the parser, before any write, in the same transaction — a NULL there means genuinely unattributable, not “awaiting backfill”.

Rebuild-from-JSONL-alone: the test that proves the moat is real

A dual-path design is only as good as its proof that the fallback path actually works under a real outage, not just in the happy path where both signals agree. Phase 3’s exit gate names three release-blocking correctness tests, and the third one is specifically about this DAG:

“Three P0 tests green & merge-blocking (Σtoken_usage==JSONL exact; double-replay byte-identical DB; DAG rebuild from JSONL alone after a simulated outage); hierarchy ≥95% vs the labeled corpus even without SubagentStart; missing-Stop→unknown; PreCompact reprices vs baseline; no priceless model; 12-scenario negative catalogue green.” (development-plan.md §3, Phase 3 row)

The relevant work packages that assemble and run this test:

WP Role
WP-IN10 “Replay-on-startup + deterministic full projection rebuild. Double-replay → byte-identical events_raw and projected DB.” (deps IN2, IN6, IN7, IN8, IN9)
WP-IN13 “Reconciliation / idempotency / DAG-rebuild suite (P0 blockers). All three P0 tests green in CI and blocking.” (deps IN10, IN9, X1)
WP-X3 “Three P0 reconciliation release-blocker tests. Σtoken_usage==JSONL exact; double-replay byte-identical; DAG-rebuild-from-JSONL-alone.” (deps X2, IN10, IN7, D1)

What “rebuild from JSONL alone” means concretely: simulate the outage of the live hook stream entirely — the first path in the dual-path diagram above goes dark — and confirm that replaying only the ~/.claude/projects/*.jsonl transcript through the projection and WP-IN8’s JSONL Agent/Workflow spawn-chain path still reconstructs the same persisted orchestration_edges rows. This is the release-blocking proof that the fallback path in WP-IN8 is not a paper guarantee. It sits on the development plan’s own “moat spine” — the sub-chain called out as the schedule-critical thread independent of alerting:

“The moat spine (independent of alerts) is the sub-chain …D4 → IN1 → IN6 → IN7 → IN8 → IN9 → IN10, landing the three P0 reconciliation tests at wave 16 — protect that on schedule.” (development-plan.md §4)

The Phase 3 exit gate also requires the reconstructed hierarchy to hit ≥95% accuracy against a hand-labeled real session even without SubagentStart firing at all — i.e. the JSONL-only path has to carry the tree on its own, not just contribute alongside a healthy hook stream, to clear the gate.

As built: the three P0 tests exist as real test files in the server suite and are green — token-sum exactness, double-replay idempotence (re-ingesting an unchanged session writes nothing new), and DAG-rebuild-from-JSONL-alone. The last one is no longer a fallback proof: since hooks never feed edges, JSONL-alone is the only branch, and the test asserts the system’s normal operation, not an outage mode.

On the “merge-blocking” in the exit gate quoted above: that phrase is development-plan.md’s own wording and is left exactly as written. Read against this repository it was only half-true until 2026-08-25 — the tests did fail the CI run, but main carried no branch-protection rule, so a red run withheld nothing. main is branch-protected now: the required status-check context is ci (lowercase — the job id in .github/workflows/ci.yml; the workflow’s CI display name is not the context), and force-pushes to main and deletion of main are refused for everyone. So a red P0 test does withhold the merge button — from a contributor. It does not withhold it from the repository owner: enforce_admins is deliberately off, because agenthropic has exactly one maintainer whose normal working mode is a direct push to main, and admin enforcement would lock the sole maintainer out of his own repository. The same reading applies to WP-IN13’s “green in CI and blocking” in the table above. Full write-up: the standing correction.

The ≥95% bar has not been passed — it has not been measured. An earlier revision of this note said it “was passed with margin … on the hand-labeled corpus,” and both halves of that were wrong. The 0.000% orphaned agents / 100% usage attribution figures come from the Phase-0 desktop probe, in which the parser scored its own output against transcripts nobody had labelled — a self-check, not an accuracy measurement, and PROVISIONAL for exactly that reason. The hand-labeled corpus does not exist: the five spike/corpus/sessions/<short>/LABEL-ME.md trees are still templates with their GROUND TRUTH sections blank, awaiting the owner. The exit gate is built and runs, and because it has nothing to score against it prints NOT CERTIFIED at n = 0 rather than a number it cannot compute — it requires n ≥ 52 hand-labelled edges with a one-sided Wilson lower bound clearing 95% (packages/test-fixtures/src/annotations/score.ts). Treat parser accuracy on real data as unmeasured, not as proven. Replay-on-startup is the ingest watcher’s first tick over the whole corpus; idempotent whole-session re-ingest replaces the byte-identical-substrate formulation (JSONL never lands in events_raw — see ingest & reconciliation).

Contrast with rivals

DESIGN.md §6 gives an “honest read of the state of the art” that draws the line between table-stakes and the moat precisely:

  Scope Edge origin Aggregation Verdict (DESIGN.md §6)
simple10 Session-scoped Event-derived, computed at render time via buildAgentTree()/layoutTree() (parent→child, orphan-reparenting, root synthesis) plus a dependency-free N-body force graph (physics.ts) Per-instance “Table-stakes… the model to match” for the session-scoped tree — validate against one real subagent-heavy session before committing
hoangsonww Multi-session-looking, but really a post-hoc reconstruction Reconstructed post-hoc on SubagentStop; not a persisted edge table Type-aggregated (OrchestrationDAG.tsx is a 3–4 layer diagram of categories) “Oversold” — worth studying for its D3 Sankey/aggregate polish, not as a DAG implementation to copy
agenthropic Cross-session, global Persisted at ingest/projection time (orchestration_edges, dual-path via WP-IN8) Per-instance The moat — “unclaimed ground” (DESIGN.md §2 item 1)

Two things worth being precise about, since it is easy to blur them:

Confirmed shape of orchestration_edges (as built: migration 13 is the authority)

Beyond WP-D7’s Done-when and DESIGN.md §4, the canonical decision CD-4 (concept-analysis-v2.md) pins the column set explicitly:

“persisted orchestration_edges (self-ref parent_agent_id, instance/host_id, derived_from_event_id, idempotent)” (concept-analysis-v2.md, CD-4)

So the confirmed constraints on the table are:

The migration is now written, so the synthesis era is over. The table was introduced by migration 5, given endpoint indexes by migration 12, and rebuilt by migration 13 to widen the provenance CHECK. This is its current shape (apps/server/src/db/migrations.ts):

CREATE TABLE orchestration_edges (
  id              INTEGER PRIMARY KEY,
  session_id      TEXT NOT NULL,
  parent_agent_id TEXT NOT NULL,
  child_agent_id  TEXT NOT NULL,
  source          TEXT NOT NULL CHECK (source IN ('tool_use','directory','task_notification','queue_operation','legacy_explore')),
  instance        TEXT NOT NULL,
  host_id         TEXT NOT NULL,
  created_at      TEXT,
  UNIQUE (session_id, parent_agent_id, child_agent_id)
);
CREATE INDEX idx_orchestration_edges_session_id ON orchestration_edges(session_id);
CREATE INDEX idx_orchestration_edges_parent_agent_id ON orchestration_edges(parent_agent_id);
CREATE INDEX idx_orchestration_edges_child_agent_id ON orchestration_edges(child_agent_id);

Two later migrations are worth reading as decisions rather than as maintenance. Migration 12 added the two endpoint indexes because the global-DAG query walks edges by parent_agent_id and child_agent_id, and with only the session index present every such walk degraded to a full table scan; the table stays correct without them, which is exactly why the fix could wait — a dropped index costs speed, never truth. Migration 13 could not simply widen the CHECK, because SQLite cannot alter one: it creates the new table, copies every row, drops the old one, renames, and recreates all three indexes inside a single transaction. The cheaper alternative would have been to let the parser emit legacy_explore edges under a disguised tool_use label, and provenance honesty is precisely the property this CHECK exists to defend.

Where the built table departs from the CD-4 sketch, and why:

The fifth path, and why it is labelled rather than blended

Four of the five provenance values are structural: they name a fact the transcript states outright — a matched toolUseId, a directory containing a child, a task notification, a queue operation. The fifth, legacy_explore, is not. It exists for transcripts written before Claude Code 2.1.71, where a bare Explore sidecar carries neither a toolUseId nor a spawnDepth, so every structural anchor misses. The parser tries it only after all four structural paths have failed, and only when two independent conditions hold at once: the sidecar matches the bare legacy shape by key presence, and some other progress record names the child’s hex id as its own structural top-level agentId. Nested data.agentId values are deliberately not scanned — that direction drifts toward the substring joins the parser’s own gate #5 forbids.

The consequence is a genuine inference sitting in the same table as four observations, so it is given its own label rather than being blended into tool_use. Nothing in the read path collapses the two: a consumer that wants only observed edges can filter on source, and a reader looking at a rendered tree can ask where any particular edge came from.

Its status is implemented but not measured. The bare-Explore shape does not occur anywhere in the corpus available to this project, so the path is exercised by fixtures only, and its scope stays PROVISIONAL until a real pre-2.1.71 transcript ratifies it. That distinction also bounds the probe figures reported earlier on this page: they were produced by the four structural paths alone, with no legacy_explore edge anywhere in the set the probe ran over. It does not mean the ≥95% bar was cleared without legacy_explore — the bar has not been cleared at all, by any path, because it has never been scored against a labelled corpus.

Roadmap placement

The DAG moat is not a Phase 1 deliverable. DESIGN.md §9 places the moat extensions (ELK/Graphviz layout, fleet aggregation) at Phase 5+, but the moat’s core artifact — the persisted orchestration_edges table itself, dual-path derivation, and the rebuild-from-JSONL-alone proof — lands earlier, in Phase 3:

“3 — Projection, the DAG moat, reconciliation, cost (P0 blockers) | Pure normalizer + projection; dual-path orchestration_edges; reconciliation + backfill; replay-on-startup; watchdog ‘unknown’; CostEngine with compaction repricing + delegation-savings.” (development-plan.md §3)

Reading the two roadmap layers together:

As built: the Phase 3 core (persisted table, JSONL edge derivation, P0 proofs) and the Phase 4 read side (tree + global-DAG endpoints and the SPA views over them) both landed in the 2026-07 implementation wave. The Phase 5+ extensions remain unbuilt: no ELK/Graphviz layout, no fleet UI — the instance/host_id columns stay schema-ready, not shipped features.

First layout extension: ELK/Graphviz

Once the persisted graph exists and is queryable, its rendering can still improve independently of its correctness. DESIGN.md §6 names the specific next step, framed explicitly as future work, not something scoped into the current build:

“First extension when needed: ELK/Graphviz layout over the persisted tree.” (DESIGN.md §6)

The framing matters: this is a layout improvement (how the already-correct, already-persisted graph is laid out on screen — automatic layered/hierarchical positioning, the kind ELK or Graphviz specialize in) rather than a data-model change. It is explicitly deferred (“when needed”), with no work package or phase assignment yet in development-plan.md — there is nothing further to source about scope, timing, or which library wins until that decision is made.

What’s undecided

(What was open when this page was written; the as-built resolutions follow each item.)

See also