How to read this page. The four views described below are built and shipped in
apps/web. The page keeps the design-era text it was first written as — a description of intended behaviour derived from the design basis (docs/ai/DESIGN.md) and the build plan (docs/analysis/development-plan.md), written before any application code existed — and amends it in place rather than replacing it. So a value marked (planned) or (leaning — unconfirmed) records what was still open when the plan was written, not what is open today; where the built thing settled the question, an As built box under the paragraph says how. The security invariants were binding then and are binding now.
Update — 2026-07 (as built). All four views described below are built and shipped in
apps/web. What actually runs: a React SPA behind a hand-rolled hash router (#/live,#/sessions,#/dag,#/cost—apps/web/src/router.ts) whose shell renders nothing until a token is present, held insessionStorageonly (apps/web/src/token.ts). The four views areLiveView(WP-U6),SessionsView(WP-U7),DagView(WP-U8) andCostView(WP-U9). Corrections to the text below, each verified against the repository:
agents.statushas five values, not three and not four:working | waiting | completed | error | unknown(apps/web/src/views/status.ts). The enumeration mismatch §1 calls unresolved schema work is resolved. Anullpersisted status is a separate state rendered asunrecorded— it is deliberately not folded intounknown, becauseunknownis a state the watchdog actively assigns whilenullmeans nothing was ever recorded.- The routes are real and prefixed
/api:GET /api/sessions/:id/tree,GET /api/dag/global,GET /api/cost/summary. They are no longer illustrative naming, and the API reference is no longer a stub.- Only the cost view uses D3.
CostView’s sankey layout comes fromd3-sankey(apps/web/src/views/layout/cost-flow.ts); the session tree and global DAG are drawn as hand-rolled SVG over a pure layered-layout module (apps/web/src/views/layout/layered.ts) — no force graph was built.- An unpriced model is not “a red build.” At runtime the per-session cost-analysis endpoint raises
PricingError→ HTTP 422, and the DB rollups surface the affected tokens as a separateunpricedTokensfigure. Nothing is ever silently costed at $0.model_pricinghas noverified_oncolumn. Its primary key is(model, bucket, effective_from)(migration 7,apps/server/src/db/migrations.ts), and the seeded rates are explicitly PROVISIONAL / awaiting ratification.- The UI honesty rules are implemented, not aspirational — always-rendered status buckets including zero counts, a permanent observed/inferred edge legend, declared-not-drawn dropped edges, a truncation banner with real numbers, and an unpriced-token KPI. They are listed per view below.
This page documents the four views the dashboard’s browser SPA is designed to render,
the five daily questions each one exists to answer, and how the SPA authenticates and
stays live. The key takeaway: every view is a read-only projection over data that
already exists in SQLite before the browser ever opens — the session tree and the
global DAG are queries over the persisted orchestration_edges table, not anything
reconstructed in JavaScript from an event stream (per DESIGN.md §6 and WP-U3/WP-U8’s
own Done-when clauses) — and the SPA itself renders nothing until it has proven
possession of the mandatory dashboard token (WP-U5).
The dashboard’s entire scope is driven by five concrete questions, not feature-parity
envy with the audited rivals. They are recorded verbatim as decision D3 in
docs/analysis/implementation-plan.md (“v1 daily-questions — the MVP definition of
value”), and every downstream work package that builds a view cites one or more of them
by number:
| # | Daily question (verbatim, D3) | Answered by |
|---|---|---|
| Q1 | What is the subagent tree of this session, and which branch is still running? | Session subagent tree (§2) for the tree shape; live status (§1) for “still running” |
| Q2 | Which agent/subagent burned the most tokens (and roughly what did it cost)? | Cost / Sankey / delegation-savings (§4) |
| Q3 | Did any session get stuck / error without me noticing? | Live status board (§1) |
| Q4 | What did today/this week cost, and how much did Haiku/Sonnet routing save? | Cost / Sankey / delegation-savings (§4) |
| Q5 | Show me last night’s sessions — persisted, after a restart. | Live status board’s session list (§1), backed by SQLite WAL persistence |
The backend read endpoints that serve these are grouped by the same numbering in
development-plan.md: WP-U3 (“Session/agent/subagent-tree endpoints — daily
Q1/Q3/Q5”) and WP-U4 (“Cost, delegation & global-DAG endpoints — daily Q2/Q4”).
Note that WP-U4 bundles the global-DAG read endpoint in with the cost/delegation ones
even though the global DAG (§3 below) isn’t itself one of the five named questions — it
is the moat feature that cuts across all of them, not an answer to a single one (see
the moat). This grouping is the development plan’s own, quoted
as-is rather than re-derived.
Phase 4’s exit gate (docs/analysis/development-plan.md, Phase 4 row; restated in the
roadmap)
is explicit that shipping means all five are answerable from the UI, and that a
new session’s story is understandable in under 30 seconds — the numeric bar every
view below is designed against, not a vague “should feel snappy.”
That second half of the gate is UNMEASURED. All five questions are answerable from the shipped UI, and the per-session cost analysis (§4) was built precisely because Q2/Q4 were answerable from the server and not from the browser. But nobody has yet sat a reader in front of a session they had not seen before and timed them, so “under 30 seconds” remains the target it was written as. Treat it as a design intention this page is written against, not as a result anyone can quote back.
SQLite (WAL) — projected tables, written once at ingest/projection
│
┌─────────────────────────────────────────────┐
│ orchestration_edges agents token_usage │
│ (moat artifact) (tree) (ground-truth) │
└───────────────┬───────────────┬──────────────┘
│ query, never render-time │ query
│ reconstruction (WP-U3, WP-U8) │ (WP-U4)
▼ ▼
Read API (loopback, token-gated, WP-U2) ── + ── RealtimeHub (SSE, WP-U1)
│ │
└───────────────┬────────────────┘
▼
React/Vite SPA (WP-U5) — loads only behind the token gate
┌───────────────┬───────────────┬───────────────┬─────────────────────┐
│ (a) Live │ (b) Session │ (c) Global │ (d) Cost / Sankey / │
│ status board │ subagent tree │ persistent │ delegation-savings │
│ (WP-U6) │ (WP-U7) │ DAG (WP-U8) │ (WP-U9) │
└───────────────┴───────────────┴───────────────┴─────────────────────┘
This is the same hook-ingest → SQLite (WAL) → SSE → browser SPA loop described on
the architecture overview; the dashboard is the read
side of that pipeline, not a second data path.
Answers: Q3 (“did any session get stuck / error without me noticing?”) and, alongside §2, the “which branch is still running” half of Q1; its session list also backs Q5 (“show me last night’s sessions”).
What it is. A board of every agent/subagent, each showing one of a small set of
states, built to be read at a glance rather than studied. WP-U6’s Done-when names the
target directly:
Live status view (working/unknown/done) — the <30s at-a-glance. A newly-stuck agent flips to “unknown” live via SSE within the window. (
development-plan.md,WP-U6)
What backs it. The unknown state is not a UI label invented for the board — it is
a real, persisted status produced by the missing-SubagentStop watchdog (WP-IN12):
“a missing SubagentStop → ‘unknown’ within the window, never a permanent
‘working’ forever” (see the DAG moat).
The board reads that same agents.status column the tree and DAG views read — it does
not maintain a separate notion of “stuck.” One naming note worth flagging rather than
smoothing over: DESIGN.md §4’s original schema sketch for agents.status enumerates
working/waiting/completed/error, with no unknown value — the three-state
working/unknown/done framing is the later, more specific one named by WP-U6 and
WP-IN12 together, and this page follows the newer, work-package-level framing as the
one actually built to. The exact reconciliation between the two enumerations is schema
work not yet written (owned by WP-D4). (As built: it is written. See the box below.)
As built: the reconciliation landed as a five-value enumeration —
working | waiting | completed | error | unknown— i.e.DESIGN.md§4’s four original values plus the watchdog’sunknown, rather than either sketch winning.apps/web/src/views/status.tsis the single place that fixes the order and the glyphs (● working,◌ waiting,✓ done,✕ error,▲ unknown).One distinction the design docs never drew, and the code insists on: a row whose
statusisnullis notunknown. It renders as its own· unrecordedstate, becauseunknownis a claim the watchdog actively makes (“this should have stopped and didn’t”) whilenullonly means no status was ever written. Folding them together would fake certainty. On the API side, however, agents with anullstatus are counted into theunknownbucket of a session’s rollup — “an absent status IS unknown, and hiding it would fake certainty” (apps/server/src/api/queries.ts).There is a sixth thing the board can show, and it exists to survive a version skew: a status string this build of the SPA has never heard of.
statusMeta()renders it as? unrecognised (<the raw value>)rather than coercing it into the nearest known state or dropping the row. If a future server persists a word this UI predates, you will see the word — which is the only outcome that does not quietly misreport it.
How to read it. WP-U6 framed the board as three states read left to right as
urgency: a session/agent still executing is working; one whose expected completion
signal never arrived within the watchdog window is unknown — the state that should draw
your eye first, since it is exactly the “stuck without me noticing” case Q3 asks about;
one that reported a clean stop is done. That urgency ordering survived into the built
board, but the vocabulary grew: what actually renders is the five persisted statuses in
the fixed order working · waiting · done · error · unknown, plus unrecorded for a row
whose status was never written. Because the flip to unknown arrives over SSE (§5 below)
rather than on a page reload, a session that goes stale while the board is already open
updates in place.
As built:
LiveViewrenders all five buckets for every session, including the ones sitting at zero (they get a dimmedbucket-zeroclass rather than being filtered out) — a board that hides empty buckets would let a reader infer “no errors” from an absent error count. It fetchesGET /api/sessions(page size 50) and then stays live off three typed SSE frames:session-ingestedtriggers a refetch,agent-status-changedpatches the affected session’s counts in place and falls back to a refetch when the patch does not match anything it is holding, andingest-failedraises a dismissible banner naming the quarantined session and the failure reason. Heartbeats are SSE comment frames and never surface to the client. Where a session has tokens that could not be priced, the row shows~ n unpricedalongside its dollar figure rather than absorbing them into it.Who writes each state — and what you see with no hooks installed. Reading a transcript proves activity, never termination, so ingest only ever writes
working.completedcomes from theSubagentStophook,waitingfrom theStophook (which fires at the end of every turn, so it means “idle right now”, not “finished”), andunknownfrom the watchdog afterDASHBOARD_WATCHDOG_MINUTESof silence. If you have not installed the hooks (hooks/install.mjs), nothing on this board will ever readdone— agents moveworking→unknownand sit there. That is not a bug in the board; it is the board declining to claim an ending nobody observed. See the status lifecycle.
AMENDED 2026-09-23 (J-4). The sentence above - “the row shows ~ n unpriced” - described a
two-way rendering, and that was what the code did when it was written: the clause was gated on
unpricedTokens > 0, so a count that arrived non-finite (NaN, Infinity) failed that test in
exactly the same way a measured zero does, and the row printed its dollar figure with nothing
beside it. Five sites shared that gate. They now share one module instead
(apps/web/src/views/unpriced.tsx), used by LiveView, SessionsView and DagView, and the
rendering is three-way:
The served unpricedTokens |
What the row shows |
|---|---|
| a positive number | ~ n unpriced beside the dollar figure - unchanged |
exactly 0 |
nothing; a measured zero is the one value with nothing to disclose |
non-finite (NaN, Infinity) |
an explicit unpriced: tokens unreadable |
The third row is the one that did not exist before. dto-guards.ts validates a DTO’s shape,
not the sanity of its numbers - its own docblock says so - so a non-finite count reaches the
view intact; under the old test the dollar amount beside it looked complete precisely when
nobody could say whether it was. The same module supplies the hover text on tree and DAG nodes
(unpricedTitleSuffix), so the four remaining sites cannot drift apart from each other again.
A second reading is affected the same way. formatRelativeTime used to return null for a
timestamp it could not parse, which is what it also returns for a timestamp that was never
recorded - so an unparseable value rendered as an absence the server had not claimed. It now
returns timestamp unreadable (UNREADABLE_TIMESTAMP in apps/web/src/format.ts, alongside
UNREADABLE_TOKENS and UNREADABLE_USD), leaving null meaning exactly one thing: nothing
was recorded.
Answers: Q1 (the tree shape of “what is the subagent tree of this session”).
What it is. A D3-rendered tree/force graph of one session’s agents and subagents,
parent→child. WP-U7’s Done-when:
Session-scoped subagent tree view (D3 force+tree, live). From a real fixture the tree matches the labeled hierarchy (≥95%). (
development-plan.md,WP-U7)
What backs it. This is the point where the moat’s core discipline is most visible
in the UI: the tree is not walked from a raw event log by the browser at render
time. WP-U3’s backend endpoint states this explicitly — GET /sessions/:id/tree
(the work package’s own illustrative naming, not yet a built or confirmed route)
(As built: it is a real route, at GET /api/sessions/:id/tree.) is
“built from a query over orchestration_edges” (per
the DAG moat),
the same persisted table DESIGN.md §4 requires to be written once, at ingest/
projection time, rather than recomputed on every page load. The ≥95% accuracy bar is
the same figure Phase 3’s exit gate holds the underlying reconstruction to against a
hand-labeled real session, “even without the dedicated subagent-start signal”
(docs/site/guide/roadmap.md#phase-3--projection-the-dag-moat-reconciliation-cost) —
the tree view inherits that correctness bar rather than defining a separate, looser one
for display purposes.
The ≥95% bar has not been cleared, because it has not been run. Measuring it needs a hand-labelled corpus — the LABEL-ME task, at least 52 labelled agents across real sessions — and that corpus does not exist yet. The harness that would score against it is built and reports NOT CERTIFIED rather than a number, which is the correct output for “no ground truth to compare against” and not a failure. Everything the Phase-0 probe measured about hierarchy accuracy therefore stays PROVISIONAL until the owner labels the corpus and ratifies the result. Read the tree as the persisted edges it draws — each one carries its own provenance, see the box below — rather than as a shape certified to be 95% right.
How to read it. Root is the main agent; each edge is a persisted parent→child
orchestration_edges row, not an inferred nesting guess; each node’s status (see §1) is
what tells you which branch is still running versus finished versus stuck.
DESIGN.md §6 names simple10’s buildAgentTree()/layoutTree() (parent→child,
orphan-reparenting, root synthesis) plus its dependency-free N-body force graph as the
rendering model to match for this view specifically — table-stakes to hit, not the
moat itself (the moat is §3, next).
As built:
SessionsViewrenders hand-written SVG over a deterministic layered layout (apps/web/src/views/layout/layered.ts), not a force simulation — the force-graph model above was studied and not adopted. What the view adds beyond the design text is edge provenance and three honesty affordances:
- Every edge is drawn solid when observed (
source = tool_use) and dashed when inferred (directory,task_notification,queue_operation,legacy_explore), with the provenance also in each edge’s<title>and a legend that is permanently on screen rather than behind a hover or a toggle. The legend names all five sources by their stored word, so an edge you can see is an edge whose evidence you can name — the four dashed sources are four different strengths of guess, not one undifferentiated “inferred”.legacy_exploreis the weakest of them and is kept separate for exactly that reason (see the API reference).- Edges whose endpoint is not in the payload are counted and declared in text (“reference agents outside this payload and are not drawn”), never drawn to a node that isn’t there and never silently dropped.
- Cyclic nodes are likewise declared rather than laid out as if acyclic, and the
unattributedbucket is always rendered — including when it is empty.
Two further things the view renders that the text above never described, because neither existed when it was written:
SessionsView builds an “observed agent outcomes” list from tree.agents (the
payload as served, not the layout, so an agent the layout could not place still has its
outcome said) and names one row per agent whose outcomeCause is non-null. The list is read
from the same field the node’s hover <title> carries, which matters because the SVG is
role="img": a cause that lived only in a <title> would be a fact the picture knows and a
screen-reader user is never told. Four renderings are kept apart deliberately
(outcomeCauseText, apps/web/src/views/SessionsView.tsx): null renders nothing at
all - never “ok”, never “succeeded”, never a dash a reader could take for a zero, because
no observed outcome is not a claim of success; a known cause renders verbatim, because a
friendlier paraphrase would have to decide whether a refused spawn is a failure and it is
not; a cause word this build does not know renders as unrecognised (<raw value>); and a
field that is not a string at all renders as unrecognised (<absent>), since “no cause word
was sent” and “no outcome was observed” are different facts and must not be spelled the same
way. The cause is shown for any agent that carries one, not only for status: error -
on the measured corpus most observed causes are concurrency_limit (a spawn that was
refused, so the agent never ran) or user_interrupt (a person stopped it), and filtering to
error would have hidden them while promoting them would have invented failures.~ n unpriced when positive, silence at a measured zero, and
unpriced: tokens unreadable when the count arrives non-finite. A withdrawn clause here
would say the unattributed dollars are the whole of it.Answers: no single one of the five questions by name — it is the cross-cutting moat feature, not a Q-numbered item (see the mapping caveat above) — but it is what makes “which branch across all my recent sessions is doing the work” answerable at all, something none of the five audited rivals can do (per the moat).
What it is. The same kind of parent→child graph as §2, except it spans every
session for this instance, not just one. WP-U8’s Done-when:
Global persistent per-instance orchestration DAG view (the moat). Spans multiple sessions, sourced from a query over persisted edges. (
development-plan.md,WP-U8)
What backs it. This is precisely the feature the DAG moat
page exists to explain in depth: orchestration_edges rows carry a non-null
instance/host_id column from the very first migration (WP-D7’s Done-when; CD-4 in
concept-analysis-v2.md), which is what makes a query spanning multiple sessions
possible without a schema change — the same table §2’s session-scoped tree reads, just
queried without a session_id filter. No other audited project has this: simple10’s
tree is session-scoped and event-derived at render time; hoangsonww’s
“DAG cockpit” (OrchestrationDAG.tsx) is a type-aggregated 3–4-layer diagram
reconstructed post-hoc on SubagentStop, not a per-instance persisted graph
(DESIGN.md §6). agenthropic’s version is per-instance (one node per actual agent run,
never collapsed into a “type” category) and persisted (written once, queried many
times) — see the DAG moat’s contrast table
for the full rival-by-rival comparison.
Empirical footing. The persisted, per-instance DAG this view queries is not a speculative design bet. The Phase-0 corpus probe read the real
~/.claude/projects/corpus and empirically pre-answered CD-1 asCONDITIONAL-GO(confidence 85): JSONL is a trustworthy, outage-surviving single source of truth for the persisted subagent DAG — provided the parser keys on theAgent/Workflowspawn tools (notTask), walks both on-disk layouts (85% of agent files are nested), and indexes subagents as parents. That de-risks, but does not replace, the formal Phase-0 spike.The last clause of that paragraph used to read “no production code ships before the
WP-S7GO gate,” and it is no longer what happened: implementation began on 2026-07-11 by an explicit owner override of that condition, recorded as such. The override moved the build; it did not ratify the numbers.CONDITIONAL-GOandconfidence 85are still the probe’s own PROVISIONAL figures, unsigned, and the hierarchy-accuracy bar above them is still uncertified for want of a labelled corpus.
How to read it. Same visual grammar as §2 (nodes = agent instances, edges =
persisted parent→child relationships, node status = working/unknown/done from §1), but
the frame is “everything this Mac Mini has run,” not one session. DESIGN.md §6 names
the first planned layout improvement for this view once it exists — “ELK/Graphviz
layout over the persisted tree” — as future work with no committed timing, not
something scoped into the current view (WP-U8 builds the query and the render; the
layout upgrade is separately deferred).
As built:
DagViewreadsGET /api/dag/globaland sharesSessionsView’s layered layout, its five-source solid/dashed provenance legend and its declared-not-drawn dropped edges. The one thing unique to it is truncation honesty: the view asks for 1000 nodes — the endpoint’s own default, with a hard ceiling of 5000 (both PROVISIONAL constants inpackages/shared/src/schemas/common.ts, not ratified limits) — and when the cap bites the view shows a banner with the real figures — “Truncated: showing n of N agents and m of M edges (node limit 1000)” — instead of presenting a partial graph as the whole picture. The slice the server returns is the most recently active agents, so a truncated graph is a recency window and the banner says so rather than letting it read as “everything there is.” The ELK/Graphviz layout upgrade is still not built.
Answers: Q2 (which agent/subagent burned the most tokens/cost) and Q4 (today/this week’s cost and Haiku/Sonnet routing savings).
What it is. A Sankey-style flow of token spend, plus a delegation-savings figure.
WP-U9’s Done-when:
Cost / Sankey / delegation-savings view (daily Q2/Q4). Every displayed dollar traces to ground-truth tokens × dated price. (
development-plan.md,WP-U9)
What backs it. Every number here is the read-side of the equation
the cost model describes in full: cost = Σ
(tokens_in_bucket × dated_rate_for_bucket), computed server-side by CostEngine from
token_usage (ground-truth, copied verbatim from ~/.claude/projects/*.jsonl, never
inferred) and model_pricing (a dated, versioned table — effective_from/
verified_on (As built: the shipped table has effective_from only; no
verified_on column was ever created.) — never a single hardcoded constant).
Delegation-savings is the same
equation run twice per row and diffed: Σ max(0, top-tier-equiv − actual) against
whatever model actually ran (see
cost model §7).
A model observed in a fixture with no priced row is a red build, never a silent
“estimated” label (WP-C6, the staleness-fails-CI gate) (As built: the enforcement
is at runtime, not only in CI — see the box below.) — so a dollar figure this view
shows is never a guess, and the delegation-savings metric is explicitly tied to the
model-routing decision it’s meant to inform, not displayed as a decoration
(WP-C5; see cost model §7
on the named vanity-metric risk and its mitigation).
How to read it. Sankey flow width is proportional to token volume, colored/grouped
by model or bucket (speed/inference_geo/service_tier — DESIGN §4); the
delegation-savings figure sits alongside it as “what routing to a cheaper model already
saved you,” re-priced against the top-tier rate that would otherwise have applied.
DESIGN.md §6 names hoangsonww’s D3 Sankey as the rendering technique worth studying
directly — its aggregation shortcut (type-aggregated categories instead of a
per-instance graph) is what agenthropic does not copy for §3’s DAG, but the Sankey
rendering idea itself is fair game for this cost view.
As built:
CostViewreadsGET /api/cost/summary(default top-5 sessions) and lays out amodel → all cost → sessionsankey withd3-sankey. Three things differ from, or are more specific than, the text above:
- Unpriced tokens are their own KPI, captioned “no price row matched — not counted in $”, and an
Unpricedcolumn appears in the per-model, per-day and top-session tables. That column renders~ nor a plain0— never a blank cell, which a reader could mistake for “none.”- A model with usage but a $0 price is listed in text, under “Not in the flow (usage but $0 priced)”, rather than drawn as a zero-width flow that would be invisible. The same applies to the “other sessions” remainder outside the top-N.
- Two different failure modes, not one. The DB rollups behind this view never halt: unpriced tokens contribute $0 to the total and are surfaced separately. The per-session
GET /api/sessions/:id/cost-analysisendpoint does the opposite — it raisesPricingErrorand returns HTTP 422 rather than serve a figure it cannot justify, and its delegation-savings number carries an explicitisEstimate: true. The seeded rates themselves are still PROVISIONAL and unratified (apps/server/src/db/migrations.ts), so treat the dollar amounts as correctly-computed from numbers that have not yet been signed off.
AMENDED 2026-09-23 (J-4). The first bullet above says the Unpriced column “renders ~ n
or a plain 0 - never a blank cell, which a reader could mistake for ‘none.’” That is still
what the column does, and a non-finite count renders there as ~ tokens unreadable rather
than as a number. The tiles around it did not hold to the same rule, and now do:
| Where | Old rendering | Now |
|---|---|---|
| Total cost tile | the priced tokens only caption was gated on unpricedTokens > 0, so an unreadable count silently removed it while the Unpriced tokens tile two columns over printed tokens unreadable about the same field |
an unreadable count gets its own caption, “the unpriced-token count came back unreadable - what this leaves out is unknown” (kpi-total-unpriced-unknown) |
| Today (UTC) and Last 7 days (UTC) tiles | the same > 0 gate, and a second way to fail it: computeCostWindows sums whole perDay rows with +, so one row whose unpricedTokens arrived non-finite makes the whole window sum NaN, and NaN > 0 is false |
the note is three-way (UnpricedWindowNote) - ~ n unpriced when positive, silence at a measured zero, and an explicit unreadable caption otherwise (kpi-today-unpriced-unknown, kpi-today-partial-unpriced-unknown, kpi-week-unpriced-unknown) |
The total itself is untouched in each case: it is a real sum of real priced tokens. What changed is that nothing beside it now claims to know what it leaves out. The stale-tab arm of the Today tile gained the same note for its own reason - it printed cost and tokens with only a staleness caveat, and a reader told about one gap reasonably assumes there is no other.
AMENDED 2026-09-23 (J-4). The second bullet’s “the same applies to the ‘other sessions’
remainder outside the top-N” now has a fuller answer above the table, because the endpoint
serves two facts the view did not previously have: sessionCount (how many sessions carry any
usage at all - the population the slice was cut from) and hasMore (whether the slice is
truncated). They are separate facts and the scope line uses them separately:
hasMore true, the line reads “the n costliest of N sessions with recorded
usage”, and a remainder is attributed to the sessions the table does not list.hasMore false, it reads “every session with recorded usage is listed here - all
N”, and a remainder above USD_EPSILON is reported as unaccounted for, in those words:
there is no unlisted session to attribute it to, so the served total and the served rows
disagree and the page says so rather than inventing a bucket.listedSessionCount === 0, it says none of the N sessions reached
the table.The empty state is split along the same seam rather than being one sentence. “No sessions
recorded yet” is now only said when sessionCount is 0 and the totals report no usage;
when the totals do report usage over a zero session count, the empty state names the
contradiction and both served figures instead of denying the usage the KPIs above it are
rendering. That is the same disagreement the remainder sentence handles at the other end of
the table, and it gets the same treatment: state it, name both figures, attribute nothing.
Q4 asks what today and this week cost, and the view answers with two tiles above the
sankey: Today (UTC) and Last 7 days (UTC). The parenthetical is not decoration.
The server buckets usage by the UTC calendar date of the timestamp on the JSONL line, so
a tile labelled plainly “today” would mean UTC-today to the server and local-today to
whoever is reading it — the same figure, quietly meaning two different windows. Both
tiles therefore name their own boundary, print the exact dates they cover
(YYYY-MM-DD, and weekStart → today for the week), and are followed by a line saying
outright that these are UTC calendar days over the recorded usage timestamps, not your
local timezone.
The seven-day width is a PROVISIONAL constant (WEEK_WINDOW_DAYS), and the window is
inclusive of today, so it spans today − 6 … today. Two edges are handled deliberately
rather than swept up: usage dated after today is excluded from both windows, because a
future date means two machines disagree about the clock and folding it in would inflate a
window it does not belong to (the “today” boundary is the later of the page’s clock tick
and the moment the snapshot was read, so a read landing just after UTC midnight does not
file the new day’s rows as future-dated while the tick lags); and usage carrying no timestamp at all lands in the
per-day table’s literal unknown row, sits outside every window, and is disclosed in
that same note when it is non-zero. Neither is dropped, and neither is folded into a
window to make the tiles add up.
One staleness caveat, on record as review item M-10: the view has no clock tick and no SSE-driven refetch, so on a tab left open across UTC midnight the “today” boundary is only as fresh as the last render.
Amended 2026-09-25 (KK3): the clock does tick now, and the page states its own age: a “Read … ago” line above the tiles says every figure is as of that read, and a Refresh costs button re-reads the summary, the burners and the savings together (the summary’s error state has a Retry). There is still no SSE-driven refetch - the page never refreshes itself.
The rest of this section predates the ticking clock and still holds: the tile does not leave the reader to work out the consequence. When the page’s UTC day has moved past the day the snapshot was read on, the Today tile takes one of two arms instead of printing its usual figure:
not measured, names the date it was read on and points at the Refresh costs control. $0.00 would
be the worse outcome of the two - a number nobody measured, in the format reserved for
numbers somebody did;The Last 7 days tile is deliberately weaker in the same situation: six of its days are measured whatever happened to the seventh, so it prints its figure with a note rather than withholding it.
Q2 asks which agent burned the most. The sankey answers it in SVG <title> tooltips,
which keyboard and assistive-technology users cannot reach at all, so the view also
renders a plain ranked table (review item M-8). It reads a second endpoint,
GET /api/dag/global, and has its own loading/error state: a failure there degrades that
one section instead of taking the whole cost view down.
The ranking key is totalTokens, not costUsd — and that choice is the point.
totalTokens sums every usage row including the unpriced ones, so it is the true burn
even for an agent whose model has no price row; ranking by dollars would quietly demote
exactly the agents whose cost is unknown, which is the same class of lie as a silent
$0.00. Each row still shows the dollar figure and the unpriced gap alongside. Ties break
by cost and then by id, so the order is stable across refetches rather than reshuffling
equal rows as if data had moved.
Three scope facts are printed above the table rather than left for the reader to assume:
And when the DAG endpoint truncates, the table says what that does to the ranking: the server slices by recency, not by burn, so a truncated slice may not contain the biggest burner at all. The banner states that in words instead of presenting a partial ranking as a global one. The same holds when the answer is not marked truncated but the number of agents returned disagrees with the count the same read serves: a separate banner says the ranking covers a slice of unknown size. Whenever the list is partial, the scope line drops “All” and says “in the returned slice”.
Clicking a session in the Top sessions table opens the analysis panel underneath it —
the browser-side consumer of GET /api/sessions/:id/cost-analysis. It is opt-in per
session by design: that endpoint re-reads transcripts off disk, which is far too
expensive to fire for every row of a summary table. Until you pick one, the panel says so
plainly rather than rendering an empty frame.
It shows two things that must not be read the same way:
deltaUsd is not a saving — on a complete substrate it should be about
zero. The panel therefore prints the delta with an explicit sign and, at or above one
cent (DELTA_SIGNAL_USD, PROVISIONAL), calls it what it is: a discrepancy worth
looking at, a signal that the substrate or the pricing is incomplete. Sub-cent deltas
are rounding and are not dressed up as findings.isEstimate field to the
literal true so it cannot be switched off, because the cache profile of the run that
never happened is not observable. The panel renders it with a ~, an explicit
estimate badge and the hypothetical model named — never as a bare dollar amount
sitting flush with the measured figures above it. Subagents with no resolvable
top-tier model are excluded from the estimate rather than guessed at, and their
count is shown next to it.The panel’s failure text is likewise not interchangeable. A 503, a 404 and a 422 each imply a different action by the reader, so each keeps its own sentence, and the 503/422 cases quote the server’s own message instead of restating a single hard-coded cause — an earlier version told readers “this server has no corpus configured” about a machine whose corpus was merely missing. Naming the wrong cause is worse than naming none: it sends the reader to fix something that is not broken. The full list of answers that endpoint can give is on the API reference.
The dashboard is not a public page with a login screen bolted on — it renders nothing
without the token. WP-U5’s Done-when states this as the shell’s own contract:
React/Vite SPA shell + token auth + resilient SSE client. Loads only behind the token gate; no token → no data/stream. (
development-plan.md,WP-U5)
Two things follow directly from “no token → no data/stream”:
No unauthenticated read. Every view above is served by an endpoint WP-U2 (Read
API foundation) requires to be timing-safe auth-guarded — “every read route
auth-guarded (timing-safe),” stricter than treating reads as automatically safe (the
exact gap cast’s unauthenticated GETs left open — see
security model, rule 5).
The token itself is a DASHBOARD_TOKEN compared with Node’s timingSafeEqual, never
a plain ===, and the server refuses to start at all if it is unset — never a
no-op-when-unset fallback (hoangsonww’s mistake; see
security model, rule 2).
A sample env line always uses a placeholder, never a real value:
DASHBOARD_TOKEN=<token>
Realtime is SSE, not WebSocket, and it is “resilient.” The transport decision is
already settled (CD-5, concept-analysis-v2.md): a server→browser-only feed, so SSE
— which auto-reconnects and needs no same-origin WS handshake — was chosen over
WebSocket, and DESIGN.md §8’s older “WebSocket” wording is superseded by that later,
canonical decision (see architecture overview, transport).
WP-U1 (RealtimeHub SSE endpoint) is done-when “a cross-origin Origin on
/api/stream is rejected; no wildcard CORS” — the same-origin enforcement is a
contract test, not a convention (WP-F7). “Resilient” in WP-U5’s framing and the
Phase 4 “what ships” wording (“served over a resilient, reconnecting stream” —
docs/site/guide/roadmap.md#phase-4--read-api-the-dashboard-and-the-five-daily-questions)
means the client reconnects the SSE feed on its own after a drop; the exact reconnect
backoff mechanics are not specified anywhere in the source docs and are treated here
as (planned), not fixed.
As built: the token gate and the SSE client both exist, with two specifics the design text left open and one requirement that is not met as stated:
- The SPA holds the token in
sessionStorageonly — neverlocalStorage, never a cookie, never logged and never rendered (apps/web/src/token.ts). It travels as aBearerheader on API calls; the?token=query form is used only on/api/stream, becauseEventSourcecannot set headers, and the server’s log serializer strips it.- The client exposes four connection states —
connecting → open → reconnecting → closed— surfaced in the header chip. Malformed frames are dropped silently rather than crashing the view, and exactly three typed server frames exist:session-ingested,agent-status-changedandingest-failed(the shared list inpackages/shared/src/realtime/event-types.tsis authoritative for both ends).- “Resilient” resolved to the browser’s built-in
EventSourcereconnect, hinted by a server-sentretry:field — there is no custom backoff schedule, and there is no replay: a reconnect re-subscribes to the live feed, it does not redeliver frames missed while disconnected.WP-U5’s “resumable” wording is therefore not satisfied in the literal sense; the views compensate by refetching their snapshot on reconnect.
The full endpoint and stream reference — routes, payload shapes field by field, the health payload, and the reconnection semantics — lives on the API reference, which documents the twelve routes the server actually registers. When this page was first written that reference was still an unwritten Phase 4 deliverable, which is why several paragraphs above hedge about “illustrative naming”; it is written now, and it is the authority wherever the two pages differ.
The dashboard is never reachable at a routable address, from a phone, a second laptop,
or anywhere off the Mac Mini itself, except through an SSH local port-forward or a
Tailscale tunnel terminating at the loopback socket — the bind stays 127.0.0.1
regardless of which carrier you use, and the mandatory token check applies identically
whether the request originated at the Mac Mini’s own keyboard or arrived over either
tunnel (DESIGN.md §8; see remote access for the
step-by-step setup of both options, and security model for why
a reverse proxy to the open port is never an acceptable substitute).
This is a design-basis page, not a shipped-system page — stated plainly rather than glossed over. (As built: four of these five are now settled; each carries its resolution inline.)
GET /sessions/:id/tree and similar names above are WP-U3’s own illustrative
naming in its Done-when text, not a confirmed, frozen API surface) — the authoritative
reference is the API reference, a Phase 4 deliverable not yet written.
(As built: resolved. The routes are real, /api-prefixed, TypeBox-schema’d with
additionalProperties: false, and documented on the API reference.)agents.status enumeration mismatch between DESIGN.md §4’s original sketch
(working/waiting/completed/error) and the later working/unknown/done
framing WP-U6/WP-IN12 build the live status board to — see §1 above. The concrete
migration DDL that resolves this is WP-D4’s deliverable, not yet written.
(As built: resolved as the union — five values,
working | waiting | completed | error | unknown — plus a distinct null =
“unrecorded” state that is deliberately not merged into unknown.)WP-U5) are named as a
requirement, not yet specified in detail. (As built: resolved by choosing the
minimum — EventSource’s own reconnect plus a server retry: hint. No custom
backoff, and no replay of frames missed while disconnected.)DESIGN.md §6 names ELK/Graphviz as the first extension “when needed,” with
no committed timing or library choice. (As built: still open. The DAG ships on a
hand-written deterministic layered layout; ELK/Graphviz was not adopted and has no
committed timing.)apps/web throughout this page are leanings per
DESIGN.md §10, not locked decisions. (As built: locked. pnpm monorepo with
apps/server, apps/web, packages/shared, packages/core,
packages/test-fixtures, hooks/; React + Vite; D3 only via d3-sankey, and only
in the cost view.)orchestration_edges
mechanics behind the session tree (§2) and global DAG (§3) views.