agenthropic

ADR-0010: CD-8 — Phase 0 is a throwaway GO/NO-GO feasibility spike

Empirical update — 2026-07-04 desktop probe

A read-only 8-agent probe of the real ~/.claude/projects/ corpus (phase0-probe.md) pre-answers CD-1 as CONDITIONAL-GO → build (confidence 85/100). This de-risks — but does not replace — this spike: the formal Phase-0 / WP-S1…WP-S7 still confirm it on the paired-capture corpus with Ivan’s tree sign-off, and the WP-S7 GO gate below still stands (no production code before it). Three corrections apply to the acceptance criteria as written:

As-built update — 2026-07-30

Verdict: overridden. Not passed.

This is the one ADR in this directory whose gate was bypassed rather than cleared, and it must be read that way. The WP-S7 verdict was CONDITIONAL-GO, meaning conditions attached. Implementation began on 2026-07-11, while those conditions were still open, by an explicit instruction from the project owner (Ivan Baev) in chat, given after he was told that the outstanding CD-8 conditions were precisely what was blocking the work. Two further dispatch overrides followed (2026-07-18, 2026-07-29).

The distinction matters more here than anywhere else in this corpus. This ADR exists because both external reports reached their conclusions by default, without running the experiment. An override that gets recorded as a pass would reproduce exactly that failure mode one level up: a gate that was skipped, remembered as a gate that was met. So, stated without softening:

What the built system does independently evidence. Three merge-blocking P0 proofs now assert, against fixture corpora, what G0.1 and G0.4 set out to measure: the DAG rebuilds from JSONL alone after a simulated outage; a double replay is byte-identical; Σ token_usage matches an independently written in-test reader. G0.2 is answered by construction — the shipped installer wires four hooks (UserPromptSubmit, Stop, SubagentStop, PreCompact), and SubagentStart does not exist, confirming the §4.2 suspicion this ADR was written to test.

That is real evidence and it points the same way the probe did. It is still not the thing this ADR asked for. The spike was meant to produce a ratified verdict on a paired-capture corpus with the owner’s tree sign-off, before the schema was poured. The schema is poured. The sign-off is outstanding. The honest summary is: the architecture was validated after being committed to, by tests the same author wrote, on fixtures rather than a hand-labelled corpus — which is better than EXPANDED’s “validate at the Phase-4 UI walkthrough,” and worse than what CD-8 specified.

As-built update — 2026-08-15

Two things to record, and neither of them improves this ADR’s standing.

First, a narrowing. “Three merge-blocking P0 proofs” should be read as “three CI-failing P0 proofs, merge-blocking for everyone but the owner”. They run on every push and pull request and fail the run; since 2026-08-25 main is branch-protected on the ci check, so the failure withholds a contributor’s merge, while the owner is exempt by design (enforce_admins: false). See the standing correction.

(As built — 2026-09-22: apps/server/test/p0/ holds four P0 proofs; the fourth, p0-five-daily-questions.test.ts (2026-08-07), proves the CD-10 five-questions exit gate over real HTTP through the booted server and is not one of the three G0.1/G0.4 proofs counted above.)

Second, and more to the point for an ADR about a gate that was overridden: no part of the override has been retired. WP-S7 has still not run. WP-S3 / G0.1b has still not been run as a formal probe. No hand-labelled corpus exists, so the ≥95% hierarchy criterion still has nothing to measure against and the exit gate still refuses to certify rather than reporting a number it cannot defend. Every Phase-0 figure remains PROVISIONAL and LABEL-ME.

The evidence base has grown since 2026-07-30 — the coverage bar is now 100 across five packages (ADR-0009), the P0 proofs still hold, and six further migrations have landed without disturbing them (eleven by 2026-09-19 — eighteen migrations in all, the P0 proofs still holding). But it has grown in the same direction it already pointed: more tests, written by the same author, against fixtures rather than the ratified paired-capture corpus this ADR asked for. That is worth having and it is not what was specified. An override does not become a pass by ageing well.

Context

Two load-bearing assumptions are unverified on paper: whether JSONL carries the subagent linkage well enough to rebuild the DAG (ADR-0001, LB1), and whether the “twelve” Claude Code lifecycle hooks assumed by both external reports actually all fire — in particular whether SubagentStart exists at all (concept-analysis-v2.md §4.2: “SubagentStart is probably not a real hook”). EXPANDED’s own version of a Phase 0 degrades this into paperwork, validating the linkage only at a Phase-4 UI walkthrough — by which point the normalizer and schema are already built around the unverified premise.

Decision

Phase 0 is a throwaway GO/NO-GO feasibility spike with a hard ❌ stop:

No production code until green.

Acceptance criteria

From concept-analysis-v2.md §7 (“Open questions → Phase-0 inputs”) and development-plan.md Track S:

Consequences

Alternatives considered