agenthropic

Telegram alerts

How to read this page. Unlike the rest of the usage section, nothing described here is built. The dashboard around it is: the server, the ingest pipeline, the persisted DAG, the cost engine and the four dashboard views all exist and run. The alerting core and the Telegram sink do not, and — as the box below explains — they are not on the schedule either. So read this as a design contract to check a future pull request against, not as a feature you can configure. Values marked (planned) or (leaning — unconfirmed) are the state of the design when it was written; the security invariants are binding and will not change.

Update — 2026-07 (as built): nothing on this page exists, and it is not on the path to existing. Read the whole page as a design record, not as a feature you can configure. Specifically:

The security rules on this page do hold today, because they are project-wide and not alerting-specific: the server dials nothing derived from an ingested payload, and secrets never reach SQLite, SSE or logs. They hold trivially, in the sense that the outbound-dispatch code path they constrain has never been written — the first rule holds because the server cannot dial anything.

This page covers the designed alerting core — the rule engine, the no-SSRF webhook dispatcher, the secret-handling contract, and the delivery guarantees — and its one concrete delivery adapter, Telegram, relaying to @baev_bot_bot. The key takeaway: alerts are a security-sensitive dispatch path by construction — the dispatcher only ever dials operator-configured targets (never a URL taken from an event payload, per security model §6), the bot token is held by reference and never touches SQLite, SSE, or logs, and a real triggering condition is designed to produce exactly one throttled notification, not a flood.

What it is, and why it’s a moat feature

Telegram alerting is one of the five features confirmed absent across all six audited rival dashboards — see the moat §2.3 and DESIGN §2, item 3: “Telegram alert sink → @baev_bot_bot. Graft hoangsonww’s formatTelegram webhook provider.” hoangsonww is the one audited project with a ready-made Telegram bridge — a formatTelegram webhook provider plus a full alert_rules / webhook_targets / webhook_deliveries schema — and because it is MIT-licensed with a real LICENSE file, that pattern is grafted with attribution, not clean-room reimplemented (the-moat §5; concept-analysis-v2 CD-9). It is deliberately isolated from the rest of that codebase, in particular its RCE spawner (/api/run) and its type-aggregated DAG cockpit — neither of those is grafted, ever.

Turning the dashboard from “something you have to look at” into “something that tells you when it needs attention” is the roadmap’s own framing of Phase 5’s goal (see the roadmap).

Architecture at a glance

        projection (agents, sessions, token_usage — Phase 3 output)
                             │
                             ▼
              ┌───────────────────────────┐
              │   Alert rules engine       │   WP-A5
              │  cost_threshold / stuck_   │   reads the projection only —
              │  agent / error             │   never raw events_raw payloads
              └─────────────┬──────────────┘
                             │ rule fires → alert_events row (dedupe_key)
                             ▼
              ┌───────────────────────────┐
              │   AlertSink port           │   WP-A1 (packages/shared,
              │   pure interface           │   no server/driver import)
              └─────────────┬──────────────┘
                             ▼
              ┌───────────────────────────┐
              │  Webhook dispatcher        │   WP-A4 — dials ONLY rows in
              │  + no-SSRF guard           │   webhook_targets (operator-
              └─────────────┬──────────────┘   configured); never a payload URL
                             ▼
              ┌───────────────────────────┐
              │  Telegram AlertSink        │   WP-A6 — token resolved via
              │  adapter                   │   token_ref (WP-A3), never inline
              └─────────────┬──────────────┘
                             ▼
                    @baev_bot_bot (Telegram)
                             │
                             ▼
              ┌───────────────────────────┐
              │  Delivery log              │   WP-A7 — retry/backoff, dedupe,
              │  webhook_deliveries        │   rate-limit → exactly one
              └───────────────────────────┘   throttled notification per condition

Every box above is a named work package in docs/analysis/development-plan.md’s Track A (WP-A1…WP-A10) — none of it is running code yet (see Current status).

The three rule kinds

WP-A5 (“Alert rules engine”) fixes the rule taxonomy at exactly three kinds, evaluated over the projection (Phase 3’s agents/sessions/token_usage tables), never directly over raw events_raw:

Kind (alert_rules.kind) What it observes Operator configuration Source
cost_threshold A session’s or agent’s accumulated dollar cost (ground-truth tokens × dated price, per the cost model) crossing an operator-set dollar limit A threshold amount, e.g. {"threshold_usd": 5.00} — set through the operator alerts API once WP-A8 ships WP-A5’s Done-when states this boundary explicitly: “cost_threshold fires exactly at the operator limit (boundary tested)”
stuck_agent An agent or session that has stopped making forward progress No documented config shape yet — flagged open below WP-A5’s Done-when names stuck agent as one of the three kinds; no further config schema is specified in any source in scope
error An agent or session that resolves to an error condition No documented config shape yet — flagged open below WP-A5’s Done-when names error as the third kind

Two of the three kinds are, honestly, thinly specified beyond their name — stated here rather than papered over, per the project’s own documentation style:

The no-SSRF dispatcher

This is the single most security-load-bearing piece of the alerting core, and it holds regardless of anything else on this page: the webhook dispatcher (WP-A4) only ever dials operator-configured targets. It never constructs an outbound URL from data that arrived inside an ingested event.

                    ┌──────────────────────────┐
   ALLOWED  ───────▶│ webhook_targets row       │──────▶ dial (Telegram, etc.)
                     │ (operator-configured,     │
                     │  created via WP-A8 API)   │
                     └──────────────────────────┘

                    ┌──────────────────────────┐
   NEVER   ─── ✗ ───│ a URL/host read out of an │        (disler's bug —
                     │ events_raw payload /      │         never built here)
                     │ hook event / JSONL line   │
                     └──────────────────────────┘

Secret handling: token_ref, never the secret

The Telegram bot token is a secret, and the design treats it as one at every layer — this is CD-10 in docs/analysis/concept-analysis-v2.md, stated verbatim: “Telegram token via token_ref → launchd env / chmod-600 (never in SQLite, never to the browser).”

See also configuration for how secrets fit into the broader config surface — that page is no longer a stub: it enumerates the environment variables the server actually reads, including the one secret that is live today, DASHBOARD_TOKEN, which is held to the same never-in-SQLite, never-to-the-browser rule token_ref is designed around.

Delivery guarantees: exactly one throttled notification

WP-A7 (“Delivery log with retry/backoff + dedupe & rate-limit”) fixes the outcome a real triggering condition must produce: exactly one throttled notification, not a flood of duplicate pings every time the underlying condition is re-observed. This is restated as a release-blocking exit gate for Phase 5 in docs/analysis/development-plan.md: “a real error/stuck condition yields exactly one throttled notification.”

The designed webhook_deliveries table (Phase 5, not yet built) carries the columns that make this mechanical, not aspirational:

CREATE TABLE alert_events (
  id           TEXT PRIMARY KEY,
  rule_id      TEXT NOT NULL REFERENCES alert_rules(id),
  session_id   TEXT,
  agent_id     TEXT,
  fired_at     TEXT NOT NULL,
  dedupe_key   TEXT NOT NULL   -- rate-limit/dedupe boundary, WP-A7
);

CREATE TABLE webhook_deliveries (
  id               TEXT PRIMARY KEY,
  target_id        TEXT NOT NULL REFERENCES webhook_targets(id),
  alert_event_id   TEXT NOT NULL REFERENCES alert_events(id),
  status           TEXT NOT NULL CHECK(status IN ('pending','sent','failed')),
  attempt_count    INTEGER NOT NULL DEFAULT 0,
  next_retry_at    TEXT,
  delivered_at     TEXT
);

The designed setup flow (planned)

No operator-facing setup flow exists yet — the operator alerts API (WP-A8) and alerts UI (WP-A9) are Phase 6 work, not yet built. The shape below is the designed sequence, marked planned throughout since exact endpoint paths, request/response shapes, and UI copy are not fixed by any source:

As built: steps 3 and 4 will not happen in this form. WP-A8 and WP-A9 were cut, so there is no API to register a target through and no UI to configure a rule in — the surviving intent is that configuration would live in a file. Read the numbered steps for the contract they fix (a target is operator-configured; the request carries a token_ref name, never a secret), not for the mechanism they name. Steps 1 and 2 are external to agenthropic and unaffected.

  1. Register a Telegram bot and obtain a chat id (planned; external to agenthropic) — done through Telegram’s own bot-registration flow (e.g. BotFather), entirely outside agenthropic’s own surface. agenthropic never talks to Telegram’s bot-management API on the operator’s behalf; it only ever sends messages once a target is configured.
  2. Store the bot token by reference, not by value (planned) — the operator places the token where WP-A3’s resolver expects it (a launchd environment variable, or a chmod 600 dotfile), never inside a request body or a database row.
  3. Register a webhook_targets row through the authenticated operator alerts API (planned, WP-A8) — auth-gated and timingSafeEqual-protected, same as the one write endpoint that exists today (POST /api/hooks/event); the cross-origin rejection applies to the SSE stream (see security model). The request supplies the token_ref name, never the secret value.
  4. Configure an alert_rules row: a rule kind, its config, and the target to fire through (planned, WP-A8) — e.g. a cost_threshold rule with a dollar limit, pointed at the webhook_targets row created in step 3.
  5. A real condition fires the rule, producing exactly one throttled delivery (§ Delivery guarantees) to @baev_bot_bot.

Once WP-A9’s alerts UI ships, steps 3–4 are expected to move from raw API calls to UI forms — but that UI does not exist yet, and no source specifies its exact fields beyond “rule config, target registration, delivery-log view.”

Work packages behind this page

For traceability, every claim above maps to one of these ten work packages in Track A of docs/analysis/development-plan.md (Phase 5–6):

WP Owner Delivers
WP-A1 backend AlertSink port + alert domain types (packages/shared) — pure interface, no server/driver import
WP-A2 data Alert & webhook schema migration (clean-room-safe, hoangsonww-attributed); forward-only, idempotent
WP-A3 security token_ref resolver (launchd env / chmod-600) + redaction + static gate; a >0600 dotfile is rejected
WP-A4 security Webhook dispatcher + no-SSRF guard; no code path reads a URL from a payload
WP-A5 backend Rules engine (cost threshold / stuck agent / error) over the projection
WP-A6 backend Telegram AlertSink adapter (hoangsonww-attributed), token via token_ref, per AlertKind
WP-A7 backend Delivery log + retry/backoff + dedupe & rate-limit → exactly one throttled notification
WP-A8 backend Operator alerts API — auth-gated CRUD for rules + targets
WP-A9 frontend Alerts UI — rule config, target registration, delivery-log view; token_ref name only
WP-A10 qa Alerts negative-test corpus + coverage hardening — SSRF and secret-leak tests

Current status

The paragraph that stood here described a pre-Phase-0 project waiting on a feasibility verdict. That is no longer where things are, and the change is worth stating precisely, because it moves alerting further from being built rather than closer.

The Phase-0 spike returned CONDITIONAL GO, and the prerequisite phases this page said had not landed have landed: the security spine, the ingest substrate, the persisted projection and the read API all exist and run. Every dependency alerting was waiting on is therefore satisfied. What changed at the same time is the schedule those dependencies were supposed to feed. v1.0 is defined as the persisted cross-session DAG plus dollar-accurate cost, explicitly “no alerts”, with a hard date of 2026-12-01; alerting was moved wholesale into v2.0, and v2.0 is entered only through kill checkpoint KC-5 — fourteen consecutive days of the author actually using v1.0 daily, plus at least three dated friction-log entries wishing for a notification. Nothing schedules it; only evidence admits it. Two of the ten work packages below, WP-A8 (operator alerts API) and WP-A9 (alerts UI), were cut outright and stay cut even if KC-5 passes, on the reasoning that a single operator does not need CRUD screens for himself.

So the old “longest tail of the whole graph” framing has inverted. Alerting is not the last thing blocking a release; it is deliberately outside the release, and the honest current status is that it may never be built at all — and the roadmap counts that outcome as a success, since it would mean the project declined to build a notification system for a dashboard nobody opens. Everything above this section remains a binding design commitment to check a future pull request against, and none of it is a description of running code.

See also