ChangeGraphChangeGraph

Local development

Incident pipeline, Tier-1 path, layer stack, MCP boundaries, and stack.

Architecture overview

This docs site (apps/docs) is product how-to documentation. The product app (incident dashboard) is separate. Run it with Docker and pnpm alongside these docs.

Visual walkthrough: Architecture · per-tool tabs: Tools reference · OSS runtimes: Open source tools.

Incident pipeline

Incident → Ingest → Correlate → Candidates → Pack → Mastra → UI
   │          │          │            │              │            │
 Sentry    GitHub/    LKG +        Scorer v2    explain_incident Timeline
 issue     Sentry/    windows      ranks PRs    investigate_     + evidence
           Vercel                  /commits     incident         + explain
  1. Incident: Sentry issue/event opens a time window around the spike.
  2. Ingest: Normalize providers into tenant-scoped entities (HMAC, idempotent delivery claims, natural keys). See Where the running app reads data.
  3. Correlate: Map release → commit when known; resolve last-known-good; expand 24h → 72h → 7d.
  4. Candidates: Deterministic scorer ranks changes with fixed weights (no LLM in the loop).
  5. Pack: Redact secrets, then pack a budgeted compact digest (frozen ranks). See Agent context. Optional Headroom compress after pack; the packer works without it.
  6. Runtime: in-repo Mastra-compatible explain_incident / investigate_incident (src/lib/agents/mastra-runtime, package @mastra/core) plus gated propose_incident / critic_incident / durable investigate resume. Agent.generate throws; model HTTP is src/lib/llm/providers.ts. AgentConfig.model is a string or function. Read-only capped tools; never mutates ranks. Requires a provider (not_configured when unset).
  7. UI: Timeline, scores, evidence package; packed Mastra explanation.

Tier-1 production path

GitHub (PR/commit/files)
        │
        ▼
Vercel deploy (SHA + project)
        │
        ▼
Sentry error (issue/event + stack)
        │
        ▼
   ChangeGraph (graph + scorer + UI)

Adapters only. Scoring stays provider-agnostic. Roadmap Tier-2/3 (Datadog, Slack/Jira, K8s, cloud) plug in as evidence sources, not core domain types.

Deterministic vs Mastra vs LLM

LayerRoleMutates scores?
Deterministic coreIngestion, scorer v2, LKG, evidence packageN/A (owns ranks)
Context packerRedact → budgeted compact JSON (frozen ranks); optional Headroom after packNo
investigate_incidentMulti-step read-only workflow on the in-repo Mastra-compatible runtime (src/lib/agents/mastra-runtime; on by default; INVESTIGATION_AGENT_ENABLED=0 to opt out). Model HTTP is src/lib/llm/providers.ts.No
explain_incidentOne-shot explain of the packed package on that same runtime (OpenAI / Anthropic / xAI / Groq / Fireworks via the provider router)No
Mastra critic_incidentOptional second-pass HYPOTHESIS notes after investigate ok (FEATURE_CRITIC_LOOP, default off)No
Mastra durable supervisorOptional single-incident investigate checkpoint/resume (FEATURE_DURABLE_SUPERVISOR, default off)No

The same packed evidence package feeds the UI, MCP tools, and the in-repo runtime. Deterministic scoring always runs. Investigation needs a configured LLM provider and fails closed as not_configured.

Headroom

Optional post-pack compress after redaction. The packer works without it. If the proxy is down, ChangeGraph uses the packed prompt (fallback: true). Packer + explain/investigate work without Headroom. Env: HEADROOM_URL (alias HEADROOM_BASE_URL), HEADROOM_API_KEY, HEADROOM_TIMEOUT_MS. Upstream: github.com/headroomlabs-ai/headroom. Catalog: Open source tools.

Bidirectional MCP

DirectionStatusPurpose
MCP server (shipped)External agents → ChangeGraphlist_incidents, get_incident, get_evidence_package (stdio, read-only)
MCP client (gated)ChangeGraph → external toolsRead adapter (FEATURE_MCP_CLIENT, default off). Fails closed as not_configured without a registry. The full multi-server registry is deferred. Scoring stays MCP-agnostic

Auth: CHANGEGRAPH_MCP_TOKEN (alias MCP_AUTH_TOKEN) when exposing beyond localhost.

Separation of concerns

LayerResponsibility
IngestionWebhooks/API, HMAC, idempotency, cursors, natural keys
CorrelationDeterministic scoring + LKG + window ladder
PresentationSetup / Incidents / Detail UI (product app)
ExplainMastra explain_incident over packed packages
AgentsMCP server (read); Mastra investigate_incident (capped tools, non-mutating)
Docs siteMarketing + how-to (apps/docs on Vercel)

Stack (MVP)

Next.js App Router + TypeScript · pnpm · Drizzle + Postgres schema (where the running app reads data) · Zod · Vitest · Tailwind CSS (shadcn-style primitives, no shadcn package) · Biome · in-repo Mastra-compatible runtime (src/lib/agents/mastra-runtime). Catalog: Open source tools.

Multi-tenant rule

All domain tables carry tenant_id. No cross-tenant queries in MVP. Secrets in integration_secrets ciphertext only.

API sketch

  • Webhooks: POST /api/webhooks/github|sentry|vercel
  • Product API: /api/v1/… (incidents, setup/backfill, investigate)
  • MCP: stdio tools list_incidents, get_incident, get_evidence_package

Out of MVP

K8s / cloud inventory, Datadog, Jira/Slack, feature flags as core types, AST/code graph, auto-remediation, autonomous prod changes.

Engineering docs

Contributor-facing markdown under repo /docs:

  • docs/SPEC.md
  • docs/MASTRA_AGENTS.md
  • docs/AGENT_CONTEXT.md
  • docs/INTEGRATIONS_OSS.md
  • docs/ARCHITECTURE.md
  • docs/SCENARIOS.md
  • docs/SECURITY.md
  • docs/ALIGNMENT.md
  • docs/REPLAY.md
  • docs/AUTH.md

This docs app is the product-facing experience; it does not replace those files.