Use ChangeGraph
UI walkthrough, score breakdown, why-not-others, and Run investigation agent.
Investigate an incident
When production breaks, ChangeGraph builds an incident from Sentry (or fixtures), then ranks recent changes with a deterministic scorer. Engineers review evidence; the product never presents correlation as proven causation.
UI walkthrough
Open Incidents, then pick a scenario (A–D locally) or a live incident.
| Panel | What it shows |
|---|---|
| Header | Title, status, release, mapped SHA (nullable), LKG when known |
| Timeline | Releases, deploys, commits/PRs around the investigation window |
| Candidates | Ranked PRs/commits with score and rank |
| Score breakdown | Per-signal contributions (temporal, deployment, file overlap, service, historical) + penalties |
| Evidence | Grounded facts: overlapping paths, timing, release membership |
| Why not others | Counter-evidence for lower-ranked or distractor candidates |
| Explanation | Mastra explain_incident over packed context (ok / not_configured / error) |
| Investigation Agent | Mastra investigate_incident. Does not change scores |
FACT vs CORRELATION vs HYPOTHESIS vs UNKNOWN
- FACT: observed provider data (file in PR, stack path, deploy time)
- CORRELATION: scored relationship (“this PR overlaps the stack and landed in the release window”)
- HYPOTHESIS: Mastra / agent phrasing constrained to the packed evidence package
- UNKNOWN: missing mapping, empty candidate set, thin overlap. Valid outcomes
Never treat a top-ranked candidate as proven root cause.
Score breakdown (what to look for)
Scorer v2 (max 100) uses:
| Signal | Max | Meaning |
|---|---|---|
| Temporal | 25 | Closeness to incident time |
| Deployment | 25 | On mapped release / deploy membership |
| File / stack overlap | 25 | Intersection with stack frame paths |
| Service | 15 | Culprit / path-depth / service·transaction proxy |
| Historical | 10 | Same-release / prior deploy membership |
Penalties: no file overlap (−35), only test/docs (−10). See Scoring & LKG.
Why not others
Good investigations show why distractors lost:
- Closer in time but zero stack overlap (scenario B)
- Partial overlap only (scenario C)
- Same files but off the release line
- Docs/README-only changes
If the UI cannot explain “why not,” treat the ranking with more skepticism.
Run investigation agent
The Investigation Agent is the in-repo Mastra-compatible workflow investigate_incident (src/lib/agents/mastra-runtime), on by default:
INVESTIGATION_AGENT_ENABLED=1
Set INVESTIGATION_AGENT_ENABLED=0 to opt out. When enabled, incident detail shows Run investigation. Behavior:
- Read-only Mastra steps: load context → expand evidence (capped tools) → pack → explain → validate
- Does not mutate scores, candidates, or production
- Requires a configured LLM provider (
not_configureduntil you add a key) - Optional
FEATURE_CRITIC_LOOP(default off) may add HYPOTHESIS critique notes after investigateok. The investigate panel lists those notes whencritiqueis on the response, and showsnot_configuredwhen that is the critic status. Gate off hides the critic UI. It does not change ranks. Explain does not run critic. - Optional
FEATURE_DURABLE_SUPERVISOR(default off) may checkpoint/resume a single investigate run. Resume reloads ranks from the evidence package. Never auto-propose or git.
Boundaries and MCP relationship: MCP server.
Scoring window
Windows expand via resolveInvestigationWindow: 24h → 72h → 7d when the candidate set is thin. LKG further bounds “since healthy.” Defaults are correlation aids, not hard truth.
LLM provider
Configure a provider so explain_incident can run. Model HTTP is the ChangeGraph provider router (src/lib/llm/providers.ts); Agent.generate on the in-repo runtime throws. If no provider is set, the UI shows not_configured and asks you to add a key. Timeline, candidates, and evidence remain visible. See Agent context packing and Open source tools.