ChangeGraphChangeGraph

Use ChangeGraph

Scorer v2 weights 25/25/25/15/10, penalties, windows, and LKG.

Scoring & last-known-good

Scores are CORRELATION only, never causation. Re-running on the same inputs must yield stable rankings (scorer_version: v2).

Weight table (max 100)

KeyWeightRole
temporal25Time proximity to the incident
deployment25Membership on the mapped release / deploy
file_overlap25Overlap between changed paths and stack frames
service15Culprit / path-depth / service·transaction proxy
historical10Same-release / prior deploy membership

Defined in src/lib/scoring/weights.ts and locked by unit tests.

Penalties

PenaltyValueWhen
no_file_overlap−35Candidate paths do not intersect stack
only_test_or_docs−10Changes look like tests/docs only

Penalties keep time-proximate distractors (README PRs) from winning on recency alone.

Investigation windows

Ladder (hours): 24 → 72 → 168 (7d).

  • Start narrow when LKG / release context is rich
  • Expand when the candidate set is empty or too thin
  • UI may surface windowExpanded / expansion reason

Last-known-good (LKG)

LKG bounds “what changed since things were healthy”:

  • Prefer a prior known-good deploy (lkg_sha, environment, deployed_at)
  • Prefer candidates between LKG and the incident over ancient history
  • When LKG is unknown, fall back to the configured time window

Treat LKG as a correlation aid, not proof of root cause. Fixtures include an LKG deploy for scenario A (good000lkg… in e2e expectations).

Evidence package

Only fields the scorer (and sanitizer / packer) allow are packaged for the UI and Mastra. Packed context: Agent context.

  • Overlapping paths
  • Timing / release membership
  • Counter-evidence
  • LKG snapshot fields when present

Secrets and raw tokens are stripped.

Interpreting ranks

  • Top-1 / Top-3 are decision supports for humans (pnpm replay reports accuracy)
  • Low scores or empty candidate lists are valid (scenarios B / D)
  • Never invent a culprit PR when evidence is thin
  • Tie-breakers are deterministic (score ↓, pull number ↓, commit SHA ↑)

What never affects the score

  • LLM output
  • MCP tool calls
  • Investigation Agent narratives
  • Integration secrets