Use ChangeGraph
Scorer v2 weights 25/25/25/15/10, penalties, windows, and LKG.
Scoring & last-known-good
Scores are CORRELATION only, never causation. Re-running on the same inputs must yield stable rankings (scorer_version: v2).
Weight table (max 100)
| Key | Weight | Role |
|---|---|---|
temporal | 25 | Time proximity to the incident |
deployment | 25 | Membership on the mapped release / deploy |
file_overlap | 25 | Overlap between changed paths and stack frames |
service | 15 | Culprit / path-depth / service·transaction proxy |
historical | 10 | Same-release / prior deploy membership |
Defined in src/lib/scoring/weights.ts and locked by unit tests.
Penalties
| Penalty | Value | When |
|---|---|---|
no_file_overlap | −35 | Candidate paths do not intersect stack |
only_test_or_docs | −10 | Changes look like tests/docs only |
Penalties keep time-proximate distractors (README PRs) from winning on recency alone.
Investigation windows
Ladder (hours): 24 → 72 → 168 (7d).
- Start narrow when LKG / release context is rich
- Expand when the candidate set is empty or too thin
- UI may surface
windowExpanded/ expansion reason
Last-known-good (LKG)
LKG bounds “what changed since things were healthy”:
- Prefer a prior known-good deploy (
lkg_sha, environment, deployed_at) - Prefer candidates between LKG and the incident over ancient history
- When LKG is unknown, fall back to the configured time window
Treat LKG as a correlation aid, not proof of root cause. Fixtures include an LKG deploy for scenario A (good000lkg… in e2e expectations).
Evidence package
Only fields the scorer (and sanitizer / packer) allow are packaged for the UI and Mastra. Packed context: Agent context.
- Overlapping paths
- Timing / release membership
- Counter-evidence
- LKG snapshot fields when present
Secrets and raw tokens are stripped.
Interpreting ranks
- Top-1 / Top-3 are decision supports for humans (
pnpm replayreports accuracy) - Low scores or empty candidate lists are valid (scenarios B / D)
- Never invent a culprit PR when evidence is thin
- Tie-breakers are deterministic (score ↓, pull number ↓, commit SHA ↑)
What never affects the score
- LLM output
- MCP tool calls
- Investigation Agent narratives
- Integration secrets