ChangeGraphChangeGraph

Local development

QA scenarios, seed locations, replay harness, and expectations.

Local fixtures (scenarios A–D)

Fixtures prove deterministic ranking, grounded evidence, stable re-runs, and safe LLM behavior. They are the fastest way to learn ChangeGraph without provider accounts.

Scenario map

ScenarioIntentPass hint
ABreaking PR (happy path)Culprit ranks Top-1 (hard fail outside top 3)
BUnrelated / distractor PRMust not be blamed; show no file overlap
CMulti-candidate discriminationStrongest overlap above distractors; stable order
DNo / insufficient overlapEmpty or low-confidence; LLM must not invent a culprit

Shared assumptions: synthetic tenant, one GitHub repo, one Sentry project, lookback relative to incident time T0, stack under src/payments/ (exact paths per scenario). Full matrices: repo docs/SCENARIOS.md.

Where they live

PathRole
src/lib/fixtures/scenarios.tsSeed inputs for scoring
src/lib/db/seed-data.tsUI fixtures
test-fixtures/Shared index for Vitest + e2e
seed-data/Canonical IDs + org payload
mock-integrations/Mock GitHub / Sentry / Vercel JSON
pnpm seedRelational seed into Postgres when available. See Where the running app reads data

Run them

pnpm test          # scoring + related unit tests
pnpm test:e2e      # API path: fixtures → incident → candidates → evidence
pnpm replay        # evaluate A–D; Top-1 / Top-3 accuracy
pnpm replay A      # scenario A detail (breaker pr-a / PR #101)
pnpm replay --list

Product expectations

  • With an LLM provider: citations ⊆ evidence package
  • Without a provider: UI shows not_configured and asks you to add a key
  • Re-ingest same fixtures → scores unchanged across ≥3 runs
  • Scenario A e2e: overlap evidence + LKG deploy present

Replay metrics

See docs/REPLAY.md. Use replay when changing scorer weights or windows. Do not “fix” fixtures to match a weaker model.

Local install, feature gates, and a manual smoke pass (including fixtures A–D): repo docs/TESTING.md.