Local development
QA scenarios, seed locations, replay harness, and expectations.
Local fixtures (scenarios A–D)
Fixtures prove deterministic ranking, grounded evidence, stable re-runs, and safe LLM behavior. They are the fastest way to learn ChangeGraph without provider accounts.
Scenario map
| Scenario | Intent | Pass hint |
|---|---|---|
| A | Breaking PR (happy path) | Culprit ranks Top-1 (hard fail outside top 3) |
| B | Unrelated / distractor PR | Must not be blamed; show no file overlap |
| C | Multi-candidate discrimination | Strongest overlap above distractors; stable order |
| D | No / insufficient overlap | Empty or low-confidence; LLM must not invent a culprit |
Shared assumptions: synthetic tenant, one GitHub repo, one Sentry project, lookback relative to incident time T0, stack under src/payments/ (exact paths per scenario). Full matrices: repo docs/SCENARIOS.md.
Where they live
| Path | Role |
|---|---|
src/lib/fixtures/scenarios.ts | Seed inputs for scoring |
src/lib/db/seed-data.ts | UI fixtures |
test-fixtures/ | Shared index for Vitest + e2e |
seed-data/ | Canonical IDs + org payload |
mock-integrations/ | Mock GitHub / Sentry / Vercel JSON |
pnpm seed | Relational seed into Postgres when available. See Where the running app reads data |
Run them
pnpm test # scoring + related unit tests
pnpm test:e2e # API path: fixtures → incident → candidates → evidence
pnpm replay # evaluate A–D; Top-1 / Top-3 accuracy
pnpm replay A # scenario A detail (breaker pr-a / PR #101)
pnpm replay --list
Product expectations
- With an LLM provider: citations ⊆ evidence package
- Without a provider: UI shows
not_configuredand asks you to add a key - Re-ingest same fixtures → scores unchanged across ≥3 runs
- Scenario A e2e: overlap evidence + LKG deploy present
Replay metrics
See docs/REPLAY.md. Use replay when changing scorer weights or windows. Do not “fix” fixtures to match a weaker model.
Local install, feature gates, and a manual smoke pass (including fixtures A–D): repo docs/TESTING.md.