Refreshing a Benchmark from New Signals
Use this playbook when customer-provided production examples, Contribution findings, source changes, or target-system changes show that the current Benchmark no longer represents the intended behavior.
Classify the signal before editing
Identify whether the signal changes the correctness standard, coverage vocabulary, available Cases, selected Dataset membership, or candidate behavior. These changes have different owners and version consequences.
Refresh path
- Record the signal, source, effective date, and latest Benchmark Version it may affect.
- Route a source or rule change to Reference Materials, Expert Contributions, and Correctness Governance.
- Route a changed behavior axis to Coverage Facets and review the impact on existing classifications.
- Route missing examples to Assets Cases, Synthesis, or Coverage Management Case Foundry; review candidates in Case Review.
- Change selected membership under Benchmark Datasets and create a new Dataset Snapshot.
- Create the Benchmark Version that represents the revised Dataset and governed evaluator boundary.
- Run the same saved Harness Version against the new boundary when you need to isolate the benchmark change. Run a new Harness Version separately when candidate behavior also changed.
- Explain results using the two named boundaries; Compare does not implicitly normalize evidence across different Benchmark Versions.
Preserve history
Do not edit an older Snapshot or Run to resemble current reality. Staleness means the evidence may no longer answer the current question; it does not erase what the old version measured.
Worked example
Refund policy change
A support team adds a new outage-credit exception. The existing Benchmark remains valid for the former rule but has no Cases for the new exception. Experts approve revised applicability and Rubrics; Case Foundry prepares eligible, ineligible, and missing-account-evidence Cases; Case Review admits them into the current Dataset. A new Snapshot and Benchmark Version show that the unchanged Harness still passes ordinary refunds but fails the new outage boundary.
Evidence to collect
- The new signal, its source, and the benchmark version it affects.
- Cases, reference responses where supported, Policies, Rubrics, or Coverage Facets changed by the signal.
- Human approval for whether the change updates standards, coverage, or both.
- Previous and refreshed Snapshot and Benchmark Version IDs, saved Harness Version, settings, and Run Metadata.
- Case-level evidence that explains movement caused by the revised boundary rather than candidate behavior.
Related docs
Source confidence
Doctrine-backed: the approved lifecycle routes new signals to the artifact that owns the change and preserves historical evidence. Linked code-backed pages define the current Coverage, Dataset, Snapshot, governance, and Evaluation operations.