Teammately Docs
Docs menu

recipe

Refreshing a Benchmark from New Signals

Update benchmark coverage and standards when new cases, review findings, or product changes appear.

Refreshing a Benchmark from New Signals

Use this playbook when customer-provided production examples, Contribution findings, source changes, or target-system changes show that the current Benchmark no longer represents the intended behavior.

Classify the signal before editing

Identify whether the signal changes the correctness standard, coverage vocabulary, available Cases, selected Dataset membership, or candidate behavior. These changes have different owners and version consequences.

Refresh path

  1. Record the signal, source, effective date, and latest Benchmark Version it may affect.
  2. Route a source or rule change to Reference Materials, Expert Contributions, and Correctness Governance.
  3. Route a changed behavior axis to Coverage Facets and review the impact on existing classifications.
  4. Route missing examples to Assets Cases, Synthesis, or Coverage Management Case Foundry; review candidates in Case Review.
  5. Change selected membership under Benchmark Datasets and create a new Dataset Snapshot.
  6. Create the Benchmark Version that represents the revised Dataset and governed evaluator boundary.
  7. Run the same saved Harness Version against the new boundary when you need to isolate the benchmark change. Run a new Harness Version separately when candidate behavior also changed.
  8. Explain results using the two named boundaries; Compare does not implicitly normalize evidence across different Benchmark Versions.

Preserve history

Do not edit an older Snapshot or Run to resemble current reality. Staleness means the evidence may no longer answer the current question; it does not erase what the old version measured.

Worked example

Refund policy change

A support team adds a new outage-credit exception. The existing Benchmark remains valid for the former rule but has no Cases for the new exception. Experts approve revised applicability and Rubrics; Case Foundry prepares eligible, ineligible, and missing-account-evidence Cases; Case Review admits them into the current Dataset. A new Snapshot and Benchmark Version show that the unchanged Harness still passes ordinary refunds but fails the new outage boundary.

Evidence to collect

  • The new signal, its source, and the benchmark version it affects.
  • Cases, reference responses where supported, Policies, Rubrics, or Coverage Facets changed by the signal.
  • Human approval for whether the change updates standards, coverage, or both.
  • Previous and refreshed Snapshot and Benchmark Version IDs, saved Harness Version, settings, and Run Metadata.
  • Case-level evidence that explains movement caused by the revised boundary rather than candidate behavior.

Source confidence

Doctrine-backed: the approved lifecycle routes new signals to the artifact that owns the change and preserves historical evidence. Linked code-backed pages define the current Coverage, Dataset, Snapshot, governance, and Evaluation operations.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context