Teammately Docs
Docs menu

error

Benchmark Results Changed Unexpectedly

Diagnose result changes across target behavior, benchmark cases, standards, and versions.

Benchmark Results Changed Unexpectedly

Use this when benchmark results change and the team is not sure whether the cause is model behavior or an artifact change.

Symptom

A pass rate, policy-level result, rubric result, or case-level outcome changes between runs even though the expected candidate behavior did not obviously change.

Likely causes

  • Benchmark membership changed through case import, promotion, removal, or refresh.
  • Policy, rubric, or applicability versions changed between runs.
  • A different saved Harness Version, execution setting, or imported output-only Run was inspected.
  • Run metadata or benchmark version selection differs from the previous run.

Diagnostic checks

  • Compare benchmark version, case count, policy versions, rubric versions, and applicability boundary.
  • Inspect changed Cases and confirm whether each result came from a managed Run or imported output-only Run.
  • Check run metadata for candidate, prompt, retrieval, or model differences.
  • Open policy/rubric result changes and trace them to exact cases.

Fix

  • If artifacts changed, label the comparison as an artifact-boundary change rather than a pure behavior regression.
  • If the candidate identity or settings changed, run the intended saved Harness Version with the intended settings.
  • If the Benchmark boundary changed, create and name the appropriate Snapshot and Benchmark Version; do not rewrite the older Run.
  • If the cause remains unclear, hold downstream action until the changed evidence can be explained.

Prevention

  • Record benchmark, case, policy, rubric, and candidate versions for every run.
  • Use comparison views before summarizing score movement.
  • Treat mapping, coverage, and standard changes as review-context boundaries.
  • Keep previous runs reproducible for audit.

Source confidence

Code-backed: Dataset Snapshots, Compare, Run detail, and Run Metadata expose the identities and boundaries needed to separate Dataset, candidate, configuration, and imported-output changes.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context