# Benchmark Results Changed Unexpectedly
Generated: 2026-09-13T04:42:16.523Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: troubleshooting.benchmark-results-changed
title: Benchmark Results Changed Unexpectedly
summary: Diagnose result changes across target behavior, benchmark cases, standards, and versions.
kind: error
product_area: troubleshooting
status: stable
updated: 2026-08-23
canonical: /docs/troubleshooting/benchmark-results-changed-unexpectedly
---
# Benchmark Results Changed Unexpectedly
Use this when benchmark results change and the team is not sure whether the cause is model behavior or an artifact change.
## Symptom
A pass rate, policy-level result, rubric result, or case-level outcome changes between runs even though the expected candidate behavior did not obviously change.
## Likely causes
- Benchmark membership changed through case import, promotion, removal, or refresh.
- Policy, rubric, or applicability versions changed between runs.
- A different saved Harness Version, execution setting, or imported output-only Run was inspected.
- Run metadata or benchmark version selection differs from the previous run.
## Diagnostic checks
- Compare benchmark version, case count, policy versions, rubric versions, and applicability boundary.
- Inspect changed Cases and confirm whether each result came from a managed Run or imported output-only Run.
- Check run metadata for candidate, prompt, retrieval, or model differences.
- Open policy/rubric result changes and trace them to exact cases.
## Fix
- If artifacts changed, label the comparison as an artifact-boundary change rather than a pure behavior regression.
- If the candidate identity or settings changed, run the intended saved Harness Version with the intended settings.
- If the Benchmark boundary changed, create and name the appropriate Snapshot and Benchmark Version; do not rewrite the older Run.
- If the cause remains unclear, hold downstream action until the changed evidence can be explained.
## Prevention
- Record benchmark, case, policy, rubric, and candidate versions for every run.
- Use comparison views before summarizing score movement.
- Treat mapping, coverage, and standard changes as review-context boundaries.
- Keep previous runs reproducible for audit.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Compare Harness Versions](/docs/benchmark-evaluations/compare)
- [Read run results](/docs/benchmark-evaluations/inspect-results)
- [Inspect execution settings](/docs/benchmark-evaluations/execution-settings)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Versions, staleness, and resolution](/docs/object-model/versions-staleness-and-resolution)
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Benchmark versioning](/docs/governance/benchmark-versioning)
{% /related-card-grid %}
## Source confidence
Code-backed: Dataset Snapshots, Compare, Run detail, and Run Metadata expose the identities and boundaries needed to separate Dataset, candidate, configuration, and imported-output changes.
---
id: benchmark-evaluations.compare
title: Compare Harness Versions
summary: Compare two or more saved Harness Versions in a symmetric evidence matrix across Cases, evaluators, and Coverage Facets.
kind: task
product_area: benchmark_evaluations
status: stable
updated: 2026-09-13
canonical: /docs/benchmark-evaluations/compare
---
# Compare Harness Versions
Compare shows aggregate observed evidence for saved Harness Versions. Choose the Versions, compatible evaluation configuration, and measurement to compare. Use **Inspect individual Run comparisons** for the detailed result matrix.
## Prerequisites
Saved Harness Versions with evaluation results are shown for the selected Benchmark Version. Choose one or more Versions to inspect; Run counts may differ. An output-only imported reference Run cannot become a Harness column because it has no executable saved Version.
## Select a row mode
The individual Run matrix offers row modes including Cases, all results, Policies, Rubrics, Dimension ontology values, Project Topics, Topic Groups, and Case Construction Patterns. Use Cases to inspect concrete disagreement, Policies or Rubrics to locate correctness movement, and Coverage Facets to see whether gains concentrate in one behavior slice.
## Steps
1. Confirm the immutable Benchmark Version and choose at least two visible Harness Versions.
2. Select an aggregate measurement, or open individual Run comparisons and choose a row mode.
3. Check evidence completeness for each Harness column. A blank or incomplete cell is not a failure.
4. Locate rows with material disagreement and connect them back to Case and evaluator evidence.
5. Preserve regressions and required-criterion failures next to gains.
6. Use the exact candidate and row evidence when starting an Improvement Session or requesting an Expert Contribution.
Compare reads existing evidence and does not mutate Runs. Selecting aggregate Harness Versions recomputes their comparison on compatible evidence; individual matrix visibility is local presentation. It does not activate Harnesses or choose a winner.
{% example-demo title="Example: facet-local improvement" %}
Three Harness Versions look similar overall. The Project Topic row mode shows that Version 14 improves source-authority Topics but regresses escalation Topics. Switching to Cases identifies two regressions, and the team starts Improve with those exact failures instead of claiming a uniform improvement.
{% /example-demo %}
## Common mistakes
- Comparing different Benchmark Versions as though only the candidate moved.
- Treating missing evidence as a failed cell.
- Reading a facet aggregate without checking the distinct Cases behind it.
- Describing an imported output-only Run as a Harness Version.
- Selecting the newest Version solely because it is newest.
## Object and state changes
Compare reads existing evidence. Selecting aggregate Versions reads their observed evidence; selecting individual matrix row modes changes presentation; it does not activate a Harness, mutate a Run, or retain a candidate.
## Success criteria
- At least two exact Harness Versions share the same Benchmark Version.
- Incomplete cells remain distinct from failed evidence.
- Material movement resolves to Cases, evaluators, or Coverage Facets.
## Common failure modes
- Comparing moved evidence boundaries as candidate-only change.
- Treating local column visibility as product configuration.
- Using an output-only Run as a Harness column.
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Arena and Rankings](/docs/benchmark-evaluations/arena-and-rankings)
- [Harnesses](/docs/assets/harnesses)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
{% /related-card-grid %}
## Source confidence
Code-backed: the active Compare route and Evaluation Matrix define Harness columns, local visibility, symmetric comparison, and the current evidence row modes.
---
id: governance.reproducibility
title: Reproducibility
summary: Preserve enough source context to explain and repeat correctness decisions.
kind: reference
product_area: governance
status: stable
updated: 2026-08-23
canonical: /docs/governance/reproducibility
---
# Reproducibility
## Definition
Reproducibility means preserving enough exact identity and observable configuration to explain what was evaluated and to repeat the supported execution path. It does not mean every future execution will produce an identical stochastic output. It means a reader can distinguish changes in candidate, evidence, evaluator, sampling, and runtime metadata instead of attributing every result difference to the model.
## Fields, states, or lifecycle rules
- Preserve Project, Benchmark, Benchmark Version, Dataset Snapshot, and Run identity.
- Preserve the exact saved Harness Version rather than an editable draft or display label.
- Preserve admitted Case, Policy, and Rubric version boundaries through the Benchmark Version.
- Record Run Group, attempt, sampling profile, evaluator set, and visible execution settings.
- Retain run metadata and measured telemetry when captured; missing values remain unknown.
- Record completeness, incomplete Cases, and terminal state beside scores.
- Use canonical evaluation receipts for Improvement Session candidate claims.
- Do not claim private worker activity, hidden reasoning, infrastructure internals, or unavailable traces as reproducibility evidence.
## Related objects
Dataset Snapshots preserve the evidence set. Benchmark Evaluations preserves candidate, Run, settings, results, and available telemetry. Compare and Arena interpret candidates inside compatible evidence boundaries. Improve adds Goal Contract, candidate, and canonical receipt identity when evaluation drives code or Harness changes.
{% example-demo title="Diagnosing a score change" %}
Two Runs use the same Harness Version but report different pass rates. The operator confirms that one Run used Benchmark Version 6 and the other used Version 7, which added source-conflict Cases and a revised grounding Rubric. The version and completeness record explains the movement. The team avoids filing a candidate regression until it compares Runs inside the same evidence boundary.
{% /example-demo %}
## Source confidence
Code-backed: Snapshot, Run detail, and run-metadata surfaces expose the immutable evidence boundary, candidate identity, status, counts, settings, and available metadata needed for supported reproducibility. They do not promise deterministic model output or unrestricted execution traces.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Benchmark Versioning](/docs/governance/benchmark-versioning)
- [Compare Harness Versions](/docs/benchmark-evaluations/compare)
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Product quickstart](/docs/quickstart)
- [Task index](/docs/operating-manual/task-index)
{% /related-card-grid %}
---
id: troubleshooting.benchmark-runs
title: Benchmark run troubleshooting
summary: Diagnose a Run that cannot start, has no usable outputs, or produces results that cannot be compared safely.
kind: error
product_area: troubleshooting
status: stable
updated: 2026-09-07
canonical: /docs/troubleshooting/benchmark-runs
---
# Benchmark run troubleshooting
## Symptoms
- The evaluation surface has no Benchmark Version to run.
- Output import or mapping cannot identify a Case or output column.
- A Run is created but remains empty, incomplete, or failed.
- Results appear under the wrong candidate label or Harness Version.
- Two Runs show a score change but do not share a comparable evidence boundary.
## Likely causes
- No immutable Snapshot exists for the intended evidence set.
- Imported outputs are mapped to the wrong Case or column.
- The saved Harness Version or candidate metadata does not match the evaluated system.
- The Run is partial, failed, or being compared across different Benchmark Versions.
## Check the boundary before the failure
1. Confirm the URL and page identify the intended Project, Benchmark, and Benchmark Version.
2. Open the Version or Snapshot and verify it contains the expected Cases and evaluators. If it does not, repair coverage and create a new Snapshot; do not edit the historical Run.
3. In Run setup, confirm the saved Harness Version and candidate metadata describe the system that produced the outputs.
4. If importing outputs, inspect the mapping preview. Match the Case identifier and candidate-output column deliberately; do not use a reference-output column as candidate behavior.
5. Open Run detail and inspect status, Case count, errors, metadata, and per-Case results before trusting aggregates.
## Fix
- **No runnable version:** finish Case selection and create a Snapshot first.
- **No mapped outputs:** correct Case identifiers or column mapping, then submit again under the intended Run.
- **Wrong Harness or metadata:** create a correctly configured Run. Do not relabel completed evidence to represent a different system.
- **Partial failure:** preserve successful per-Case evidence when the product does, correct the failed input or execution boundary, and rerun using a clearly named attempt.
- **Confusing comparison:** compare the Benchmark Version, Harness Version, candidate metadata, and evaluator boundary. Qualify or avoid the comparison when more than the intended variable changed.
## Prevention
Create a Snapshot before execution, save the exact Harness Version, preview output mapping on representative Cases, and name candidate metadata consistently. Before comparing, confirm that every difference between the two Runs is intentional and visible.
## Recovery check
Open the recovered Run and sample several Case results. Confirm the displayed input, candidate output, reference output where present, applicable Rubrics, and metadata all belong together. A completed status alone does not prove correct mapping.
{% example-demo title="Example: scores drop after output import" %}
A team imports a new candidate file and sees a sudden score collapse. Run detail shows that the column containing reference outputs was mapped as candidate output. The operator creates a new Run, maps the actual candidate column, preserves the mistaken Run as an identifiable failed attempt, and compares only the corrected Run with the prior candidate under the same Benchmark Version.
{% /example-demo %}
## Source confidence
Code-backed: Runs index, setup, output mapping, Run detail, and metadata-display implementations establish the identifiers, mapping choices, and evidence shown during diagnosis. Backend-provider errors and customer Harness behavior may require additional operational logs outside this page.
## Related task pages
{% related-card-grid title="Related workflows" %}
- [Run an evaluation](/docs/benchmark-evaluations/run-evaluation)
- [Inspect results](/docs/benchmark-evaluations/inspect-results)
- [Compare Harness Versions](/docs/benchmark-evaluations/compare)
- [Run Metadata](/docs/benchmark-evaluations/run-metadata)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference" %}
- [Benchmarks and versions](/docs/concepts/benchmarks-and-versions)
- [Benchmark Snapshots](/docs/coverage-engineering/benchmark-snapshots)
- [Outputs](/docs/object-model/outputs)
{% /related-card-grid %}