---
id: benchmark-evaluations.compare
title: Compare Harness Versions
summary: Compare two or more saved Harness Versions in a symmetric evidence matrix across Cases, evaluators, and Coverage Facets.
kind: task
product_area: benchmark_evaluations
status: stable
updated: 2026-09-13
canonical: /docs/benchmark-evaluations/compare
---

# Compare Harness Versions

Compare shows aggregate observed evidence for saved Harness Versions. Choose the Versions, compatible evaluation configuration, and measurement to compare. Use **Inspect individual Run comparisons** for the detailed result matrix.

## Prerequisites

Saved Harness Versions with evaluation results are shown for the selected Benchmark Version. Choose one or more Versions to inspect; Run counts may differ. An output-only imported reference Run cannot become a Harness column because it has no executable saved Version.

## Select a row mode

The individual Run matrix offers row modes including Cases, all results, Policies, Rubrics, Dimension ontology values, Project Topics, Topic Groups, and Case Construction Patterns. Use Cases to inspect concrete disagreement, Policies or Rubrics to locate correctness movement, and Coverage Facets to see whether gains concentrate in one behavior slice.

## Steps

1. Confirm the immutable Benchmark Version and choose at least two visible Harness Versions.
2. Select an aggregate measurement, or open individual Run comparisons and choose a row mode.
3. Check evidence completeness for each Harness column. A blank or incomplete cell is not a failure.
4. Locate rows with material disagreement and connect them back to Case and evaluator evidence.
5. Preserve regressions and required-criterion failures next to gains.
6. Use the exact candidate and row evidence when starting an Improvement Session or requesting an Expert Contribution.

Compare reads existing evidence and does not mutate Runs. Selecting aggregate Harness Versions recomputes their comparison on compatible evidence; individual matrix visibility is local presentation. It does not activate Harnesses or choose a winner.

{% example-demo title="Example: facet-local improvement" %}
Three Harness Versions look similar overall. The Project Topic row mode shows that Version 14 improves source-authority Topics but regresses escalation Topics. Switching to Cases identifies two regressions, and the team starts Improve with those exact failures instead of claiming a uniform improvement.
{% /example-demo %}

## Common mistakes

- Comparing different Benchmark Versions as though only the candidate moved.
- Treating missing evidence as a failed cell.
- Reading a facet aggregate without checking the distinct Cases behind it.
- Describing an imported output-only Run as a Harness Version.
- Selecting the newest Version solely because it is newest.

## Object and state changes

Compare reads existing evidence. Selecting aggregate Versions reads their observed evidence; selecting individual matrix row modes changes presentation; it does not activate a Harness, mutate a Run, or retain a candidate.

## Success criteria

- At least two exact Harness Versions share the same Benchmark Version.
- Incomplete cells remain distinct from failed evidence.
- Material movement resolves to Cases, evaluators, or Coverage Facets.

## Common failure modes

- Comparing moved evidence boundaries as candidate-only change.
- Treating local column visibility as product configuration.
- Using an output-only Run as a Harness column.

## Related reference pages

{% related-card-grid title="Related reference pages" %}
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Arena and Rankings](/docs/benchmark-evaluations/arena-and-rankings)
- [Harnesses](/docs/assets/harnesses)
{% /related-card-grid %}

## Related troubleshooting pages

{% related-card-grid title="Related troubleshooting pages" %}
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
{% /related-card-grid %}

## Source confidence

Code-backed: the active Compare route and Evaluation Matrix define Harness columns, local visibility, symmetric comparison, and the current evidence row modes.
