# Benchmark Run Metadata
Generated: 2026-09-13T04:39:06.434Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: benchmark-evaluations.run-metadata
title: Benchmark Run Metadata
summary: Interpret benchmark-level descriptive fields without confusing them with project settings or version identity.
kind: reference
product_area: benchmark_evaluations
status: stable
updated: 2026-09-07
canonical: /docs/benchmark-evaluations/run-metadata
---
# Benchmark Run Metadata
## Definition
Benchmark Run Metadata is descriptive context attached to benchmark-level evaluation work. It helps operators interpret a Run without replacing the exact Benchmark Version, Harness Version, Run Group, or published Regime Version that defines the evidence boundary.
Project-level Run Metadata templates are retired. When the current evaluation surface offers metadata fields, manage them at the benchmark or Run setup boundary and keep the values specific to the evidence being created.
## Fields, states, or lifecycle rules
- Metadata describes a benchmark evaluation context; it is not a Policy, Rubric, Case, Harness Version, Benchmark Version, or Regime Version.
- Existing Runs retain the metadata and exact version identities recorded with their evidence.
- Benchmark-level fields can be managed from the benchmark context when that surface exposes the control.
- External model configuration is declared when the Run is created rather than through a project-level template.
- Metadata can help compare or interpret Runs, but it does not make an external metric a Teammately-verified result.
## Choose the right boundary
Put executable candidate behavior in a Harness and its saved Version. Put Case content and materials in Assets and Benchmark Datasets. Put evaluation scoring behavior in the published Regime and governed Policies and Rubrics. Use Run Metadata only for descriptive context that should travel with a particular benchmark evaluation.
{% example-demo title="Example: benchmark-level experiment context" %}
Two Runs use the same Benchmark Version but different saved Harness Versions. Their benchmark-level metadata records the experiment labels and external model configuration needed to interpret the comparison. The metadata does not change either candidate identity or the published Regime used to score the evidence.
{% /example-demo %}
## Source confidence
Code-backed: the current Run presentation and setup surface expose benchmark-level metadata context and explicitly fence off retired project-level templates. Exact fields depend on the benchmark evaluation surface in use.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation)
- [Inspect evaluation results](/docs/benchmark-evaluations/inspect-results)
- [Compare evaluation results](/docs/benchmark-evaluations/compare)
{% /related-card-grid %}
---
id: benchmark-evaluations.overview
title: Benchmark Evaluations
summary: Run and inspect exact Harness Versions against an immutable Benchmark Version through Dashboard, List, Arena, and Compare.
kind: concept
product_area: benchmark_evaluations
status: stable
updated: 2026-09-13
canonical: /docs/benchmark-evaluations
---
# Benchmark Evaluations
Benchmark Evaluations is the version-scoped workspace for executing and comparing candidate systems. The active top-level tabs are **Dashboard**, **List**, **Arena**, and **Compare**. Every managed Run binds an exact saved Harness Version to the immutable Benchmark Version shown in the route.
> Evaluation boundary
>
> Interpret evidence inside its recorded Benchmark Version, Harness Version, Run or Run Group, evaluator set, and metadata. Run counts belong to launches. Additional launches add evidence without rewriting earlier Runs.
## Surfaces and objects
Dashboard summarizes progress, leaderboards, rank progression across Runs, and available resource telemetry. List is segmented into **Runs**, **Evaluation results**, and **Traces / Spans**. The results segment summarizes Case outcomes and Policy or Rubric failures. Arena compares candidate pairs across governed metrics. Compare is a symmetric matrix of Harness Versions across selected evidence rows.
A Run Group can collect one standard attempt or repeated attempts. A Run records one candidate execution and its per-Case progress. Evaluation results record the admitted Policy and Rubric outcomes. Costs, tokens, and latency are telemetry only when the provider or execution path captured them.
> Traces / Spans capability fence
>
> The List navigation exposes Traces / Spans, but the current benchmark API does not expose evaluation execution traces. Do not claim that trajectories, spans, private reasoning, or tool traces can be inspected from Benchmark Evaluations today.
## Decision checkpoint
| Need | Open | Evidence to preserve |
| --- | --- | --- |
| Configure and launch managed Runs | Evaluation Settings and New evaluation run | Machine, saved Harness Versions, and per-Harness Run counts |
| Start candidate execution | Run modal | Exact Harness and Benchmark Versions |
| Inspect status and output summaries | List → Runs or Evaluation results | Run Group, attempt, Case counts, incomplete state |
| Compare candidate pairs | Arena | Metric family, pair count, only-A, only-B, shared failures |
| Compare many candidates by governed rows | Compare | Harness columns and chosen Case or facet row mode |
| Admit external reference outputs | Output mapping | Case mapping, attempt assignment, insert/update report |
## Rankings and repeated sampling
Dashboard aggregates compatible observed Runs for each saved Harness Version across launches. Average score weights Runs equally. Supported binary views report passed at least once or passed every time over the observed case outcomes. Counts and missing evidence are shown; unequal counts do not prevent comparison. Historical group metrics retain their recorded meanings.
Ranking is a routing signal. A candidate can lead overall while failing required Policy or high-impact Rubric evidence. Use Arena or Compare to locate the disagreement and List to confirm completeness before starting Improve work.
## External outputs
Uploaded or API-supplied reference outputs create output-only Runs that can be scored and inspected in List. They are not saved Harness Versions and therefore cannot be optimized in Improve or selected as Harness columns in Compare or Arena.
{% example-demo title="Example: repeated evaluation without evidence drift" %}
A team launches three Runs of Harness Version 8 and one Run of Version 11 against the same Benchmark Version. Both appear with their evidence counts. A later launch of Version 11 adds two Runs to its aggregate evidence without changing either launch group. The team can inspect individual Runs before deciding whether more evidence is useful.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Configure evaluation execution](/docs/benchmark-evaluations/execution-settings)
- [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation)
- [Inspect evaluation results](/docs/benchmark-evaluations/inspect-results)
- [Use Arena and rankings](/docs/benchmark-evaluations/arena-and-rankings)
- [Compare Harness Versions](/docs/benchmark-evaluations/compare)
- [Map external outputs](/docs/benchmark-evaluations/output-mapping)
{% /related-card-grid %}
## Source confidence
Code-backed: the active version-scoped workspace, settings, Run modal, List segments, Dashboard, Arena, and Compare routes define the current evaluation model and capability fences.
---
id: benchmark-evaluations.execution-settings
title: Evaluation Execution Settings
summary: Configure machines and choose saved Harness Versions and Run counts at launch.
kind: reference
product_area: benchmark_evaluations
status: stable
updated: 2026-09-13
canonical: /docs/benchmark-evaluations/execution-settings
---
# Evaluation Execution Settings
## Definition
Evaluation Settings contains machine settings for managed execution. Harness selection and repetition are choices made when launching an evaluation.
## Fields, states, or lifecycle rules
- Choose an exact saved project Harness Version when launching. Benchmark activation is not required.
- Choose a whole number of Runs from 1 through 50 for each selected Harness. The default is one, and different Harnesses may have different counts.
- Each managed Harness launch creates one Run Group containing the requested Runs, including a one-Run group.
- Existing Runs keep their immutable configuration and evidence. Launching more Runs creates another group rather than rewriting previous membership.
## Reliability choices
Dashboard, Compare, and Arena aggregate compatible observed Runs for each saved Harness Version across launch groups. Average score weights individual Runs equally. Passed at least once and passed every time summarize observed binary case outcomes where the evaluation framework supports them.
Run counts can differ. The product displays the counts and coverage and warns about unequal evidence without requiring another launch. Missing or infrastructure-failed observations are not numerical successes or failures.
## Before launching
Check the exact saved Version, machine selection, and requested execution volume. Runtime preparation and required access still apply. Imported external outputs remain a separate flow because they have no executable Harness Version.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Manage Harnesses](/docs/assets/harnesses)
- [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation)
{% /related-card-grid %}
## Source confidence
Code-backed: Evaluation Settings defines machine configuration; the shared Run modal defines saved-Version selection and per-launch Run counts.
---
id: benchmark-evaluations.compare
title: Compare Harness Versions
summary: Compare two or more saved Harness Versions in a symmetric evidence matrix across Cases, evaluators, and Coverage Facets.
kind: task
product_area: benchmark_evaluations
status: stable
updated: 2026-09-13
canonical: /docs/benchmark-evaluations/compare
---
# Compare Harness Versions
Compare shows aggregate observed evidence for saved Harness Versions. Choose the Versions, compatible evaluation configuration, and measurement to compare. Use **Inspect individual Run comparisons** for the detailed result matrix.
## Prerequisites
Saved Harness Versions with evaluation results are shown for the selected Benchmark Version. Choose one or more Versions to inspect; Run counts may differ. An output-only imported reference Run cannot become a Harness column because it has no executable saved Version.
## Select a row mode
The individual Run matrix offers row modes including Cases, all results, Policies, Rubrics, Dimension ontology values, Project Topics, Topic Groups, and Case Construction Patterns. Use Cases to inspect concrete disagreement, Policies or Rubrics to locate correctness movement, and Coverage Facets to see whether gains concentrate in one behavior slice.
## Steps
1. Confirm the immutable Benchmark Version and choose at least two visible Harness Versions.
2. Select an aggregate measurement, or open individual Run comparisons and choose a row mode.
3. Check evidence completeness for each Harness column. A blank or incomplete cell is not a failure.
4. Locate rows with material disagreement and connect them back to Case and evaluator evidence.
5. Preserve regressions and required-criterion failures next to gains.
6. Use the exact candidate and row evidence when starting an Improvement Session or requesting an Expert Contribution.
Compare reads existing evidence and does not mutate Runs. Selecting aggregate Harness Versions recomputes their comparison on compatible evidence; individual matrix visibility is local presentation. It does not activate Harnesses or choose a winner.
{% example-demo title="Example: facet-local improvement" %}
Three Harness Versions look similar overall. The Project Topic row mode shows that Version 14 improves source-authority Topics but regresses escalation Topics. Switching to Cases identifies two regressions, and the team starts Improve with those exact failures instead of claiming a uniform improvement.
{% /example-demo %}
## Common mistakes
- Comparing different Benchmark Versions as though only the candidate moved.
- Treating missing evidence as a failed cell.
- Reading a facet aggregate without checking the distinct Cases behind it.
- Describing an imported output-only Run as a Harness Version.
- Selecting the newest Version solely because it is newest.
## Object and state changes
Compare reads existing evidence. Selecting aggregate Versions reads their observed evidence; selecting individual matrix row modes changes presentation; it does not activate a Harness, mutate a Run, or retain a candidate.
## Success criteria
- At least two exact Harness Versions share the same Benchmark Version.
- Incomplete cells remain distinct from failed evidence.
- Material movement resolves to Cases, evaluators, or Coverage Facets.
## Common failure modes
- Comparing moved evidence boundaries as candidate-only change.
- Treating local column visibility as product configuration.
- Using an output-only Run as a Harness column.
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Arena and Rankings](/docs/benchmark-evaluations/arena-and-rankings)
- [Harnesses](/docs/assets/harnesses)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
{% /related-card-grid %}
## Source confidence
Code-backed: the active Compare route and Evaluation Matrix define Harness columns, local visibility, symmetric comparison, and the current evidence row modes.
---
id: governance.reproducibility
title: Reproducibility
summary: Preserve enough source context to explain and repeat correctness decisions.
kind: reference
product_area: governance
status: stable
updated: 2026-08-23
canonical: /docs/governance/reproducibility
---
# Reproducibility
## Definition
Reproducibility means preserving enough exact identity and observable configuration to explain what was evaluated and to repeat the supported execution path. It does not mean every future execution will produce an identical stochastic output. It means a reader can distinguish changes in candidate, evidence, evaluator, sampling, and runtime metadata instead of attributing every result difference to the model.
## Fields, states, or lifecycle rules
- Preserve Project, Benchmark, Benchmark Version, Dataset Snapshot, and Run identity.
- Preserve the exact saved Harness Version rather than an editable draft or display label.
- Preserve admitted Case, Policy, and Rubric version boundaries through the Benchmark Version.
- Record Run Group, attempt, sampling profile, evaluator set, and visible execution settings.
- Retain run metadata and measured telemetry when captured; missing values remain unknown.
- Record completeness, incomplete Cases, and terminal state beside scores.
- Use canonical evaluation receipts for Improvement Session candidate claims.
- Do not claim private worker activity, hidden reasoning, infrastructure internals, or unavailable traces as reproducibility evidence.
## Related objects
Dataset Snapshots preserve the evidence set. Benchmark Evaluations preserves candidate, Run, settings, results, and available telemetry. Compare and Arena interpret candidates inside compatible evidence boundaries. Improve adds Goal Contract, candidate, and canonical receipt identity when evaluation drives code or Harness changes.
{% example-demo title="Diagnosing a score change" %}
Two Runs use the same Harness Version but report different pass rates. The operator confirms that one Run used Benchmark Version 6 and the other used Version 7, which added source-conflict Cases and a revised grounding Rubric. The version and completeness record explains the movement. The team avoids filing a candidate regression until it compares Runs inside the same evidence boundary.
{% /example-demo %}
## Source confidence
Code-backed: Snapshot, Run detail, and run-metadata surfaces expose the immutable evidence boundary, candidate identity, status, counts, settings, and available metadata needed for supported reproducibility. They do not promise deterministic model output or unrestricted execution traces.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Benchmark Versioning](/docs/governance/benchmark-versioning)
- [Compare Harness Versions](/docs/benchmark-evaluations/compare)
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Product quickstart](/docs/quickstart)
- [Task index](/docs/operating-manual/task-index)
{% /related-card-grid %}