Benchmark Run Metadata
Definition
Benchmark Run Metadata is descriptive context attached to benchmark-level evaluation work. It helps operators interpret a Run without replacing the exact Benchmark Version, Harness Version, Run Group, or published Regime Version that defines the evidence boundary.
Project-level Run Metadata templates are retired. When the current evaluation surface offers metadata fields, manage them at the benchmark or Run setup boundary and keep the values specific to the evidence being created.
Fields, states, or lifecycle rules
- Metadata describes a benchmark evaluation context; it is not a Policy, Rubric, Case, Harness Version, Benchmark Version, or Regime Version.
- Existing Runs retain the metadata and exact version identities recorded with their evidence.
- Benchmark-level fields can be managed from the benchmark context when that surface exposes the control.
- External model configuration is declared when the Run is created rather than through a project-level template.
- Metadata can help compare or interpret Runs, but it does not make an external metric a Teammately-verified result.
Choose the right boundary
Put executable candidate behavior in a Harness and its saved Version. Put Case content and materials in Assets and Benchmark Datasets. Put evaluation scoring behavior in the published Regime and governed Policies and Rubrics. Use Run Metadata only for descriptive context that should travel with a particular benchmark evaluation.
Worked example
Example: benchmark-level experiment context
Two Runs use the same Benchmark Version but different saved Harness Versions. Their benchmark-level metadata records the experiment labels and external model configuration needed to interpret the comparison. The metadata does not change either candidate identity or the published Regime used to score the evidence.
Source confidence
Code-backed: the current Run presentation and setup surface expose benchmark-level metadata context and explicitly fence off retired project-level templates. Exact fields depend on the benchmark evaluation surface in use.