Teammately Docs
Docs menu

concept

Benchmark Evaluations

Run and inspect exact Harness Versions against an immutable Benchmark Version through Dashboard, List, Arena, and Compare.

Benchmark Evaluations

Benchmark Evaluations is the version-scoped workspace for executing and comparing candidate systems. The active top-level tabs are Dashboard, List, Arena, and Compare. Every managed Run binds an exact saved Harness Version to the immutable Benchmark Version shown in the route.

Surfaces and objects

Dashboard summarizes progress, leaderboards, rank progression across Runs, and available resource telemetry. List is segmented into Runs, Evaluation results, and Traces / Spans. The results segment summarizes Case outcomes and Policy or Rubric failures. Arena compares candidate pairs across governed metrics. Compare is a symmetric matrix of Harness Versions across selected evidence rows.

A Run Group can collect one standard attempt or repeated attempts. A Run records one candidate execution and its per-Case progress. Evaluation results record the admitted Policy and Rubric outcomes. Costs, tokens, and latency are telemetry only when the provider or execution path captured them.

Decision checkpoint

NeedOpenEvidence to preserve
Configure and launch managed RunsEvaluation Settings and New evaluation runMachine, saved Harness Versions, and per-Harness Run counts
Start candidate executionRun modalExact Harness and Benchmark Versions
Inspect status and output summariesList → Runs or Evaluation resultsRun Group, attempt, Case counts, incomplete state
Compare candidate pairsArenaMetric family, pair count, only-A, only-B, shared failures
Compare many candidates by governed rowsCompareHarness columns and chosen Case or facet row mode
Admit external reference outputsOutput mappingCase mapping, attempt assignment, insert/update report

Rankings and repeated sampling

Dashboard aggregates compatible observed Runs for each saved Harness Version across launches. Average score weights Runs equally. Supported binary views report passed at least once or passed every time over the observed case outcomes. Counts and missing evidence are shown; unequal counts do not prevent comparison. Historical group metrics retain their recorded meanings.

Ranking is a routing signal. A candidate can lead overall while failing required Policy or high-impact Rubric evidence. Use Arena or Compare to locate the disagreement and List to confirm completeness before starting Improve work.

External outputs

Uploaded or API-supplied reference outputs create output-only Runs that can be scored and inspected in List. They are not saved Harness Versions and therefore cannot be optimized in Improve or selected as Harness columns in Compare or Arena.

Worked example

Example: repeated evaluation without evidence drift

A team launches three Runs of Harness Version 8 and one Run of Version 11 against the same Benchmark Version. Both appear with their evidence counts. A later launch of Version 11 adds two Runs to its aggregate evidence without changing either launch group. The team can inspect individual Runs before deciding whether more evidence is useful.

Source confidence

Code-backed: the active version-scoped workspace, settings, Run modal, List segments, Dashboard, Arena, and Compare routes define the current evaluation model and capability fences.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context