Teammately Docs
Docs menu

task

Arena and Rankings

Interpret pairwise candidate disagreement, governed metric families, repeated-sampling ranks, and uncertainty.

Arena and Rankings

Prerequisites

  • At least two Harness Versions with comparable results for one Benchmark Version.
  • Enough complete pairs to interpret the selected metric.

Arena explains pairwise candidate movement. It summarizes overall, required-Policy, preferred-Policy, Case, Rubric, and coverage metrics, then reports disagreement counts such as only A passed, only B passed, shared failures, incomplete pairs, and total comparable pairs.

Use Arena

  1. Confirm both Harness Versions and the Benchmark Version.
  2. Choose the metric family that matches the decision. Required-Policy evidence should not be hidden behind overall performance.
  3. Check comparable and incomplete pair counts before reading the direction.
  4. Inspect only-A and only-B rows to locate tradeoffs. Shared failures identify work neither candidate solves.
  5. Move to Compare or List when the pair summary needs Case, Rubric, or Coverage Facet explanation.

Arena does not conduct a new subjective preference interview and does not expose private trajectories. It computes pair evidence from the admitted evaluation results.

Read leaderboard metrics

Arena and Dashboard summarize observed Runs across launches of each saved Harness Version. Average score gives each evaluated Run equal weight. Passed at least once and passed every time summarize observed binary case outcomes where supported. These are descriptions of the collected evidence, not estimates of guaranteed future success. Counts may differ, and the notice about unequal evidence does not block comparison.

Uncertainty such as a Wilson interval communicates the limits of the observed sample. A small lead with overlapping uncertainty and many incomplete pairs is not a robust decision. Resource telemetry can add cost, token, and latency context when captured, but missing values remain unknown.

Worked example

Example: reliability tradeoff

Harness A has three observed Runs and B has one. A passes more Cases at least once, while B passes more Cases in every observed Run. The team inspects the unequal evidence counts and individual results before deciding whether another launch would help.

Object and state changes

Arena and leaderboard controls read existing evidence. They do not run candidates, approve a winner, or change frontier retention. A follow-up Improve Session is a separate object.

Success criteria

  • Metric family, pair count, incomplete count, and uncertainty are reported.
  • Only-A, only-B, and shared failures guide concrete inspection.
  • Observed-run metrics have explicit labels, Run counts, and coverage. Historical group-specific pass@n and pass^n remain distinguishable.

Common failure modes

  • Hiding required-Policy regressions behind overall rank.
  • Treating overlapping uncertainty as a decisive lead.
  • Equating missing telemetry with zero resource use.

Source confidence

Code-backed: the active Arena route, scoreboard, and leaderboard model define the pair metrics, disagreement counts, repeated-sampling summaries, uncertainty, and telemetry presentation.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context