{"query":"Arena and Rankings","corpusVersion":"local","generatedAt":"2026-09-13T04:41:03.945Z","results":[{"blockId":"benchmark-evaluations.arena-rankings#arena-and-rankings","pageId":"benchmark-evaluations.arena-rankings","title":"Arena and Rankings","pageTitle":"Arena and Rankings","url":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings#arena-and-rankings","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","sectionId":"arena-and-rankings","kind":"task","productArea":"benchmark_evaluations","score":2005.146487324227,"reasons":["search_match","title_match","display_title_match","term_match"],"markdown":"# Arena and Rankings"},{"blockId":"benchmark-evaluations.arena-rankings#use-arena","pageId":"benchmark-evaluations.arena-rankings","title":"Use Arena","pageTitle":"Arena and Rankings","url":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings#use-arena","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","sectionId":"use-arena","kind":"task","productArea":"benchmark_evaluations","score":1219.4179948531246,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Use Arena\n\n1. Confirm both Harness Versions and the Benchmark Version.\n2. Choose the metric family that matches the decision. Required-Policy evidence should not be hidden behind overall performance.\n3. Check comparable and incomplete pair counts before reading the direction.\n4. Inspect only-A and only-B rows to locate tradeoffs. Shared failures identify work neither candidate solves.\n5. Move to Compare or List when the pair summary needs Case, Rubric, or Coverage Facet explanation.\n\nArena does not conduct a new subjective preference interview and does not expose private trajectories. It computes pair evidence from the admitted evaluation results."},{"blockId":"benchmark-evaluations.overview#rankings-and-repeated-sampling","pageId":"benchmark-evaluations.overview","title":"Rankings and repeated sampling","pageTitle":"Benchmark Evaluations","url":"https://teammately.ai/docs/benchmark-evaluations.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations#rankings-and-repeated-sampling","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations.md","sectionId":"rankings-and-repeated-sampling","kind":"concept","productArea":"benchmark_evaluations","score":883.1972323463709,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Rankings and repeated sampling\n\nDashboard aggregates compatible observed Runs for each saved Harness Version across launches. Average score weights Runs equally. Supported binary views report passed at least once or passed every time over the observed case outcomes. Counts and missing evidence are shown; unequal counts do not prevent comparison. Historical group metrics retain their recorded meanings.\n\nRanking is a routing signal. A candidate can lead overall while failing required Policy or high-impact Rubric evidence. Use Arena or Compare to locate the disagreement and List to confirm completeness before starting Improve work."},{"blockId":"benchmark-evaluations.arena-rankings#object-and-state-changes","pageId":"benchmark-evaluations.arena-rankings","title":"Object and state changes","pageTitle":"Arena and Rankings","url":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings#object-and-state-changes","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","sectionId":"object-and-state-changes","kind":"task","productArea":"benchmark_evaluations","score":805.5765324408388,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Object and state changes\n\nArena and leaderboard controls read existing evidence. They do not run candidates, approve a winner, or change frontier retention. A follow-up Improve Session is a separate object."},{"blockId":"benchmark-evaluations.arena-rankings#source-confidence","pageId":"benchmark-evaluations.arena-rankings","title":"Source confidence","pageTitle":"Arena and Rankings","url":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings#source-confidence","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","sectionId":"source-confidence","kind":"task","productArea":"benchmark_evaluations","score":603.5596792844007,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Source confidence\n\nCode-backed: the active Arena route, scoreboard, and leaderboard model define the pair metrics, disagreement counts, repeated-sampling summaries, uncertainty, and telemetry presentation."},{"blockId":"benchmark-evaluations.arena-rankings#prerequisites","pageId":"benchmark-evaluations.arena-rankings","title":"Prerequisites","pageTitle":"Arena and Rankings","url":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings#prerequisites","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","sectionId":"prerequisites","kind":"task","productArea":"benchmark_evaluations","score":599.3362626241424,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Prerequisites\n\n- At least two Harness Versions with comparable results for one Benchmark Version.\n- Enough complete pairs to interpret the selected metric.\n\nArena explains pairwise candidate movement. It summarizes overall, required-Policy, preferred-Policy, Case, Rubric, and coverage metrics, then reports disagreement counts such as only A passed, only B passed, shared failures, incomplete pairs, and total comparable pairs."},{"blockId":"benchmark-evaluations.arena-rankings#read-leaderboard-metrics","pageId":"benchmark-evaluations.arena-rankings","title":"Read leaderboard metrics","pageTitle":"Arena and Rankings","url":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings#read-leaderboard-metrics","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","sectionId":"read-leaderboard-metrics","kind":"task","productArea":"benchmark_evaluations","score":594.7051824375335,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Read leaderboard metrics\n\nArena and Dashboard summarize observed Runs across launches of each saved Harness Version. Average score gives each evaluated Run equal weight. Passed at least once and passed every time summarize observed binary case outcomes where supported. These are descriptions of the collected evidence, not estimates of guaranteed future success. Counts may differ, and the notice about unequal evidence does not block comparison.\n\nUncertainty such as a Wilson interval communicates the limits of the observed sample. A small lead with overlapping uncertainty and many incomplete pairs is not a robust decision. Resource telemetry can add cost, token, and latency context when captured, but missing values remain unknown.\n\n{% example-demo title=\"Example: reliability tradeoff\" %}\nHarness A has three observed Runs and B has one. A passes more Cases at least once, while B passes more Cases in every observed Run. The team inspects the unequal evidence counts and individual results before deciding whether another launch would help.\n{% /example-demo %}"},{"blockId":"benchmark-evaluations.arena-rankings#success-criteria","pageId":"benchmark-evaluations.arena-rankings","title":"Success criteria","pageTitle":"Arena and Rankings","url":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings#success-criteria","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/arena-and-rankings.md","sectionId":"success-criteria","kind":"task","productArea":"benchmark_evaluations","score":577.8715250731142,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Success criteria\n\n- Metric family, pair count, incomplete count, and uncertainty are reported.\n- Only-A, only-B, and shared failures guide concrete inspection.\n- Observed-run metrics have explicit labels, Run counts, and coverage. Historical group-specific pass@n and pass^n remain distinguishable."}]}