{"query":"Benchmark Evaluations","corpusVersion":"local","generatedAt":"2026-09-13T04:40:14.547Z","results":[{"blockId":"benchmark-evaluations.overview#benchmark-evaluations","pageId":"benchmark-evaluations.overview","title":"Benchmark Evaluations","pageTitle":"Benchmark Evaluations","url":"https://teammately.ai/docs/benchmark-evaluations.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations#benchmark-evaluations","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations.md","sectionId":"benchmark-evaluations","kind":"concept","productArea":"benchmark_evaluations","score":1051.7421287153168,"reasons":["search_match","title_match","display_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Benchmark Evaluations\n\nBenchmark Evaluations is the version-scoped workspace for executing and comparing candidate systems. The active top-level tabs are **Dashboard**, **List**, **Arena**, and **Compare**. Every managed Run binds an exact saved Harness Version to the immutable Benchmark Version shown in the route.\n\n> Evaluation boundary\n>\n> Interpret evidence inside its recorded Benchmark Version, Harness Version, Run or Run Group, evaluator set, and metadata. Run counts belong to launches. Additional launches add evidence without rewriting earlier Runs."},{"blockId":"benchmark-evaluations.run#run-a-benchmark-evaluation","pageId":"benchmark-evaluations.run","title":"Run a Benchmark Evaluation","pageTitle":"Run a Benchmark Evaluation","url":"https://teammately.ai/docs/benchmark-evaluations/run-evaluation.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/run-evaluation#run-a-benchmark-evaluation","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/run-evaluation.md","sectionId":"run-a-benchmark-evaluation","kind":"task","productArea":"benchmark_evaluations","score":493.2791623697115,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Run a Benchmark Evaluation\n\nLaunch a managed evaluation when the immutable Benchmark Version, governed evaluators, and candidate runtimes are ready."},{"blockId":"object-model.benchmarks#benchmarks","pageId":"object-model.benchmarks","title":"Benchmarks","pageTitle":"Benchmarks","url":"https://teammately.ai/docs/object-model/benchmarks.md","humanUrl":"https://teammately.ai/docs/object-model/benchmarks#benchmarks","markdownUrl":"https://teammately.ai/docs/object-model/benchmarks.md","sectionId":"benchmarks","kind":"reference","productArea":"object_model","score":440.2588243146173,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Benchmarks"},{"blockId":"benchmark-datasets.overview#benchmark-datasets","pageId":"benchmark-datasets.overview","title":"Benchmark Datasets","pageTitle":"Benchmark Datasets","url":"https://teammately.ai/docs/benchmark-datasets.md","humanUrl":"https://teammately.ai/docs/benchmark-datasets#benchmark-datasets","markdownUrl":"https://teammately.ai/docs/benchmark-datasets.md","sectionId":"benchmark-datasets","kind":"concept","productArea":"benchmark_datasets","score":393.51017044852154,"reasons":["search_match","term_match"],"markdown":"# Benchmark Datasets\n\nBenchmark Datasets defines the evidence set for one benchmark through **Cases**, **Representation**, and **Snapshots**.\n\nThe current dataset is editable. It selects reusable project Cases and reflects current facet, policy, rubric, and contributor facts. A Snapshot freezes the exact dataset state needed by a Benchmark Version and its evaluations. These are deliberately different surfaces: editing the current set must not rewrite historical evidence."},{"blockId":"benchmark-evaluations.run-metadata#benchmark-run-metadata","pageId":"benchmark-evaluations.run-metadata","title":"Benchmark Run Metadata","pageTitle":"Benchmark Run Metadata","url":"https://teammately.ai/docs/benchmark-evaluations/run-metadata.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/run-metadata#benchmark-run-metadata","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/run-metadata.md","sectionId":"benchmark-run-metadata","kind":"reference","productArea":"benchmark_evaluations","score":377.34106623268553,"reasons":["search_match","term_match"],"markdown":"# Benchmark Run Metadata"},{"blockId":"object-model.overview#benchmark-artifacts","pageId":"object-model.overview","title":"Benchmark artifacts","pageTitle":"Object model","url":"https://teammately.ai/docs/object-model.md","humanUrl":"https://teammately.ai/docs/object-model#benchmark-artifacts","markdownUrl":"https://teammately.ai/docs/object-model.md","sectionId":"benchmark-artifacts","kind":"reference","productArea":"reference","score":337.13502139522893,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"### Benchmark artifacts\n\n- **Dataset snapshot:** A reproducible selection and representation of benchmark Cases.\n- **Coverage Story:** Benchmark-scoped intent that connects coverage structure to concrete case work.\n- **Expert Contribution:** A benchmark-scoped request containing Tasks, context, statuses, and optional Checkpoints.\n- **Contributed artifact:** A policy, Rubric, Case, or coverage observation supplied through a Contribution with attributable provenance.\n- **Benchmark version:** The fixed evaluation boundary used by Runs and Improvement Sessions.\n- **Evaluation Run:** One execution with candidate, benchmark, response, Rubric outcomes, settings, mapping, and metadata identity.\n- **Improvement Session:** A goal-directed candidate exploration process with pinned evidence, receipts, trajectories, and frontier state.\n\n{% artifact-map title=\"How correctness artifacts connect\" %}\n{% /artifact-map %}"},{"blockId":"benchmark-evaluations.overview#related-workflows","pageId":"benchmark-evaluations.overview","title":"Related workflows","pageTitle":"Benchmark Evaluations","url":"https://teammately.ai/docs/benchmark-evaluations.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations#related-workflows","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations.md","sectionId":"related-workflows","kind":"concept","productArea":"benchmark_evaluations","score":329.95730443609204,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Related workflows\n\n{% related-card-grid title=\"Related workflows\" %}\n- [Configure evaluation execution](/docs/benchmark-evaluations/execution-settings)\n- [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation)\n- [Inspect evaluation results](/docs/benchmark-evaluations/inspect-results)\n- [Use Arena and rankings](/docs/benchmark-evaluations/arena-and-rankings)\n- [Compare Harness Versions](/docs/benchmark-evaluations/compare)\n- [Map external outputs](/docs/benchmark-evaluations/output-mapping)\n{% /related-card-grid %}"},{"blockId":"benchmark-evaluations.overview#surfaces-and-objects","pageId":"benchmark-evaluations.overview","title":"Surfaces and objects","pageTitle":"Benchmark Evaluations","url":"https://teammately.ai/docs/benchmark-evaluations.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations#surfaces-and-objects","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations.md","sectionId":"surfaces-and-objects","kind":"concept","productArea":"benchmark_evaluations","score":319.93973243807034,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Surfaces and objects\n\nDashboard summarizes progress, leaderboards, rank progression across Runs, and available resource telemetry. List is segmented into **Runs**, **Evaluation results**, and **Traces / Spans**. The results segment summarizes Case outcomes and Policy or Rubric failures. Arena compares candidate pairs across governed metrics. Compare is a symmetric matrix of Harness Versions across selected evidence rows.\n\nA Run Group can collect one standard attempt or repeated attempts. A Run records one candidate execution and its per-Case progress. Evaluation results record the admitted Policy and Rubric outcomes. Costs, tokens, and latency are telemetry only when the provider or execution path captured them.\n\n> Traces / Spans capability fence\n>\n> The List navigation exposes Traces / Spans, but the current benchmark API does not expose evaluation execution traces. Do not claim that trajectories, spans, private reasoning, or tool traces can be inspected from Benchmark Evaluations today."}]}