{"query":"Benchmark Datasets","corpusVersion":"local","generatedAt":"2026-09-13T04:39:51.423Z","results":[{"blockId":"benchmark-datasets.overview#benchmark-datasets","pageId":"benchmark-datasets.overview","title":"Benchmark Datasets","pageTitle":"Benchmark Datasets","url":"https://teammately.ai/docs/benchmark-datasets.md","humanUrl":"https://teammately.ai/docs/benchmark-datasets#benchmark-datasets","markdownUrl":"https://teammately.ai/docs/benchmark-datasets.md","sectionId":"benchmark-datasets","kind":"concept","productArea":"benchmark_datasets","score":1074.930698335662,"reasons":["search_match","title_match","display_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Benchmark Datasets\n\nBenchmark Datasets defines the evidence set for one benchmark through **Cases**, **Representation**, and **Snapshots**.\n\nThe current dataset is editable. It selects reusable project Cases and reflects current facet, policy, rubric, and contributor facts. A Snapshot freezes the exact dataset state needed by a Benchmark Version and its evaluations. These are deliberately different surfaces: editing the current set must not rewrite historical evidence."},{"blockId":"benchmark-datasets.cases#benchmark-dataset-cases","pageId":"benchmark-datasets.cases","title":"Benchmark Dataset Cases","pageTitle":"Benchmark Dataset Cases","url":"https://teammately.ai/docs/benchmark-datasets/cases.md","humanUrl":"https://teammately.ai/docs/benchmark-datasets/cases#benchmark-dataset-cases","markdownUrl":"https://teammately.ai/docs/benchmark-datasets/cases.md","sectionId":"benchmark-dataset-cases","kind":"task","productArea":"benchmark_datasets","score":534.3228635096168,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Benchmark Dataset Cases"},{"blockId":"benchmark-datasets.overview#before-creating-evidence","pageId":"benchmark-datasets.overview","title":"Before creating evidence","pageTitle":"Benchmark Datasets","url":"https://teammately.ai/docs/benchmark-datasets.md","humanUrl":"https://teammately.ai/docs/benchmark-datasets#before-creating-evidence","markdownUrl":"https://teammately.ai/docs/benchmark-datasets.md","sectionId":"before-creating-evidence","kind":"concept","productArea":"benchmark_datasets","score":345.48134356059234,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Before creating evidence\n\nCheck Case clarity and schema conformance, then inspect Representation for intended behavior and provenance. Confirm approved eligible evaluator links. Resolve Snapshot blockers and preserve the resulting label, version, content hash, creation time, and Case count.\n\nA Snapshot can be reproducible while still being incomplete as product coverage. Reproducibility answers which evidence was evaluated; Representation and Coverage Management answer whether that evidence supports the intended product claim.\n\n{% example-demo title=\"Example: editable set versus frozen evidence\" %}\nThe current dataset gains four Cases and a corrected Rubric link after an expert Contribution is reconciled. An earlier Run still points to its old Benchmark Version. The operator creates a new Snapshot and Version for the changed set rather than comparing the new candidate against the old Run as though only Harness behavior moved.\n{% /example-demo %}\n\n{% related-card-grid title=\"Dataset workflows\" %}\n- [Manage benchmark Cases](/docs/benchmark-datasets/cases)\n- [Inspect Representation](/docs/benchmark-datasets/representation)\n- [Create and inspect Snapshots](/docs/benchmark-datasets/snapshots)\n- [Manage coverage](/docs/coverage-management)\n{% /related-card-grid %}"},{"blockId":"coverage.plan-benchmark-coverage#plan-benchmark-coverage","pageId":"coverage.plan-benchmark-coverage","title":"Plan Benchmark Coverage","pageTitle":"Plan Benchmark Coverage","url":"https://teammately.ai/docs/coverage-engineering/plan-benchmark-coverage.md","humanUrl":"https://teammately.ai/docs/coverage-engineering/plan-benchmark-coverage#plan-benchmark-coverage","markdownUrl":"https://teammately.ai/docs/coverage-engineering/plan-benchmark-coverage.md","sectionId":"plan-benchmark-coverage","kind":"task","productArea":"coverage_engineering","score":338.26301594647214,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Plan Benchmark Coverage\n\nPlan coverage by applying reusable project facets to one benchmark and comparing the intended behavior space with the selected dataset representation."},{"blockId":"benchmark-datasets.overview#decision-checkpoint","pageId":"benchmark-datasets.overview","title":"Decision checkpoint","pageTitle":"Benchmark Datasets","url":"https://teammately.ai/docs/benchmark-datasets.md","humanUrl":"https://teammately.ai/docs/benchmark-datasets#decision-checkpoint","markdownUrl":"https://teammately.ai/docs/benchmark-datasets.md","sectionId":"decision-checkpoint","kind":"concept","productArea":"benchmark_datasets","score":333.29745529157105,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Decision checkpoint\n\n| Surface | Use it to | Evidence rule |\n| --- | --- | --- |\n| Cases | Inspect and change current benchmark membership | Selection is live until snapshotted |\n| Representation | Find concentration and absence across governed facets | Read distribution together with distinct Case counts |\n| Snapshots | Freeze Cases, evaluator links, and representation facts | Snapshot content is read-only |\n\nCoverage Management acts on gaps found in the dataset. Assets remains the project-level reusable pool. Benchmark Evaluations runs exact Harness Versions against an immutable Benchmark Version rather than an unspecified “current dataset.”"},{"blockId":"benchmark-datasets.overview#evidence-flow","pageId":"benchmark-datasets.overview","title":"Evidence flow","pageTitle":"Benchmark Datasets","url":"https://teammately.ai/docs/benchmark-datasets.md","humanUrl":"https://teammately.ai/docs/benchmark-datasets#evidence-flow","markdownUrl":"https://teammately.ai/docs/benchmark-datasets.md","sectionId":"evidence-flow","kind":"concept","productArea":"benchmark_datasets","score":333.09449045483086,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Evidence flow\n\nCases usually begin in project Assets or materialize through Case Review or Expert Contributions. Selecting them makes them part of the current benchmark dataset. Representation then summarizes the current assignments and evaluator relationships. Snapshot readiness checks whether that state can be frozen. A Snapshot supplies the immutable dataset facts used by a Benchmark Version.\n\nThis flow is one-way for historical evidence. Later edits to an Asset, facet assignment, Policy, Rubric, or current membership may improve the next Snapshot, but they do not update a previous Snapshot. Compare candidates within one Benchmark Version unless the analysis explicitly accounts for a moved evidence boundary."},{"blockId":"benchmark-datasets.overview#source-confidence","pageId":"benchmark-datasets.overview","title":"Source confidence","pageTitle":"Benchmark Datasets","url":"https://teammately.ai/docs/benchmark-datasets.md","humanUrl":"https://teammately.ai/docs/benchmark-datasets#source-confidence","markdownUrl":"https://teammately.ai/docs/benchmark-datasets.md","sectionId":"source-confidence","kind":"concept","productArea":"benchmark_datasets","score":329.78770922911286,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Source confidence\n\nCode-backed: the active dataset routes establish the editable current set, representation workspace, and immutable Snapshot boundary."},{"blockId":"object-model.overview#benchmark-artifacts","pageId":"object-model.overview","title":"Benchmark artifacts","pageTitle":"Object model","url":"https://teammately.ai/docs/object-model.md","humanUrl":"https://teammately.ai/docs/object-model#benchmark-artifacts","markdownUrl":"https://teammately.ai/docs/object-model.md","sectionId":"benchmark-artifacts","kind":"reference","productArea":"reference","score":290.60751184366217,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"### Benchmark artifacts\n\n- **Dataset snapshot:** A reproducible selection and representation of benchmark Cases.\n- **Coverage Story:** Benchmark-scoped intent that connects coverage structure to concrete case work.\n- **Expert Contribution:** A benchmark-scoped request containing Tasks, context, statuses, and optional Checkpoints.\n- **Contributed artifact:** A policy, Rubric, Case, or coverage observation supplied through a Contribution with attributable provenance.\n- **Benchmark version:** The fixed evaluation boundary used by Runs and Improvement Sessions.\n- **Evaluation Run:** One execution with candidate, benchmark, response, Rubric outcomes, settings, mapping, and metadata identity.\n- **Improvement Session:** A goal-directed candidate exploration process with pinned evidence, receipts, trajectories, and frontier state.\n\n{% artifact-map title=\"How correctness artifacts connect\" %}\n{% /artifact-map %}"}]}