# Dataset Snapshots Generated: 2026-09-13T04:41:57.043Z Source build: local Canonical docs: https://teammately.ai/docs --- id: benchmark-datasets.snapshots title: Dataset Snapshots summary: Freeze Cases, evaluator links, and representation facts as an immutable benchmark evidence boundary. kind: task product_area: benchmark_datasets status: stable updated: 2026-08-22 canonical: /docs/benchmark-datasets/snapshots --- # Dataset Snapshots ## Prerequisites - A reviewed current Case set. - Approved eligible evaluator links and no Snapshot readiness blockers. - Permission to create benchmark evidence. A Dataset Snapshot freezes the benchmark's selected Cases, eligible evaluator links, and representation facts at a point in time. The live dataset remains editable; the Snapshot opens read-only **Cases** and **Representation** views. ## Create a Snapshot The readiness check reports Case count, approved eligible Policy and Rubric counts, and blockers. Resolve every blocker before creation. Record a meaningful Snapshot label, then verify the displayed version, content hash, creation time, and Case count. Creation does not make weak input trustworthy. Review Case clarity, coverage, materials, and evaluator applicability first. After creation, do not describe later mutable classifications or links as if they were part of the frozen state. ## Evidence rules - Identify the exact Snapshot or resulting Benchmark Version in every Run and comparison. - Create a new Snapshot when Case membership, material content, or admitted evaluator relationships change in a way that affects the claim. - Do not mutate a Snapshot to “fix” historical evidence; correct the live dataset and freeze a new one. - If historical Representation is unavailable, report that limitation instead of substituting current facts. {% example-demo title="Example: preserving a coverage expansion" %} After Case Review adds eight exception-handling Cases, the team verifies approved rubric links and creates a new Snapshot. Runs against the earlier Benchmark Version remain comparable within their old boundary, while new Runs explicitly use the expanded version. {% /example-demo %} ## Object and state changes Creation adds a new immutable Snapshot with its own label, version, hash, time, Case membership, evaluator links, and representation facts. It does not lock or copy edits back into the current dataset. ## Success criteria - Readiness has no blockers. - Identity fields and Case count match the intended boundary. - Future Runs cite the resulting exact Benchmark Version. ## Common failure modes - Snapshotting weak or invalid Cases because readiness passes structurally. - Treating current classifications as part of an older Snapshot. - Comparing candidates across moved Snapshot boundaries without disclosure. ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Benchmark Datasets](/docs/benchmark-datasets) - [Benchmark versioning](/docs/governance/benchmark-versioning) - [Reproducibility](/docs/governance/reproducibility) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Dataset upload](/docs/troubleshooting/dataset-upload) - [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly) {% /related-card-grid %} ## Source confidence Code-backed: the active Snapshots route defines readiness, blockers, immutable content, identity fields, and read-only Snapshot inspection. --- id: benchmark-datasets.overview title: Benchmark Datasets summary: Select benchmark Cases, inspect representation, and freeze immutable Snapshots for reproducible evidence. kind: concept product_area: benchmark_datasets status: stable updated: 2026-08-22 canonical: /docs/benchmark-datasets --- # Benchmark Datasets Benchmark Datasets defines the evidence set for one benchmark through **Cases**, **Representation**, and **Snapshots**. The current dataset is editable. It selects reusable project Cases and reflects current facet, policy, rubric, and contributor facts. A Snapshot freezes the exact dataset state needed by a Benchmark Version and its evaluations. These are deliberately different surfaces: editing the current set must not rewrite historical evidence. ## Decision checkpoint | Surface | Use it to | Evidence rule | | --- | --- | --- | | Cases | Inspect and change current benchmark membership | Selection is live until snapshotted | | Representation | Find concentration and absence across governed facets | Read distribution together with distinct Case counts | | Snapshots | Freeze Cases, evaluator links, and representation facts | Snapshot content is read-only | Coverage Management acts on gaps found in the dataset. Assets remains the project-level reusable pool. Benchmark Evaluations runs exact Harness Versions against an immutable Benchmark Version rather than an unspecified “current dataset.” ## Evidence flow Cases usually begin in project Assets or materialize through Case Review or Expert Contributions. Selecting them makes them part of the current benchmark dataset. Representation then summarizes the current assignments and evaluator relationships. Snapshot readiness checks whether that state can be frozen. A Snapshot supplies the immutable dataset facts used by a Benchmark Version. This flow is one-way for historical evidence. Later edits to an Asset, facet assignment, Policy, Rubric, or current membership may improve the next Snapshot, but they do not update a previous Snapshot. Compare candidates within one Benchmark Version unless the analysis explicitly accounts for a moved evidence boundary. ## Before creating evidence Check Case clarity and schema conformance, then inspect Representation for intended behavior and provenance. Confirm approved eligible evaluator links. Resolve Snapshot blockers and preserve the resulting label, version, content hash, creation time, and Case count. A Snapshot can be reproducible while still being incomplete as product coverage. Reproducibility answers which evidence was evaluated; Representation and Coverage Management answer whether that evidence supports the intended product claim. {% example-demo title="Example: editable set versus frozen evidence" %} The current dataset gains four Cases and a corrected Rubric link after an expert Contribution is reconciled. An earlier Run still points to its old Benchmark Version. The operator creates a new Snapshot and Version for the changed set rather than comparing the new candidate against the old Run as though only Harness behavior moved. {% /example-demo %} {% related-card-grid title="Dataset workflows" %} - [Manage benchmark Cases](/docs/benchmark-datasets/cases) - [Inspect Representation](/docs/benchmark-datasets/representation) - [Create and inspect Snapshots](/docs/benchmark-datasets/snapshots) - [Manage coverage](/docs/coverage-management) {% /related-card-grid %} ## Source confidence Code-backed: the active dataset routes establish the editable current set, representation workspace, and immutable Snapshot boundary. --- id: benchmark-evaluations.run title: Run a Benchmark Evaluation summary: Launch exact active Harness Versions against an immutable Benchmark Version as standard or repeated Run Groups. kind: task product_area: benchmark_evaluations status: stable updated: 2026-09-13 canonical: /docs/benchmark-evaluations/run-evaluation --- # Run a Benchmark Evaluation Launch a managed evaluation when the immutable Benchmark Version, governed evaluators, and candidate runtimes are ready. ## Prerequisites - A Benchmark Version backed by the intended Dataset Snapshot. - Approved eligible Policies and Rubrics. - At least one saved project Harness Version. - Prepared Harness runtime and required secret grants. - A chosen number of Runs for each selected Harness. ## Steps 1. Open **Benchmark Evaluations** for the intended Benchmark Version. 2. Open Evaluation Settings if you need to adjust the machine configuration. 3. Start a Run and select one or more offered Harness Versions. Confirm the exact version labels rather than relying on Harness names alone. 4. Choose the number of Runs for each Harness. Counts may differ; review the total execution volume. 5. Supply any requested Run Metadata. Keep credentials out of descriptive fields. 6. Launch. Each selected Harness creates its own Run Group containing the requested independent Runs, including when the count is one. 7. Follow output and evaluation progress. Distinguish queued, running, complete, failed, cancelled, and incomplete work rather than inferring completion from partial scores. 8. Inspect List, Dashboard, Arena, or Compare only after checking which attempts and Cases are evaluable. ## Evidence created The launch creates Run Groups and Runs bound to exact Harness and Benchmark Versions. Per-Case outputs and evaluator outcomes accrue separately, so output completion can precede evaluation completion. Provider telemetry can include tokens, cost, and latency when captured; absence of telemetry is not zero usage. Dashboard aggregates compatible observed Runs across launches. Choose average score, passed at least once, or passed every time where supported. Each Run retains its own outputs and status; inspect the group and individual Runs when work is incomplete. > No Draft execution > > A managed benchmark Run does not evaluate the mutable Harness Draft. Save the candidate and select its exact saved Version when launching. {% example-demo title="Example: two candidates, three attempts" %} Harness Versions 6 and 9 are active with `n=3`. One launch creates two Run Groups and six independent Runs against the same Benchmark Version. If one attempt fails preparation, the group reports incomplete evidence instead of silently treating the remaining two as the configured cohort. {% /example-demo %} ## Object and state changes Launching creates one Run Group per Harness and one or more independent Runs. Outputs, evaluator outcomes, progress, metadata, and telemetry accrue to those records. A later launch creates new evidence and does not overwrite the cohort. ## Success criteria - Exact Harness and Benchmark Versions are recorded. - Each launch group contains the number of Runs requested for that Harness. - Output and evaluation progress reach an interpretable terminal state. - Incomplete or failed attempts remain visible. ## Common failure modes - Selecting the wrong saved Version or assuming Draft execution. - Reading partial evaluation as a complete cohort. - Treating absent telemetry as zero usage. ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Evaluation Execution Settings](/docs/benchmark-evaluations/execution-settings) - [Harnesses](/docs/assets/harnesses) - [Run Metadata](/docs/benchmark-evaluations/run-metadata) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Benchmark runs](/docs/troubleshooting/benchmark-runs) - [Missing outputs](/docs/troubleshooting/missing-outputs) - [Authentication](/docs/troubleshooting/authentication) {% /related-card-grid %} ## Source confidence Code-backed: the current Run modal, Runs workspace, and Run Group route define selection, group creation, repeated attempts, progress, and evidence identity. --- id: governance.reproducibility title: Reproducibility summary: Preserve enough source context to explain and repeat correctness decisions. kind: reference product_area: governance status: stable updated: 2026-08-23 canonical: /docs/governance/reproducibility --- # Reproducibility ## Definition Reproducibility means preserving enough exact identity and observable configuration to explain what was evaluated and to repeat the supported execution path. It does not mean every future execution will produce an identical stochastic output. It means a reader can distinguish changes in candidate, evidence, evaluator, sampling, and runtime metadata instead of attributing every result difference to the model. ## Fields, states, or lifecycle rules - Preserve Project, Benchmark, Benchmark Version, Dataset Snapshot, and Run identity. - Preserve the exact saved Harness Version rather than an editable draft or display label. - Preserve admitted Case, Policy, and Rubric version boundaries through the Benchmark Version. - Record Run Group, attempt, sampling profile, evaluator set, and visible execution settings. - Retain run metadata and measured telemetry when captured; missing values remain unknown. - Record completeness, incomplete Cases, and terminal state beside scores. - Use canonical evaluation receipts for Improvement Session candidate claims. - Do not claim private worker activity, hidden reasoning, infrastructure internals, or unavailable traces as reproducibility evidence. ## Related objects Dataset Snapshots preserve the evidence set. Benchmark Evaluations preserves candidate, Run, settings, results, and available telemetry. Compare and Arena interpret candidates inside compatible evidence boundaries. Improve adds Goal Contract, candidate, and canonical receipt identity when evaluation drives code or Harness changes. {% example-demo title="Diagnosing a score change" %} Two Runs use the same Harness Version but report different pass rates. The operator confirms that one Run used Benchmark Version 6 and the other used Version 7, which added source-conflict Cases and a revised grounding Rubric. The version and completeness record explains the movement. The team avoids filing a candidate regression until it compares Runs inside the same evidence boundary. {% /example-demo %} ## Source confidence Code-backed: Snapshot, Run detail, and run-metadata surfaces expose the immutable evidence boundary, candidate identity, status, counts, settings, and available metadata needed for supported reproducibility. They do not promise deterministic model output or unrestricted execution traces. ## Related task pages {% related-card-grid title="Related task pages" %} - [Benchmark Versioning](/docs/governance/benchmark-versioning) - [Compare Harness Versions](/docs/benchmark-evaluations/compare) - [Benchmark Evaluations](/docs/benchmark-evaluations) - [Product quickstart](/docs/quickstart) - [Task index](/docs/operating-manual/task-index) {% /related-card-grid %}