Teammately Docs
Docs menu

concept

Benchmark Datasets

Select benchmark Cases, inspect representation, and freeze immutable Snapshots for reproducible evidence.

Benchmark Datasets

Benchmark Datasets defines the evidence set for one benchmark through Cases, Representation, and Snapshots.

The current dataset is editable. It selects reusable project Cases and reflects current facet, policy, rubric, and contributor facts. A Snapshot freezes the exact dataset state needed by a Benchmark Version and its evaluations. These are deliberately different surfaces: editing the current set must not rewrite historical evidence.

Decision checkpoint

SurfaceUse it toEvidence rule
CasesInspect and change current benchmark membershipSelection is live until snapshotted
RepresentationFind concentration and absence across governed facetsRead distribution together with distinct Case counts
SnapshotsFreeze Cases, evaluator links, and representation factsSnapshot content is read-only

Coverage Management acts on gaps found in the dataset. Assets remains the project-level reusable pool. Benchmark Evaluations runs exact Harness Versions against an immutable Benchmark Version rather than an unspecified “current dataset.”

Evidence flow

Cases usually begin in project Assets or materialize through Case Review or Expert Contributions. Selecting them makes them part of the current benchmark dataset. Representation then summarizes the current assignments and evaluator relationships. Snapshot readiness checks whether that state can be frozen. A Snapshot supplies the immutable dataset facts used by a Benchmark Version.

This flow is one-way for historical evidence. Later edits to an Asset, facet assignment, Policy, Rubric, or current membership may improve the next Snapshot, but they do not update a previous Snapshot. Compare candidates within one Benchmark Version unless the analysis explicitly accounts for a moved evidence boundary.

Before creating evidence

Check Case clarity and schema conformance, then inspect Representation for intended behavior and provenance. Confirm approved eligible evaluator links. Resolve Snapshot blockers and preserve the resulting label, version, content hash, creation time, and Case count.

A Snapshot can be reproducible while still being incomplete as product coverage. Reproducibility answers which evidence was evaluated; Representation and Coverage Management answer whether that evidence supports the intended product claim.

Worked example

Example: editable set versus frozen evidence

The current dataset gains four Cases and a corrected Rubric link after an expert Contribution is reconciled. An earlier Run still points to its old Benchmark Version. The operator creates a new Snapshot and Version for the changed set rather than comparing the new candidate against the old Run as though only Harness behavior moved.

Source confidence

Code-backed: the active dataset routes establish the editable current set, representation workspace, and immutable Snapshot boundary.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context