# Benchmark Dataset Cases Generated: 2026-09-13T04:36:17.825Z Source build: local Canonical docs: https://teammately.ai/docs --- id: benchmark-datasets.cases title: Benchmark Dataset Cases summary: Inspect benchmark Case membership, coverage traces, references, and scoped bulk actions. kind: task product_area: benchmark_datasets status: stable updated: 2026-08-22 canonical: /docs/benchmark-datasets/cases --- # Benchmark Dataset Cases ## Prerequisites - A selected benchmark and permission to inspect or manage its current dataset. - Project Cases that conform to the intended Input Schema. The Cases tab is the benchmark-scoped view of the current editable case set. It shows Case content and membership together with coverage trace, output or reference mapping, and evaluator relationships. Select one or more rows to request an Expert Contribution, create another benchmark from the selection, remove the Cases from the current benchmark, or download them. Removal changes current membership; it does not delete the reusable Case from project Assets or mutate an existing Snapshot. ## Review before snapshotting 1. Confirm each Case still conforms to Project Input Schema and has the intended materials. 2. Inspect Coverage Facet assignments and source or contributor provenance. 3. Check policy and rubric application, including whether eligible evaluator links are approved. 4. Resolve missing or ambiguous output/reference mapping when the workflow requires reference outputs. 5. Use Representation to check whether the set supports the intended claim. > Membership is not evidence yet > > The editable Cases tab can change. Use a Dataset Snapshot and Benchmark Version when an evaluation, comparison, or Improvement Session must remain reproducible. ## After changing membership Open Representation and confirm that the change affected the intended facet or evaluator population. Removing redundant Cases can improve balance even when total Case count falls. Adding many near-duplicates can increase count without adding meaningful coverage. If a selected Case needs content correction, edit it through the owning Case workflow and review every future benchmark that selects it. Existing Snapshots stay unchanged. If the Case reveals an unclear standard, request an Expert Contribution before compensating with more examples. {% example-demo title="Example: scoped bulk action" %} An operator selects five Cases tied to an unresolved exception and requests one Expert Contribution. The Cases remain in the current set while the expert works. After the controlling Rubric is clarified, the team reviews membership and creates a new Snapshot with the approved evaluator links. {% /example-demo %} ## Object and state changes Selected-row removal changes current benchmark membership; creating another benchmark creates a separate benchmark; requesting a Contribution creates scoped expert work. Downloads and inspection are read-only. No action here mutates an existing Snapshot. ## Success criteria - Current membership, Case identity, coverage trace, and evaluator relationships are understood. - Any bulk action affects only the intended selected Cases. - A new Snapshot is created when changed membership must become evaluation evidence. ## Common failure modes - Treating removal from the benchmark as project-level Case deletion. - Assuming editable membership changed an old Benchmark Version. - Selecting Cases by visible text while ignoring their durable IDs. ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Benchmark Datasets](/docs/benchmark-datasets) - [Cases](/docs/assets/cases) - [Project Input Schema](/docs/project-settings/input-schema) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Dataset upload](/docs/troubleshooting/dataset-upload) - [Unclear Cases](/docs/troubleshooting/unclear-cases) - [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage) {% /related-card-grid %} ## Source confidence Code-backed: the active Cases route defines the benchmark membership table, coverage trace, selected-row operations, downloads, and output/reference presentation. --- id: benchmark-datasets.overview title: Benchmark Datasets summary: Select benchmark Cases, inspect representation, and freeze immutable Snapshots for reproducible evidence. kind: concept product_area: benchmark_datasets status: stable updated: 2026-08-22 canonical: /docs/benchmark-datasets --- # Benchmark Datasets Benchmark Datasets defines the evidence set for one benchmark through **Cases**, **Representation**, and **Snapshots**. The current dataset is editable. It selects reusable project Cases and reflects current facet, policy, rubric, and contributor facts. A Snapshot freezes the exact dataset state needed by a Benchmark Version and its evaluations. These are deliberately different surfaces: editing the current set must not rewrite historical evidence. ## Decision checkpoint | Surface | Use it to | Evidence rule | | --- | --- | --- | | Cases | Inspect and change current benchmark membership | Selection is live until snapshotted | | Representation | Find concentration and absence across governed facets | Read distribution together with distinct Case counts | | Snapshots | Freeze Cases, evaluator links, and representation facts | Snapshot content is read-only | Coverage Management acts on gaps found in the dataset. Assets remains the project-level reusable pool. Benchmark Evaluations runs exact Harness Versions against an immutable Benchmark Version rather than an unspecified “current dataset.” ## Evidence flow Cases usually begin in project Assets or materialize through Case Review or Expert Contributions. Selecting them makes them part of the current benchmark dataset. Representation then summarizes the current assignments and evaluator relationships. Snapshot readiness checks whether that state can be frozen. A Snapshot supplies the immutable dataset facts used by a Benchmark Version. This flow is one-way for historical evidence. Later edits to an Asset, facet assignment, Policy, Rubric, or current membership may improve the next Snapshot, but they do not update a previous Snapshot. Compare candidates within one Benchmark Version unless the analysis explicitly accounts for a moved evidence boundary. ## Before creating evidence Check Case clarity and schema conformance, then inspect Representation for intended behavior and provenance. Confirm approved eligible evaluator links. Resolve Snapshot blockers and preserve the resulting label, version, content hash, creation time, and Case count. A Snapshot can be reproducible while still being incomplete as product coverage. Reproducibility answers which evidence was evaluated; Representation and Coverage Management answer whether that evidence supports the intended product claim. {% example-demo title="Example: editable set versus frozen evidence" %} The current dataset gains four Cases and a corrected Rubric link after an expert Contribution is reconciled. An earlier Run still points to its old Benchmark Version. The operator creates a new Snapshot and Version for the changed set rather than comparing the new candidate against the old Run as though only Harness behavior moved. {% /example-demo %} {% related-card-grid title="Dataset workflows" %} - [Manage benchmark Cases](/docs/benchmark-datasets/cases) - [Inspect Representation](/docs/benchmark-datasets/representation) - [Create and inspect Snapshots](/docs/benchmark-datasets/snapshots) - [Manage coverage](/docs/coverage-management) {% /related-card-grid %} ## Source confidence Code-backed: the active dataset routes establish the editable current set, representation workspace, and immutable Snapshot boundary. --- id: benchmark-datasets.representation title: Dataset Representation summary: Analyze how distinct benchmark Cases are distributed across facets, evaluator rules, and provenance. kind: task product_area: benchmark_datasets status: stable updated: 2026-08-22 canonical: /docs/benchmark-datasets/representation --- # Dataset Representation ## Prerequisites - A current benchmark dataset or Snapshot with representation facts. - Coverage Facets and evaluator relationships meaningful enough to interpret. Representation groups the current or snapshotted dataset by governed facts. Available groupings include Dimension ontology values, Topic Groups, Project Topics, Case Construction Patterns, Policies, policy application, Rubrics, rubric application, presence of rubrics, and contributors. Choose **distinct Cases** when counts matter, or **Case share** when comparing proportions. Policy and rubric views can split by application state. Filters and drilldowns narrow the visible population, and the resulting table or chart can be exported as CSV. ## Reading the view - A large bar means concentration, not correctness. - An empty category can indicate a true coverage gap, an inactive facet, missing classification, or a filter that excludes the Cases. - Topic Groups do not merge their member Topics; group-level handling and Topic-level representation remain distinct. - Policy and rubric presence is not the same as approved eligible application. - Contributor distribution is provenance evidence, not a substitute for agreement or evaluator quality. Use Coverage Management when a gap should drive a Coverage Story or Case Foundry work. Use Expert Contributions when the missing evidence requires governed expert judgment. > Historical availability > > Representation is preserved when the Snapshot contains the required representation facts. Some older Snapshots may not expose this view; do not reconstruct their distribution from current mutable classifications. {% example-demo title="Example: count and share tell different stories" %} A Topic Group has twenty Cases but represents 60% of a small dataset, while a required ontology value has only two. Distinct count reveals the thin required value; Case share reveals the concentration. The operator records a Coverage Story instead of presenting the large Topic count as balanced coverage. {% /example-demo %} ## Object and state changes Grouping, metrics, filtering, splitting, drilldown, and CSV export change only the analysis view. They do not classify Cases, edit facets, or modify Snapshot content. ## Success criteria - Counts and shares use the intended Case population. - Missing, thin, and concentrated categories are distinguished. - A governed Coverage Story or follow-up owns any actionable gap. ## Common failure modes - Reading a filtered percentage as the whole dataset. - Equating high volume with representative coverage. - Reconstructing an old Snapshot from current classifications. ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Coverage Dimensions and ontology](/docs/coverage-engineering/dimensions-ontology) - [Project Topics](/docs/coverage-engineering/project-topics) - [Case Construction Patterns](/docs/coverage-engineering/case-construction-patterns) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage) - [Stale Dimensions](/docs/troubleshooting/stale-dimensions) - [Dimension classification](/docs/troubleshooting/dimension-classification) {% /related-card-grid %} ## Source confidence Code-backed: the active Representation route defines grouping, split, metric, filtering, drilldown, chart/table, and CSV behavior. --- id: assets.cases title: Cases summary: Understand canonical project cases, their input and materials, and how they become members of benchmark datasets. kind: reference product_area: assets status: stable updated: 2026-08-22 canonical: /docs/assets/cases --- # Cases ## Definition A Case is a project-level situation used for expert contribution, benchmark coverage, or candidate evaluation. It has canonical input content and may include declared supporting materials. Cases live in the Assets pool and can be selected into one or more benchmark datasets. The Project Input Schema determines how the primary input and materials are represented. A benchmark snapshot determines which selected cases belong to one versioned evidence boundary. ## Fields, states, or lifecycle rules - Canonical primary input is stored under `content.input`. - Optional supporting values or artifacts are stored under `content.case_materials` according to the project's declared keys. - `record_content.case_view` is a rendered projection used for inspection and delivery; it is not a second editable payload. - Inputs can use plain-text, chat, or structured architecture as configured by the project. - Materials can include admitted artifact families and must satisfy any required-field and file-extension rules. - A project Case is not automatically part of every benchmark. Benchmark Datasets owns selection and snapshots. - Generated or imported Cases should be reviewed for clarity, source authority, and schema conformance before they are trusted as benchmark evidence. - Static case materials and executable Worlds remain separate. A document supplied to a Harness does not become a world merely because it affects execution. ## Case identity and change Treat the persisted case identity as opaque. Do not construct IDs in client code or documentation. When case content changes materially, benchmark interpretation must use a snapshot or version boundary that makes the selected content clear. Responses produced by a Harness are evaluation outputs attached to a Run. They are not the primary case input. Expert-authored acceptable examples can inform standards or contribution work, but the current evaluation contract should remain explicit about which candidate produced each response. {% example-demo title="Example: multimaterial case" %} A chat case asks whether an exception applies. Its required `current_policy` PDF and optional `account_history` table are stored as case materials admitted by Project Input Schema. The rendered case view presents the conversation and both materials. A benchmark snapshot selects the case, and a Run records the evaluated Harness response separately. {% /example-demo %} ## Source confidence Code-backed: the active Assets Cases route and case-content services define canonical input, case materials, and the rendered case view. Public import or export APIs are outside this reference unless separately documented. ## Related task pages {% related-card-grid title="Related task pages" %} - [Configure Project Input Schema](/docs/project-settings/input-schema) - [Work with Benchmark Datasets](/docs/benchmark-datasets) - [Manage benchmark coverage](/docs/coverage-management) {% /related-card-grid %} --- id: expert-contributions.request title: Request an Expert Contribution summary: Create a focused benchmark contribution with an accountable expert, clear objectives, selected cases, attachments, and appropriate task components. kind: task product_area: expert_contributions status: stable updated: 2026-09-07 canonical: /docs/expert-contributions/request-contribution --- # Request an Expert Contribution Request a Contribution when a benchmark needs a bounded piece of specialist judgment. The request should make the expert's decision clear, prepare the relevant evidence, and choose only the task components needed to obtain an attributable answer. ## Prerequisites - A selected project and benchmark. - An expert eligible for the contribution domain. - A concrete contribution statement or unresolved correctness question. - Selected cases, attachments, or scoped statements when the question depends on them. - Project Context and Reference Materials prepared in Agent Setup; use **Assets → Review Screens** when the Contribution needs reusable expert-facing presentation. ## Steps 1. Open the benchmark and select **Expert Contributions → Contributions**. 2. Choose **Request Contribution**. 3. Complete **Objectives & Missions**. State the decision or knowledge the benchmark needs and select the application domain: Coverage Model, Benchmark Setup, or Evaluation Validation. 4. Complete **Choose Experts** and confirm that each selected expert has the right authority for the mission. 5. Complete **Contribution Components**. Available components are Curation, Comparative, Trajectory, Form, Chat, and Interview. Choose conservative, balanced, or exploratory agent behavior; Comparative accepts two to five candidates and can allow improvement. 6. Designate the relevant Cases. Select exact Case IDs and decide whether the contribution may add Cases beyond that set. 7. Add attachments and scoped statements only when they help resolve the mission. Supported attachment scopes include completed Contributions, Policies, Rubrics, Dimensions, ontology values, Project Topics or Groups, Construction Patterns, and Case candidates. 8. Review the captured attachment snapshot version and hash, generated activities, and checkpoints. Confirm that controlling evidence is frozen and consequential meaning will be reconciled. 9. Send the request and follow its state through Overview, Contributions, or Logs & Status. ## Object and state changes This task creates a benchmark-scoped Contribution, associates experts, and records missions, application domain, Case designation, attachments, scoped statements, component behavior, and improvement permission. Planning materializes activities such as Case Review, Form, Chat, and Interview. Sending or starting work moves the Contribution toward `READY` or `IN_PROGRESS`; cancellation preserves the record. ## Success criteria - The Contribution asks one coherent specialist question. - The selected expert and application domain are appropriate. - Every case or attachment is relevant to the objective. - The chosen task types match the judgment required. - Checkpoints protect decisions that should not be silently inferred. - Attachment identities, scope statements, snapshot version, and hash are visible. - The administrator can tell what artifacts may result and where they will be governed. ## Common failure modes - Asking for general review without a materializable objective. - Selecting many cases that do not illuminate the same decision. - Leaving one-time behavior directions outside the Contribution objective, components, or scoped statements. - Omitting the source or case material needed to explain a judgment. - Assuming that task completion automatically approves contributed policies or rubrics. {% example-demo title="Example: focused coverage contribution" %} The objective asks an expert to decide whether source-freshness and customer-impact should form a distinct coverage slice. The operator selects six cases spanning those facets, attaches the controlling policy, and chooses case review plus a final checkpoint. The request can yield a coverage observation and a scoped rubric without asking the expert to redesign the entire benchmark. {% /example-demo %} ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Expert Contributions](/docs/expert-contributions) - [Contributed Artifacts](/docs/expert-contributions/contributed-artifacts) - [Review Screen](/docs/assets/review-screens) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Expert Contribution problems](/docs/troubleshooting/expert-contributions) - [Permissions](/docs/troubleshooting/permissions) - [Low expert agreement](/docs/troubleshooting/low-expert-agreement) {% /related-card-grid %} ## Source confidence Code-backed: the active Contribution composer defines expert selection, objectives, cases, attachments, statements, settings, and contribution components. Exact available components can depend on project and benchmark context.