# Dataset Representation
Generated: 2026-09-13T04:36:16.624Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: benchmark-datasets.representation
title: Dataset Representation
summary: Analyze how distinct benchmark Cases are distributed across facets, evaluator rules, and provenance.
kind: task
product_area: benchmark_datasets
status: stable
updated: 2026-08-22
canonical: /docs/benchmark-datasets/representation
---
# Dataset Representation
## Prerequisites
- A current benchmark dataset or Snapshot with representation facts.
- Coverage Facets and evaluator relationships meaningful enough to interpret.
Representation groups the current or snapshotted dataset by governed facts. Available groupings include Dimension ontology values, Topic Groups, Project Topics, Case Construction Patterns, Policies, policy application, Rubrics, rubric application, presence of rubrics, and contributors.
Choose **distinct Cases** when counts matter, or **Case share** when comparing proportions. Policy and rubric views can split by application state. Filters and drilldowns narrow the visible population, and the resulting table or chart can be exported as CSV.
## Reading the view
- A large bar means concentration, not correctness.
- An empty category can indicate a true coverage gap, an inactive facet, missing classification, or a filter that excludes the Cases.
- Topic Groups do not merge their member Topics; group-level handling and Topic-level representation remain distinct.
- Policy and rubric presence is not the same as approved eligible application.
- Contributor distribution is provenance evidence, not a substitute for agreement or evaluator quality.
Use Coverage Management when a gap should drive a Coverage Story or Case Foundry work. Use Expert Contributions when the missing evidence requires governed expert judgment.
> Historical availability
>
> Representation is preserved when the Snapshot contains the required representation facts. Some older Snapshots may not expose this view; do not reconstruct their distribution from current mutable classifications.
{% example-demo title="Example: count and share tell different stories" %}
A Topic Group has twenty Cases but represents 60% of a small dataset, while a required ontology value has only two. Distinct count reveals the thin required value; Case share reveals the concentration. The operator records a Coverage Story instead of presenting the large Topic count as balanced coverage.
{% /example-demo %}
## Object and state changes
Grouping, metrics, filtering, splitting, drilldown, and CSV export change only the analysis view. They do not classify Cases, edit facets, or modify Snapshot content.
## Success criteria
- Counts and shares use the intended Case population.
- Missing, thin, and concentrated categories are distinguished.
- A governed Coverage Story or follow-up owns any actionable gap.
## Common failure modes
- Reading a filtered percentage as the whole dataset.
- Equating high volume with representative coverage.
- Reconstructing an old Snapshot from current classifications.
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Coverage Dimensions and ontology](/docs/coverage-engineering/dimensions-ontology)
- [Project Topics](/docs/coverage-engineering/project-topics)
- [Case Construction Patterns](/docs/coverage-engineering/case-construction-patterns)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
- [Stale Dimensions](/docs/troubleshooting/stale-dimensions)
- [Dimension classification](/docs/troubleshooting/dimension-classification)
{% /related-card-grid %}
## Source confidence
Code-backed: the active Representation route defines grouping, split, metric, filtering, drilldown, chart/table, and CSV behavior.
---
id: benchmark-datasets.overview
title: Benchmark Datasets
summary: Select benchmark Cases, inspect representation, and freeze immutable Snapshots for reproducible evidence.
kind: concept
product_area: benchmark_datasets
status: stable
updated: 2026-08-22
canonical: /docs/benchmark-datasets
---
# Benchmark Datasets
Benchmark Datasets defines the evidence set for one benchmark through **Cases**, **Representation**, and **Snapshots**.
The current dataset is editable. It selects reusable project Cases and reflects current facet, policy, rubric, and contributor facts. A Snapshot freezes the exact dataset state needed by a Benchmark Version and its evaluations. These are deliberately different surfaces: editing the current set must not rewrite historical evidence.
## Decision checkpoint
| Surface | Use it to | Evidence rule |
| --- | --- | --- |
| Cases | Inspect and change current benchmark membership | Selection is live until snapshotted |
| Representation | Find concentration and absence across governed facets | Read distribution together with distinct Case counts |
| Snapshots | Freeze Cases, evaluator links, and representation facts | Snapshot content is read-only |
Coverage Management acts on gaps found in the dataset. Assets remains the project-level reusable pool. Benchmark Evaluations runs exact Harness Versions against an immutable Benchmark Version rather than an unspecified “current dataset.”
## Evidence flow
Cases usually begin in project Assets or materialize through Case Review or Expert Contributions. Selecting them makes them part of the current benchmark dataset. Representation then summarizes the current assignments and evaluator relationships. Snapshot readiness checks whether that state can be frozen. A Snapshot supplies the immutable dataset facts used by a Benchmark Version.
This flow is one-way for historical evidence. Later edits to an Asset, facet assignment, Policy, Rubric, or current membership may improve the next Snapshot, but they do not update a previous Snapshot. Compare candidates within one Benchmark Version unless the analysis explicitly accounts for a moved evidence boundary.
## Before creating evidence
Check Case clarity and schema conformance, then inspect Representation for intended behavior and provenance. Confirm approved eligible evaluator links. Resolve Snapshot blockers and preserve the resulting label, version, content hash, creation time, and Case count.
A Snapshot can be reproducible while still being incomplete as product coverage. Reproducibility answers which evidence was evaluated; Representation and Coverage Management answer whether that evidence supports the intended product claim.
{% example-demo title="Example: editable set versus frozen evidence" %}
The current dataset gains four Cases and a corrected Rubric link after an expert Contribution is reconciled. An earlier Run still points to its old Benchmark Version. The operator creates a new Snapshot and Version for the changed set rather than comparing the new candidate against the old Run as though only Harness behavior moved.
{% /example-demo %}
{% related-card-grid title="Dataset workflows" %}
- [Manage benchmark Cases](/docs/benchmark-datasets/cases)
- [Inspect Representation](/docs/benchmark-datasets/representation)
- [Create and inspect Snapshots](/docs/benchmark-datasets/snapshots)
- [Manage coverage](/docs/coverage-management)
{% /related-card-grid %}
## Source confidence
Code-backed: the active dataset routes establish the editable current set, representation workspace, and immutable Snapshot boundary.
---
id: benchmark-datasets.cases
title: Benchmark Dataset Cases
summary: Inspect benchmark Case membership, coverage traces, references, and scoped bulk actions.
kind: task
product_area: benchmark_datasets
status: stable
updated: 2026-08-22
canonical: /docs/benchmark-datasets/cases
---
# Benchmark Dataset Cases
## Prerequisites
- A selected benchmark and permission to inspect or manage its current dataset.
- Project Cases that conform to the intended Input Schema.
The Cases tab is the benchmark-scoped view of the current editable case set. It shows Case content and membership together with coverage trace, output or reference mapping, and evaluator relationships.
Select one or more rows to request an Expert Contribution, create another benchmark from the selection, remove the Cases from the current benchmark, or download them. Removal changes current membership; it does not delete the reusable Case from project Assets or mutate an existing Snapshot.
## Review before snapshotting
1. Confirm each Case still conforms to Project Input Schema and has the intended materials.
2. Inspect Coverage Facet assignments and source or contributor provenance.
3. Check policy and rubric application, including whether eligible evaluator links are approved.
4. Resolve missing or ambiguous output/reference mapping when the workflow requires reference outputs.
5. Use Representation to check whether the set supports the intended claim.
> Membership is not evidence yet
>
> The editable Cases tab can change. Use a Dataset Snapshot and Benchmark Version when an evaluation, comparison, or Improvement Session must remain reproducible.
## After changing membership
Open Representation and confirm that the change affected the intended facet or evaluator population. Removing redundant Cases can improve balance even when total Case count falls. Adding many near-duplicates can increase count without adding meaningful coverage.
If a selected Case needs content correction, edit it through the owning Case workflow and review every future benchmark that selects it. Existing Snapshots stay unchanged. If the Case reveals an unclear standard, request an Expert Contribution before compensating with more examples.
{% example-demo title="Example: scoped bulk action" %}
An operator selects five Cases tied to an unresolved exception and requests one Expert Contribution. The Cases remain in the current set while the expert works. After the controlling Rubric is clarified, the team reviews membership and creates a new Snapshot with the approved evaluator links.
{% /example-demo %}
## Object and state changes
Selected-row removal changes current benchmark membership; creating another benchmark creates a separate benchmark; requesting a Contribution creates scoped expert work. Downloads and inspection are read-only. No action here mutates an existing Snapshot.
## Success criteria
- Current membership, Case identity, coverage trace, and evaluator relationships are understood.
- Any bulk action affects only the intended selected Cases.
- A new Snapshot is created when changed membership must become evaluation evidence.
## Common failure modes
- Treating removal from the benchmark as project-level Case deletion.
- Assuming editable membership changed an old Benchmark Version.
- Selecting Cases by visible text while ignoring their durable IDs.
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Benchmark Datasets](/docs/benchmark-datasets)
- [Cases](/docs/assets/cases)
- [Project Input Schema](/docs/project-settings/input-schema)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Dataset upload](/docs/troubleshooting/dataset-upload)
- [Unclear Cases](/docs/troubleshooting/unclear-cases)
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
{% /related-card-grid %}
## Source confidence
Code-backed: the active Cases route defines the benchmark membership table, coverage trace, selected-row operations, downloads, and output/reference presentation.
---
id: coverage-management.overview
title: Coverage Management
summary: Manage benchmark coverage from setup through representation, Coverage Stories, case review, Case Foundry, and contribution requests.
kind: concept
product_area: coverage_management
status: stable
updated: 2026-09-07
canonical: /docs/coverage-management
---
# Coverage Management
Coverage Management is the benchmark-scoped workspace for deciding whether the current dataset represents the behavior space the benchmark is meant to test. It connects project Coverage Facets to dataset representation, coverage guidance, Coverage Stories, case preparation, and expert contribution requests.
## Definition
Project Coverage Facets define reusable Dimensions, Project Topics, and Case Construction Patterns. Coverage Management applies those foundations to one benchmark. **Get Started** establishes benchmark coverage guidance. The overview shows representation and operational status. Coverage Stories organize meaningful slices or gaps. Case Review inspects prepared cases and materials. Case Foundry coordinates case construction work.
The goal is to make missing or thin behavior explicit before evaluation evidence is trusted. Coverage Management does not replace Benchmark Datasets; it explains and improves the representation of the selected data.
## Decision checkpoint
| Observation | Use | Next durable result |
| --- | --- | --- |
| Benchmark purpose or guidance is missing | Get Started | Saved coverage setup and readiness |
| A facet tuple is thin or absent | Representation and Coverage Stories | Named coverage need and intended evidence |
| More cases are needed | Case Foundry | Bounded construction work tied to the gap |
| Generated cases may be unclear | Case Review | Reviewed case and material quality |
| Specialist judgment is required | Contribution request from coverage context | Benchmark-scoped Expert Contribution |
| Coverage changed materially | Benchmark Datasets | Updated selection and snapshot boundary |
## Coverage Stories and case work
A Coverage Story gives a gap or behavior slice an operational narrative: why it matters, which facet combinations define it, what evidence exists, and what sourcing work remains. It should be concrete enough to guide case construction and expert attention.
Case Foundry can prepare case work from that structure. Case Review checks the resulting inputs and generated materials before they enter trusted dataset evidence. AI assistance can accelerate preparation, but selection and benchmark interpretation remain explicit human and product-state decisions.
## Coverage and correctness
Coverage gaps sometimes reveal missing correctness rather than missing cases. If experts cannot say how a represented situation should be judged, request an Expert Contribution and update policies or rubrics. If the standard is clear but no case exercises it, use Weave and Case Foundry. If cases exist but are not selected or snapshotted, use Benchmark Datasets.
This routing prevents Comparison Directions, Contribution-scoped agent behavior, and coverage structure from being mixed together. Comparison Directions guide response variation; Coverage Facets and Coverage Stories describe the benchmark behavior space.
{% example-demo title="Conflicting-source story" %}
Representation shows that the benchmark covers current-source questions but almost never combines them with a plausible superseded document. A Coverage Story names the conflict pattern, relevant source-freshness and impact facets, and the desired case count. Case Foundry prepares candidates, Case Review rejects unrealistic material, and the accepted cases enter a new dataset snapshot.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Set up benchmark coverage](/docs/coverage-management/get-started)
- [Work with Coverage Stories](/docs/coverage-management/coverage-stories)
- [Run Case Foundry](/docs/coverage-management/case-foundry)
- [Review prepared Cases](/docs/coverage-management/case-review)
- [Plan benchmark coverage](/docs/coverage-engineering/plan-benchmark-coverage)
- [Work with Benchmark Datasets](/docs/benchmark-datasets)
- [Request an Expert Contribution](/docs/expert-contributions/request-contribution)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Coverage Engineering](/docs/coverage-engineering)
- [Cases](/docs/assets/cases)
- [Comparison Directions](/docs/assets/comparison-directions)
{% /related-card-grid %}
## Source confidence
Code-backed: the active benchmark coverage routes expose setup, overview, Coverage Stories, Case Review, Case Foundry integration, realtime state, and contribution-request entry points.
---
id: coverage.dimensions-ontology
title: Dimensions and Ontology
summary: Define reusable behavior axes and their allowed values, then inspect how cases and benchmarks cover them.
kind: reference
product_area: coverage_engineering
status: stable
updated: 2026-09-07
canonical: /docs/coverage-engineering/dimensions-ontology
---
# Dimensions and Ontology
## Definition
Dimensions are reusable project-level axes for describing how cases differ. Each Dimension contains ontology values: the named members used to classify cases and measure representation. A Dimension might be **Source condition**, with values such as **Current**, **Superseded**, **Conflicting**, and **Missing**.
Use **Project Foundations → Coverage Facets → Dimension** to create, generate, inspect, and maintain them.
## Fields, states, or lifecycle rules
### What a Dimension contains
| Element | Purpose |
| --- | --- |
| Name and description | Explain the behavior axis and its boundary |
| Ontology values | Define the values used for classification |
| Examples | Show classified cases and the reason for a value assignment |
| Statistics | Show case-pool and benchmark distribution by ontology value |
| Benchmark focus | Show whether values are required, sampled, diagnostic, or ignored in benchmark setup |
Dimensions and ontology values are project foundations. A benchmark does not copy them. **Get Started** assigns benchmark-specific roles to the project values, and **Representation** reports how the selected cases cover them.

Use the table to compare each Dimension's definition with its ontology and current classification footprint before opening the detail view.
### Create or generate a schema
Create a Dimension manually when the axis and vocabulary are already understood. Use the dimension-schema generator when project context or source material should produce a reviewable proposal. Generated proposals can include a definition, why the Dimension matters, proposed ontology members, and warnings.
A proposal is not the active schema. Review each proposed Dimension and value before accepting it. Avoid accepting near-duplicates simply because they use different wording.
> Classification boundary
>
> Creating or editing a Dimension does not instantly classify every existing case. Missing or stale classifications can be queued and monitored separately. Treat unclassified cases as missing evidence, not as an implicit ontology value.
### Design rules
- Make the Dimension answer one stable question. Split axes that mix several independent concerns.
- Give every ontology value a definition that distinguishes it from neighboring values.
- Prefer values that can be applied consistently to real cases.
- Do not use a Dimension to encode case quality, policy approval, or a desired model score.
- Review distributions after changing values. A clean schema can still leave important cases unclassified.
- Delete only after checking Case Pool and benchmark usage; removal changes the project coverage vocabulary.
### Benchmark roles
In benchmark Get Started, each Dimension and ontology value can receive a focus role:
- **Required:** the benchmark is expected to cover this value.
- **Sampled:** include it as part of the desired mix.
- **Diagnostic:** track it for analysis without making it part of the main denominator.
- **Ignored:** exclude it from the benchmark coverage intention.
- **Unset:** no explicit benchmark instruction has been recorded.
Those roles shape Coverage Story generation and interpretation. They do not alter the project-level definition of the value.
{% example-demo title="Example: source condition" %}
The project defines a Source condition Dimension with Current, Superseded, Conflicting, and Missing values. One benchmark marks all four as required; another marks Current as required and the remaining values as diagnostic. The same project vocabulary supports different benchmark intentions without duplicating the Dimension.
{% /example-demo %}
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Generate a dimension schema](/docs/coverage-engineering/generate-dimension-schema)
- [Work with Project Topics](/docs/coverage-engineering/project-topics)
- [Plan Benchmark Coverage](/docs/coverage-engineering/plan-benchmark-coverage)
- [Analyze dataset representation](/docs/benchmark-datasets/representation)
{% /related-card-grid %}
## Source confidence
Code-backed: the current Dimension list, detail, settings, proposal, classification, examples, and statistics surfaces establish these fields and lifecycle boundaries.
---
id: coverage.project-topics
title: Project Topics
summary: Maintain source-grounded subject areas, editable groups, and Atlas relationships used to organize project and benchmark coverage.
kind: reference
product_area: coverage_engineering
status: stable
updated: 2026-09-07
canonical: /docs/coverage-engineering/project-topics
---
# Project Topics
## Definition
Project Topics are reusable subject areas derived from project knowledge, entered by an administrator, or contributed by an expert. They organize cases by what they are about. Unlike a Dimension, a Project Topic is not one value on a fixed behavior axis: a case can belong to several Topics.
The Project Topics workspace contains **Topics**, **Groups**, and **Atlas**.
## Fields, states, or lifecycle rules
### Topics
Each Topic has a name, description, active state, author type, case usage, benchmark usage, and source mentions where available. Source mentions retain evidence such as the originating Reference Material, Project Agent Brief, manual entry, or expert suggestion.
The Topic detail shows its definition and source grounding, linked and example cases, coverage confidence, nearby-topic signals, coverage by Dimensions and ontology values, and editable membership in Topic Groups.
An AI-created or source-extracted Topic remains an editable project artifact. Review its definition and grounding before using it to steer a benchmark.
### Groups
A Project Topic Group is an editable bundle of Topics. It does not merge or replace its members. Groups let benchmark setup express intent at a useful scale while retaining Topic-level traceability.
In **Get Started**, a Group can be handled as:
- **Cover every topic:** the benchmark should represent each member Topic.
- **Cover the group:** the group should be represented without requiring every member.
- **Use as guidance:** it can guide story construction without becoming a coverage obligation.
- **Do not use:** exclude the Group from this benchmark's setup.
- **Unset:** no explicit instruction.
Groups can start from manual work, AI suggestions, source material, or Atlas exploration. Manual edits remain significant; regenerating a suggestion should not be treated as authority to overwrite the reviewed group.
### Atlas
Atlas visualizes Topics, relationships, and Groups. Use it to inspect neighborhoods, redundancy, missing nearby Topics, and possible groupings. Relationships are analytical evidence, not a taxonomy merge. A close position or strong relation score does not mean two Topics are interchangeable.
The Topic coverage view can distinguish direct case grounding from breadth, Dimension spread, and binding confidence. When a Topic looks thin, inspect the linked cases before generating more. The problem may be missing cases, weak classification, an overly broad definition, or a duplicate Topic.
### Refresh and review
Project Topics can be refreshed from current project context and indexed sources. Refreshing can create or update Topics, source mentions, evidence, and relationship analysis. Review the resulting changes and recommendation runs before incorporating them into Groups or benchmark setup.
{% example-demo title="Example: group without flattening" %}
The project has Topics for Contract renewal, Price adjustment, and Termination notice. An administrator groups them as Agreement lifecycle. One benchmark chooses Cover every topic because each action has distinct risk. Another chooses Use as guidance because it only needs broad agreement-related examples. The individual Topics remain traceable in both benchmarks.
{% /example-demo %}
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Work with Dimensions and Ontology](/docs/coverage-engineering/dimensions-ontology)
- [Define Case Construction Patterns](/docs/coverage-engineering/case-construction-patterns)
- [Configure benchmark Get Started](/docs/coverage-management/get-started)
- [Analyze dataset representation](/docs/benchmark-datasets/representation)
{% /related-card-grid %}
## Source confidence
Code-backed: the current Topic, Group, source-mention, statistics, relationship, Atlas, suggestion, and benchmark-handling contracts establish this model.