# Plan Benchmark Coverage
Generated: 2026-09-13T04:33:46.837Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: coverage.plan-benchmark-coverage
title: Plan Benchmark Coverage
summary: Apply project Coverage Facets to one benchmark, inspect representation, and turn important gaps into concrete case or contribution work.
kind: task
product_area: coverage_engineering
status: stable
updated: 2026-09-07
canonical: /docs/coverage-engineering/plan-benchmark-coverage
---
# Plan Benchmark Coverage
Plan coverage by applying reusable project facets to one benchmark and comparing the intended behavior space with the selected dataset representation.
## Prerequisites
- A selected project and benchmark.
- A clear benchmark purpose.
- Relevant Dimensions, Project Topics, and Case Construction Patterns, or enough project knowledge to create them.
- Existing Cases or a plan for sourcing and constructing them.
## Steps
1. Review **Coverage Facets** at project scope. Confirm that Dimensions, Project Topics, and Case Construction Patterns describe reusable behavior structure rather than one benchmark's current case count.
2. Open the benchmark and select **Coverage Management → Get Started**.
3. Define the benchmark-specific coverage guidance and confirm setup readiness.
4. Open Coverage Management and inspect current dataset representation across the relevant facets and tuples.
5. Name important thin or absent combinations as Coverage Stories. Explain why each slice matters and what evidence would make it usable.
6. Route the gap according to its cause: Case Foundry or case sourcing for missing situations, Expert Contributions for missing judgment, Correctness Governance for missing standards, or Benchmark Datasets for missing selection.
7. Review generated or contributed cases in Case Review before relying on them.
8. Update dataset selection and create a new snapshot when the represented evidence changes materially.
## Object and state changes
This task can update benchmark coverage setup, representation guidance, Coverage Stories, Case Foundry work, case-review state, contribution requests, dataset selection, and snapshots. Project Coverage Facets may also change when the work discovers a reusable missing axis or construction pattern.
## Success criteria
- The benchmark purpose maps to explicit project Coverage Facets.
- Important combinations have selected evidence or a named gap.
- Each gap is routed to a responsible artifact or workstream.
- Constructed cases pass case review and Project Input Schema checks.
- Dataset snapshots make material coverage changes explicit.
## Common failure modes
- Using case count as the coverage goal.
- Creating benchmark-only tags where a reusable Dimension or Topic is needed.
- Treating response-variation guidance as coverage structure.
- Generating cases before defining which gap they should close.
- Trusting representation after selection changes without a new snapshot boundary.
{% example-demo title="Example: plan high-impact exception coverage" %}
The team maps exception type, source authority, and customer impact. Representation shows many low-impact ordinary cases but no high-impact cases with conflicting authority. A Coverage Story names the gap, an expert Contribution clarifies the controlling rule, and Case Foundry prepares cases for the missing tuple before a new snapshot is created.
{% /example-demo %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Coverage Engineering](/docs/coverage-engineering)
- [Coverage Management](/docs/coverage-management)
- [Benchmark Datasets](/docs/benchmark-datasets)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
- [Stale dimensions](/docs/troubleshooting/stale-dimensions)
- [Unrealistic synthetic cases](/docs/troubleshooting/unrealistic-synthetic-cases)
{% /related-card-grid %}
## Source confidence
Code-backed: current setup, overview, representation, Coverage Story, Case Foundry, and Case Review routes support this workflow.
---
id: coverage.overview
title: Coverage Engineering
summary: Design the behavior space a benchmark must represent and connect reusable project facets to benchmark coverage work.
kind: concept
product_area: coverage_engineering
status: stable
updated: 2026-09-07
canonical: /docs/coverage-engineering
---
# Coverage Engineering
Coverage Engineering is the capability for designing what a specialist AI benchmark must represent. It turns requirements, project knowledge, existing cases, and observed failures into a structured coverage map that guides dataset selection, case construction, expert contributions, and evaluation interpretation.
## Definition
Coverage work has two scopes. At project scope, **Coverage Facets** manages reusable Dimensions, Project Topics, and Case Construction Patterns. At benchmark scope, **Coverage Management** applies those facets to setup, representation, Coverage Stories, Case Review, Case Foundry, and contribution requests. **Benchmark Datasets** owns the selected Cases and snapshots that embody the resulting coverage.
Coverage Engineering is therefore broader than adding cases. It explains why a behavior slice matters, how it is represented, which combinations are thin, and what work should close the gap.
## Decision checkpoint
| Question | Product surface | Durable result |
| --- | --- | --- |
| Which axes distinguish important behavior? | Coverage Facets → Dimensions | Reusable coverage axes |
| Which domain subjects must be represented? | Coverage Facets → Project Topics | Project topic structure |
| How should cases be constructed repeatedly? | Coverage Facets → Case Construction Patterns | Reusable construction guidance |
| What should this benchmark cover? | Coverage Management → Get Started | Benchmark-specific coverage guidance |
| Where is the selected dataset thin? | Representation and Coverage Stories | Named gaps and sourcing work |
| Which exact cases define evidence? | Benchmark Datasets | Selection and snapshot boundary |
## Coverage map and benchmark evidence
A coverage map should identify meaningful combinations rather than isolated tags. A source-freshness Dimension may be well populated overall while the combination of superseded source, high customer impact, and exception request remains absent. Coverage Stories make those combinations operational; Case Foundry and expert contribution requests can then target them.
Representation is evidence about the dataset, not proof that the benchmark is complete. A large or balanced count can still omit a critical boundary. Product teams should use specialist judgment to decide which gaps materially affect trust.
## Relationship to correctness and construction
Correctness Elicitation answers what should count as correct in a represented situation. Weave constructs or imports the cases and materials needed to exercise the situation. If the team cannot judge a coverage slice, request an Expert Contribution. If the judgment is clear but no case exists, use Case Foundry or other case-construction work. If cases exist but are not selected, update Benchmark Datasets.
Comparison Directions are not Coverage Facets. They guide comparative response variation and belong to **Assets → Comparison Directions**. Keep benchmark representation in Dimensions, Topics, Patterns, Stories, and dataset snapshots.
{% example-demo title="Authority-conflict coverage" %}
A project creates source authority and customer impact Dimensions, a Project Topic for eligibility exceptions, and a pattern for pairing current and superseded documents. Coverage Management shows that the high-impact conflict tuple has no selected cases. A Coverage Story justifies the gap, Case Foundry prepares cases, and the accepted cases enter a new dataset snapshot.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Plan benchmark coverage](/docs/coverage-engineering/plan-benchmark-coverage)
- [Manage Coverage](/docs/coverage-management)
- [Work with Benchmark Datasets](/docs/benchmark-datasets)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Coverage Dimensions](/docs/object-model/coverage-dimensions)
- [Ontology](/docs/object-model/ontology)
- [Cases](/docs/assets/cases)
{% /related-card-grid %}
## Source confidence
Code-backed: current navigation and benchmark coverage routes establish project Coverage Facets, benchmark Coverage Management, setup, representation, Stories, Case Review, and Case Foundry responsibilities.
---
id: coverage-management.overview
title: Coverage Management
summary: Manage benchmark coverage from setup through representation, Coverage Stories, case review, Case Foundry, and contribution requests.
kind: concept
product_area: coverage_management
status: stable
updated: 2026-09-07
canonical: /docs/coverage-management
---
# Coverage Management
Coverage Management is the benchmark-scoped workspace for deciding whether the current dataset represents the behavior space the benchmark is meant to test. It connects project Coverage Facets to dataset representation, coverage guidance, Coverage Stories, case preparation, and expert contribution requests.
## Definition
Project Coverage Facets define reusable Dimensions, Project Topics, and Case Construction Patterns. Coverage Management applies those foundations to one benchmark. **Get Started** establishes benchmark coverage guidance. The overview shows representation and operational status. Coverage Stories organize meaningful slices or gaps. Case Review inspects prepared cases and materials. Case Foundry coordinates case construction work.
The goal is to make missing or thin behavior explicit before evaluation evidence is trusted. Coverage Management does not replace Benchmark Datasets; it explains and improves the representation of the selected data.
## Decision checkpoint
| Observation | Use | Next durable result |
| --- | --- | --- |
| Benchmark purpose or guidance is missing | Get Started | Saved coverage setup and readiness |
| A facet tuple is thin or absent | Representation and Coverage Stories | Named coverage need and intended evidence |
| More cases are needed | Case Foundry | Bounded construction work tied to the gap |
| Generated cases may be unclear | Case Review | Reviewed case and material quality |
| Specialist judgment is required | Contribution request from coverage context | Benchmark-scoped Expert Contribution |
| Coverage changed materially | Benchmark Datasets | Updated selection and snapshot boundary |
## Coverage Stories and case work
A Coverage Story gives a gap or behavior slice an operational narrative: why it matters, which facet combinations define it, what evidence exists, and what sourcing work remains. It should be concrete enough to guide case construction and expert attention.
Case Foundry can prepare case work from that structure. Case Review checks the resulting inputs and generated materials before they enter trusted dataset evidence. AI assistance can accelerate preparation, but selection and benchmark interpretation remain explicit human and product-state decisions.
## Coverage and correctness
Coverage gaps sometimes reveal missing correctness rather than missing cases. If experts cannot say how a represented situation should be judged, request an Expert Contribution and update policies or rubrics. If the standard is clear but no case exercises it, use Weave and Case Foundry. If cases exist but are not selected or snapshotted, use Benchmark Datasets.
This routing prevents Comparison Directions, Contribution-scoped agent behavior, and coverage structure from being mixed together. Comparison Directions guide response variation; Coverage Facets and Coverage Stories describe the benchmark behavior space.
{% example-demo title="Conflicting-source story" %}
Representation shows that the benchmark covers current-source questions but almost never combines them with a plausible superseded document. A Coverage Story names the conflict pattern, relevant source-freshness and impact facets, and the desired case count. Case Foundry prepares candidates, Case Review rejects unrealistic material, and the accepted cases enter a new dataset snapshot.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Set up benchmark coverage](/docs/coverage-management/get-started)
- [Work with Coverage Stories](/docs/coverage-management/coverage-stories)
- [Run Case Foundry](/docs/coverage-management/case-foundry)
- [Review prepared Cases](/docs/coverage-management/case-review)
- [Plan benchmark coverage](/docs/coverage-engineering/plan-benchmark-coverage)
- [Work with Benchmark Datasets](/docs/benchmark-datasets)
- [Request an Expert Contribution](/docs/expert-contributions/request-contribution)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Coverage Engineering](/docs/coverage-engineering)
- [Cases](/docs/assets/cases)
- [Comparison Directions](/docs/assets/comparison-directions)
{% /related-card-grid %}
## Source confidence
Code-backed: the active benchmark coverage routes expose setup, overview, Coverage Stories, Case Review, Case Foundry integration, realtime state, and contribution-request entry points.
---
id: benchmark-datasets.overview
title: Benchmark Datasets
summary: Select benchmark Cases, inspect representation, and freeze immutable Snapshots for reproducible evidence.
kind: concept
product_area: benchmark_datasets
status: stable
updated: 2026-08-22
canonical: /docs/benchmark-datasets
---
# Benchmark Datasets
Benchmark Datasets defines the evidence set for one benchmark through **Cases**, **Representation**, and **Snapshots**.
The current dataset is editable. It selects reusable project Cases and reflects current facet, policy, rubric, and contributor facts. A Snapshot freezes the exact dataset state needed by a Benchmark Version and its evaluations. These are deliberately different surfaces: editing the current set must not rewrite historical evidence.
## Decision checkpoint
| Surface | Use it to | Evidence rule |
| --- | --- | --- |
| Cases | Inspect and change current benchmark membership | Selection is live until snapshotted |
| Representation | Find concentration and absence across governed facets | Read distribution together with distinct Case counts |
| Snapshots | Freeze Cases, evaluator links, and representation facts | Snapshot content is read-only |
Coverage Management acts on gaps found in the dataset. Assets remains the project-level reusable pool. Benchmark Evaluations runs exact Harness Versions against an immutable Benchmark Version rather than an unspecified “current dataset.”
## Evidence flow
Cases usually begin in project Assets or materialize through Case Review or Expert Contributions. Selecting them makes them part of the current benchmark dataset. Representation then summarizes the current assignments and evaluator relationships. Snapshot readiness checks whether that state can be frozen. A Snapshot supplies the immutable dataset facts used by a Benchmark Version.
This flow is one-way for historical evidence. Later edits to an Asset, facet assignment, Policy, Rubric, or current membership may improve the next Snapshot, but they do not update a previous Snapshot. Compare candidates within one Benchmark Version unless the analysis explicitly accounts for a moved evidence boundary.
## Before creating evidence
Check Case clarity and schema conformance, then inspect Representation for intended behavior and provenance. Confirm approved eligible evaluator links. Resolve Snapshot blockers and preserve the resulting label, version, content hash, creation time, and Case count.
A Snapshot can be reproducible while still being incomplete as product coverage. Reproducibility answers which evidence was evaluated; Representation and Coverage Management answer whether that evidence supports the intended product claim.
{% example-demo title="Example: editable set versus frozen evidence" %}
The current dataset gains four Cases and a corrected Rubric link after an expert Contribution is reconciled. An earlier Run still points to its old Benchmark Version. The operator creates a new Snapshot and Version for the changed set rather than comparing the new candidate against the old Run as though only Harness behavior moved.
{% /example-demo %}
{% related-card-grid title="Dataset workflows" %}
- [Manage benchmark Cases](/docs/benchmark-datasets/cases)
- [Inspect Representation](/docs/benchmark-datasets/representation)
- [Create and inspect Snapshots](/docs/benchmark-datasets/snapshots)
- [Manage coverage](/docs/coverage-management)
{% /related-card-grid %}
## Source confidence
Code-backed: the active dataset routes establish the editable current set, representation workspace, and immutable Snapshot boundary.