Teammately Docs
Docs menu

task

Dataset Representation

Analyze how distinct benchmark Cases are distributed across facets, evaluator rules, and provenance.

Dataset Representation

Prerequisites

  • A current benchmark dataset or Snapshot with representation facts.
  • Coverage Facets and evaluator relationships meaningful enough to interpret.

Representation groups the current or snapshotted dataset by governed facts. Available groupings include Dimension ontology values, Topic Groups, Project Topics, Case Construction Patterns, Policies, policy application, Rubrics, rubric application, presence of rubrics, and contributors.

Choose distinct Cases when counts matter, or Case share when comparing proportions. Policy and rubric views can split by application state. Filters and drilldowns narrow the visible population, and the resulting table or chart can be exported as CSV.

Reading the view

  • A large bar means concentration, not correctness.
  • An empty category can indicate a true coverage gap, an inactive facet, missing classification, or a filter that excludes the Cases.
  • Topic Groups do not merge their member Topics; group-level handling and Topic-level representation remain distinct.
  • Policy and rubric presence is not the same as approved eligible application.
  • Contributor distribution is provenance evidence, not a substitute for agreement or evaluator quality.

Use Coverage Management when a gap should drive a Coverage Story or Case Foundry work. Use Expert Contributions when the missing evidence requires governed expert judgment.

Worked example

Example: count and share tell different stories

A Topic Group has twenty Cases but represents 60% of a small dataset, while a required ontology value has only two. Distinct count reveals the thin required value; Case share reveals the concentration. The operator records a Coverage Story instead of presenting the large Topic count as balanced coverage.

Object and state changes

Grouping, metrics, filtering, splitting, drilldown, and CSV export change only the analysis view. They do not classify Cases, edit facets, or modify Snapshot content.

Success criteria

  • Counts and shares use the intended Case population.
  • Missing, thin, and concentrated categories are distinguished.
  • A governed Coverage Story or follow-up owns any actionable gap.

Common failure modes

  • Reading a filtered percentage as the whole dataset.
  • Equating high volume with representative coverage.
  • Reconstructing an old Snapshot from current classifications.

Source confidence

Code-backed: the active Representation route defines grouping, split, metric, filtering, drilldown, chart/table, and CSV behavior.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context