Teammately Docs
Docs menu

reference

Glossary

Definitions for current Teammately capabilities, product surfaces, artifacts, and lifecycle states.

Glossary

Use these definitions when choosing a product surface, writing an operating instruction, or interpreting evaluation evidence. Capitalization identifies current product capabilities, surfaces, and named artifacts where it matters.

Definition

The glossary fixes the meaning of current capability, surface, artifact, and lifecycle terms. Capability names describe what Teammately enables; product-surface names describe where people operate; artifact names identify the state that must remain traceable.

Capabilities

  • Correctness Elicitation: Turns specialist judgment into attributable standards and cases through structured dialogue, tasks, and checkpoints.
  • Coverage Engineering: Represents the behavior space that a benchmark must cover through facets, stories, patterns, and concrete Cases.
  • Weave: Builds case-grounded agent behavior from project knowledge, policies, Rubrics, tools, materials, and worlds.
  • Trialground: Evaluates saved Harness versions against versioned benchmark boundaries and exposes inspectable evidence.
  • Coevolve: Improves agents and benchmark understanding together by feeding contribution, coverage, and evaluation evidence into governed iteration.

Fields, states, or lifecycle rules

Project foundations

  • Project Agent Brief: The published project-understanding document in Project Context.
  • Reference Materials: The Agent Setup surface for connecting and indexing source-backed project knowledge.
  • Reference block: An indexed unit of project knowledge with source and generation identity.
  • Comparison Direction: Project Asset that guides meaningful candidate-output differences in comparative expert work; it is not a coverage facet or approved standard.
  • Review Screen: Project Asset that configures the context and inputs presented to an expert.
  • Policy: A governed statement of required or prohibited behavior, with applicability and provenance.
  • Rubric: An executable judgment criterion, commonly binary, that tests applicable behavior against a Policy.
  • Coverage Facet: Reusable project structure for representing behavior space, including Dimensions, Project Topics, and Case Construction Patterns.
  • Case: A reusable behavior situation with canonical input, optional case materials, metadata, and an optional executable world reference.
  • Harness: The agent configuration being built and evaluated. A mutable Draft can be saved as an immutable Harness version.
  • Project Input Schema: The project-level architecture governing canonical Case input and material fields.
  • Benchmark Run Metadata: Benchmark-level descriptive context requested when an Evaluation Run is created; it is not a project-level template or version identity.

Benchmark workspace

  • Benchmark: A named workspace for a particular correctness boundary and its datasets, coverage, contributions, evaluations, and improvement work.
  • Dataset snapshot: A reproducible selection and representation of benchmark Cases.
  • Coverage Story: Benchmark-scoped intent connecting a behavior risk or need to concrete coverage work.
  • Expert Contribution: A benchmark-scoped request for specialist judgment, containing one or more Tasks and optional Checkpoints.
  • Task: A bounded unit of work inside a Contribution, such as form input, chat, interview, or Case Review.
  • Checkpoint: An explicit confirmation boundary inside a Contribution.
  • Contributed artifact: A policy, Rubric, Case, or coverage observation produced by an expert while retaining Contribution provenance.
  • Benchmark version: The fixed dataset and correctness boundary used for reproducible evaluation.
  • Evaluation Run: One execution of a saved Harness Version against a Benchmark Version, with response, Rubric outcomes, settings, mapping, and metadata. Execution trajectories are not currently exposed.
  • Compare: A symmetric Benchmark Evaluations matrix whose columns are saved Harness Versions and whose rows group governed evidence within one Benchmark Version.
  • Arena: A comparative evaluation surface inside Benchmark Evaluations.
  • Improvement Session: A benchmark-scoped process that explores candidate Harness changes against an explicit goal and pinned evidence.
  • Goal Contract: The Improvement Session definition of target evidence, success criteria, and constraints.
  • Evaluation receipt: Canonical evidence that a particular candidate was evaluated under a particular boundary.
  • Current frontier: The set of retained candidates that currently represent the session's best supported tradeoffs.

Decision checkpoint

If you mean...Use...Do not substitute...
Reusable source-backed project knowledgeReference Materials / Reference blockAn untracked attachment or the latest file without generation identity
Human specialist work for a benchmarkExpert ContributionA generic approval queue
The agent state actually evaluatedSaved Harness versionThe mutable Harness Draft
A fixed evaluation boundaryBenchmark versionA Run or a score
One execution and its evidenceEvaluation RunThe Benchmark itself
Goal-directed candidate explorationImprovement SessionAn unversioned list of suggestions

Worked example

Distinguish benchmark and run

Benchmark Version 4 fixes the selected Cases and governed standards. Harness Version 9 is the candidate. The Evaluation Run is the one execution of Harness Version 9 against Benchmark Version 4. Compare can place Harness Version 9 beside other saved Harness Versions in a symmetric evidence matrix, while an Improvement Session can use failed Case evidence as a pinned target for new candidates.

Source confidence

Docs-backed and code-aligned: the product doctrine defines the five capabilities, and current navigation establishes the product-surface names. Use the linked code-backed pages when an exact field, state transition, or route behavior matters.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context