Glossary
Use these definitions when choosing a product surface, writing an operating instruction, or interpreting evaluation evidence. Capitalization identifies current product capabilities, surfaces, and named artifacts where it matters.
Definition
The glossary fixes the meaning of current capability, surface, artifact, and lifecycle terms. Capability names describe what Teammately enables; product-surface names describe where people operate; artifact names identify the state that must remain traceable.
Capabilities
- Correctness Elicitation: Turns specialist judgment into attributable standards and cases through structured dialogue, tasks, and checkpoints.
- Coverage Engineering: Represents the behavior space that a benchmark must cover through facets, stories, patterns, and concrete Cases.
- Weave: Builds case-grounded agent behavior from project knowledge, policies, Rubrics, tools, materials, and worlds.
- Trialground: Evaluates saved Harness versions against versioned benchmark boundaries and exposes inspectable evidence.
- Coevolve: Improves agents and benchmark understanding together by feeding contribution, coverage, and evaluation evidence into governed iteration.
Fields, states, or lifecycle rules
Project foundations
- Project Agent Brief: The published project-understanding document in Project Context.
- Reference Materials: The Agent Setup surface for connecting and indexing source-backed project knowledge.
- Reference block: An indexed unit of project knowledge with source and generation identity.
- Comparison Direction: Project Asset that guides meaningful candidate-output differences in comparative expert work; it is not a coverage facet or approved standard.
- Review Screen: Project Asset that configures the context and inputs presented to an expert.
- Policy: A governed statement of required or prohibited behavior, with applicability and provenance.
- Rubric: An executable judgment criterion, commonly binary, that tests applicable behavior against a Policy.
- Coverage Facet: Reusable project structure for representing behavior space, including Dimensions, Project Topics, and Case Construction Patterns.
- Case: A reusable behavior situation with canonical input, optional case materials, metadata, and an optional executable world reference.
- Harness: The agent configuration being built and evaluated. A mutable Draft can be saved as an immutable Harness version.
- Project Input Schema: The project-level architecture governing canonical Case input and material fields.
- Benchmark Run Metadata: Benchmark-level descriptive context requested when an Evaluation Run is created; it is not a project-level template or version identity.
Benchmark workspace
- Benchmark: A named workspace for a particular correctness boundary and its datasets, coverage, contributions, evaluations, and improvement work.
- Dataset snapshot: A reproducible selection and representation of benchmark Cases.
- Coverage Story: Benchmark-scoped intent connecting a behavior risk or need to concrete coverage work.
- Expert Contribution: A benchmark-scoped request for specialist judgment, containing one or more Tasks and optional Checkpoints.
- Task: A bounded unit of work inside a Contribution, such as form input, chat, interview, or Case Review.
- Checkpoint: An explicit confirmation boundary inside a Contribution.
- Contributed artifact: A policy, Rubric, Case, or coverage observation produced by an expert while retaining Contribution provenance.
- Benchmark version: The fixed dataset and correctness boundary used for reproducible evaluation.
- Evaluation Run: One execution of a saved Harness Version against a Benchmark Version, with response, Rubric outcomes, settings, mapping, and metadata. Execution trajectories are not currently exposed.
- Compare: A symmetric Benchmark Evaluations matrix whose columns are saved Harness Versions and whose rows group governed evidence within one Benchmark Version.
- Arena: A comparative evaluation surface inside Benchmark Evaluations.
- Improvement Session: A benchmark-scoped process that explores candidate Harness changes against an explicit goal and pinned evidence.
- Goal Contract: The Improvement Session definition of target evidence, success criteria, and constraints.
- Evaluation receipt: Canonical evidence that a particular candidate was evaluated under a particular boundary.
- Current frontier: The set of retained candidates that currently represent the session's best supported tradeoffs.
Decision checkpoint
| If you mean... | Use... | Do not substitute... |
|---|---|---|
| Reusable source-backed project knowledge | Reference Materials / Reference block | An untracked attachment or the latest file without generation identity |
| Human specialist work for a benchmark | Expert Contribution | A generic approval queue |
| The agent state actually evaluated | Saved Harness version | The mutable Harness Draft |
| A fixed evaluation boundary | Benchmark version | A Run or a score |
| One execution and its evidence | Evaluation Run | The Benchmark itself |
| Goal-directed candidate exploration | Improvement Session | An unversioned list of suggestions |
Worked example
Distinguish benchmark and run
Benchmark Version 4 fixes the selected Cases and governed standards. Harness Version 9 is the candidate. The Evaluation Run is the one execution of Harness Version 9 against Benchmark Version 4. Compare can place Harness Version 9 beside other saved Harness Versions in a symmetric evidence matrix, while an Improvement Session can use failed Case evidence as a pinned target for new candidates.
Source confidence
Docs-backed and code-aligned: the product doctrine defines the five capabilities, and current navigation establishes the product-surface names. Use the linked code-backed pages when an exact field, state transition, or route behavior matters.