# Using Teammately Alongside Existing Evaluation Infrastructure
Generated: 2026-09-13T04:42:53.739Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: playbooks.alongside-existing-evals
title: Using Teammately Alongside Existing Evaluation Infrastructure
summary: Position Teammately as correctness infrastructure that can complement existing tests and metrics.
kind: recipe
product_area: playbooks
status: stable
updated: 2026-09-07
canonical: /docs/playbooks/using-teammately-alongside-existing-evaluation-infrastructure
---
# Using Teammately Alongside Existing Evaluation Infrastructure
Use this playbook when a team already has tests, traces, dashboards, or offline evals and wants Teammately to add expert-grounded correctness evidence rather than replace everything.
## Choose the integration boundary
Use this when existing infrastructure already owns candidate execution, CI status, traces, or metrics and the team wants Teammately to own deliberate coverage, expert-grounded standards, and Benchmark evidence. Keep external systems authoritative for facts Teammately does not ingest or compute.
## Two supported evidence paths
1. Inventory the external Case identity, candidate identity, metrics, CI status, and links needed by customer operators.
2. Import representative inputs as Teammately Cases and design their coverage and governed evaluator boundary.
3. For a candidate Teammately can execute, save it as a Harness Version and create a managed Run.
4. For responses produced externally, export immutable Teammately Case IDs, generate one response per intended Case, and map them into an output-only Run.
5. Store useful external candidate or run identifiers in the configured Run Metadata fields. Do not claim Teammately verified an external metric merely because its identifier is present.
6. Inspect Teammately Rubric evidence in List. Use Compare and Arena only for saved Harness Versions; output-only Runs are not Harness columns.
7. In customer-owned human review context, present Teammately evidence and external CI or telemetry as separately sourced facts.
## Ownership matrix
| Evidence | Owning system |
| --- | --- |
| Coverage Facets, Policies, Rubrics, Benchmark Versions | Teammately |
| Managed Harness response and Rubric outcomes | Teammately Run |
| Imported response mapping and Rubric outcomes | Teammately output-only Run |
| External latency, CI status, trace, or deployment fact | External system |
| Downstream operational decision | Customer |
{% example-demo title="CI benchmark handoff" %}
A team keeps latency and regression tests in CI. It exports Teammately Case IDs, produces candidate responses externally, and imports them as one output-only Run with the CI run ID in Run Metadata. Teammately reports must-level Policy failures for those responses; CI remains authoritative for latency. The customer's review context shows both facts with their sources and does not treat the imported candidate as a saved Harness Version.
{% /example-demo %}
## Evidence to collect
- External eval artifacts and the identifiers needed to trace candidates or runs.
- The Teammately-owned Cases, standards, coverage, Benchmark Version, output mapping, Rubric evidence, and Run Metadata.
- Mapping notes that explain which metrics remain outside Teammately.
- Benchmark comparisons under approved human standards.
- Review context that keeps external CI status separate from Teammately benchmark evidence.
## Related docs
{% related-card-grid title="Related docs" %}
- [Map external outputs](/docs/benchmark-evaluations/output-mapping)
- [Configure Run Metadata](/docs/benchmark-evaluations/run-metadata)
- [Compare Harness Versions](/docs/benchmark-evaluations/compare)
- [Read run results](/docs/benchmark-evaluations/inspect-results)
- [Run a benchmark](/docs/benchmark-evaluations/run-evaluation)
- [Importing cases](/docs/operating-manual/import-and-prepare-cases)
- [Agent Setup](/docs/agent-setup)
{% /related-card-grid %}
## Source confidence
Doctrine-backed: the approved product boundary keeps Teammately correctness evidence separate from external infrastructure facts and customer decisions. Linked code-backed pages define managed Runs, output-only Runs, Run Metadata, and comparison eligibility.
---
id: intro.product-boundaries
title: Product boundaries
summary: Understand what Teammately owns across correctness specification, benchmark development, evaluation, and improvement.
kind: concept
product_area: introduction
status: stable
updated: 2026-09-07
canonical: /docs/introduction/product-boundaries
---
# Product boundaries
Teammately owns the correctness system that connects specialist judgment to deliberate benchmark coverage, executable standards, constructed cases, evaluation evidence, and improvement history. This page distinguishes that system from adjacent inputs and downstream responsibilities.
> Adjacent systems are inputs
>
> Logs, traces, source repositories, model endpoints, coding environments, and external evaluation results can supply material or receive work. Their presence does not change the Teammately ownership boundary: Teammately governs the connected correctness artifacts and the evidence produced from them.
## Definition
The product boundary follows artifacts and authority. Teammately can index project knowledge, prepare an expert contribution, materialize an accepted policy or rubric, construct a case, execute an evaluation through a managed harness, and coordinate an Improvement Session. It preserves which inputs, versions, settings, and human decisions produced the resulting evidence.
Customer teams own the AI system outside that evidence graph and the action taken afterward. Teammately can prepare a scoped package for an external coding worker, but it does not claim private work performed outside the product. It can show benchmark evidence, but it does not turn that evidence into an automatic downstream decision.
## Decision checkpoint
| Area | Teammately owns | Boundary |
| --- | --- | --- |
| Project knowledge and agent context | Materials, Indexed Reference, Project Context, and Contribution-scoped agent behavior | Reference material is not automatically a governed policy or rubric |
| Expert work | Contribution scope, tasks, checkpoints, attributable responses, and contributed artifacts | Agent preparation does not substitute for the expert's judgment |
| Cases and worlds | Canonical case input, case materials, generated artifacts, and verified world references | Static materials and executable environments remain distinct |
| Evaluation | Benchmark Versions, saved Harness Versions, settings, Runs, responses, Rubric results, and comparisons | A score alone does not explain correctness; execution traces are not currently exposed |
| Improvement | Goal Contracts, candidates, evaluation receipts, frontiers, and chronology | External worker activity is represented only when returned through the defined contract |
| Downstream action | Inspectable correctness evidence and review context | The customer decides what operational action follows |
## Human and agent authority
AI agents scale preparation and exploration. They can organize source material, propose coverage structure, draft possible standards, generate cases, evaluate candidates, and suggest improvement directions. The owning surface determines when an artifact becomes durable or governed. An agent proposal does not silently acquire expert authority.
Expert Contributions make this boundary explicit. The product can prepare focused questions and relevant evidence, while the domain specialist supplies the judgment. Correctness Governance records policies and rubrics as governed project assets. Improvement Sessions can branch candidate hypotheses, but retained candidates require observable evaluation evidence.
## Data and execution boundary
Project Input Schema controls the accepted shape of case input and materials. Static context remains part of case content or case-material references. An executable or queryable environment uses a world reference and follows a separate runtime boundary. Public docs describe the behavior visible through stable product surfaces; they do not promote internal storage or service structures into customer-facing contracts.
Similarly, the presence of Harness Assets and managed Runs does not imply that Teammately owns a customer's model registry, production telemetry, or deployment system. A harness is the executable candidate boundary used by a benchmark evaluation.
{% example-demo title="External coding worker" %}
An Improvement Session starts from failed grounding Cases and a confirmed Goal Contract. Teammately prepares a scoped package for a coding worker with the pinned target and evidence. The worker's private activity is outside the product boundary. A returned Harness Version and canonical evaluation request become observable candidates; their Rubric results enter the session record, while Improve may add a safe narrated trajectory of observable session activity.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [The correctness lifecycle](/docs/introduction/correctness-lifecycle)
- [Start an Improvement Session](/docs/improve/start-improvement-session)
- [Use Reference Materials](/docs/agent-setup/reference-materials)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Human Approval Boundaries](/docs/governance/human-approval-boundaries)
- [What AI Features Can and Cannot Do](/docs/governance/what-ai-features-can-and-cannot-do)
- [Project Input Schema](/docs/project-settings/input-schema)
{% /related-card-grid %}
## Source confidence
Doctrine-backed: this page states product ownership and authority boundaries. Exact UI and execution behavior is delegated to linked code-backed pages.
---
id: integrations.overview
title: Integrations
summary: Move Cases, source materials, external outputs, and benchmark evidence across Teammately's supported product boundaries.
kind: concept
product_area: data_integrations
status: stable
updated: 2026-09-07
canonical: /docs/integrations
---
# Integrations
## Definition
An integration connects customer-owned data or execution systems to a specific Teammately workflow. The useful boundary is not “send arbitrary records.” It is preserving enough identity, context, and version information for imported material to remain reviewable and for exported evidence to remain interpretable.
## Choose the boundary
| need | Teammately boundary | preserve |
| --- | --- | --- |
| Bring behavior examples into a Project | Assets Cases or the Case import workflow | source identity, input, materials, attributes, current Project Input Schema |
| Give agents Project knowledge | Agent Setup → Reference Materials | source document identity, indexing state, intended use |
| Evaluate outputs produced elsewhere | Benchmark Evaluations → Import reference outputs | immutable Benchmark Version, exact Case IDs, output identity, measured telemetry |
| Send work to specialists | Expert Contributions and Expert UI | assignment, task context, Contribution state, Checkpoint, provenance |
| Use evidence in another system | Run and Case result context | Run, Benchmark Version, Case version, Policy version, Rubric version, incomplete states |
Each pattern has a different lifecycle. Reference Materials are not Case materials. An external response is neither a Case-scoped reference output nor a Harness Version. A Contribution assignment is not a generic notification integration.
## Data movement principles
1. Read the owning product contract before transforming data. For Cases, start with Project Input Schema; for outputs, start with the immutable Benchmark Version.
2. Use backend-issued IDs as join keys and customer IDs as correlation keys.
3. Preserve versions when a downstream result depends on mutable source objects.
4. Distinguish acceptance from completion. Upload, artifact processing, evaluation, and expert work can have separate states.
5. Keep missing values missing. Do not turn unknown telemetry or incomplete evaluation into zero or false.
6. Reconcile a small sample before moving a full collection.
## Product UI and API boundaries
The product UI owns interactive setup, mapping, preview, conflict resolution, and human decisions. Automated integrations should use only a supported external contract that preserves the same object and version boundaries.
Internal browser endpoints, event payloads, database shapes, and service-to-service handlers are not customer integration contracts. Their presence in source code does not make them stable or safe to automate against.
{% example-demo title="External candidate evaluation" %}
A team produces assistant outputs in its own evaluation pipeline. It selects an immutable Benchmark Version, downloads the exact Case IDs, and maps one output to each Case. The imported output set becomes an output-only reference Run.
The team retains its own generation ID as correlation metadata while using the Teammately Run, Benchmark Version, and Case IDs as correctness-evidence identity. Missing cost remains absent. The Run becomes interpretable after Rubric evaluation completes; successful upload alone is not a result.
{% /example-demo %}
## Fields, states, or lifecycle rules
- Projects scope reusable Cases, Project Context, Reference Materials, Harness assets, and correctness governance. Contribution-specific agent behavior and direction selection belong to benchmark-scoped Contributions.
- Benchmark Versions freeze the Case and evaluator boundary used by a Run.
- API keys belong to organization administration and do not replace Project permissions.
- Imports and evaluations can be asynchronous.
- External outputs create non-Harness reference evidence.
- Teammately evidence informs customer review; it does not record a customer deployment decision.
## Source confidence
Code-backed: current Case, Reference Materials, Expert Contributions, and Benchmark Evaluation output-mapping surfaces prove the product data-movement boundaries described here. Exact Public API serialization remains separately publication-gated until the refreshed external service implements its contract.
## Related task pages
{% related-card-grid title="Integration workflows" %}
- [Import case examples](/docs/integrations/import-case-examples)
- [Reference Materials](/docs/agent-setup/reference-materials)
- [Connect model outputs](/docs/integrations/connect-model-outputs)
- [Expert UI](/docs/integrations/reviewer-workspace)
- [Check candidate correctness](/docs/integrations/check-candidate-correctness)
{% /related-card-grid %}
---
id: object-model.benchmarks
title: Benchmarks
summary: Understand a Benchmark as the durable program that owns benchmark-scoped coverage, evidence boundaries, evaluations, and improvement work.
kind: reference
product_area: object_model
status: stable
updated: 2026-09-07
canonical: /docs/object-model/benchmarks
---
# Benchmarks
## Definition
A Benchmark is the durable project object for one intended evaluation program. It owns benchmark-scoped work across Benchmark Datasets, Coverage Management, Expert Contributions, Benchmark Evaluations, and Improve while its selected Cases, governed standards, and candidate systems evolve.
The current Benchmark Dataset is editable. A Dataset Snapshot freezes selected Case membership, and a Benchmark Version provides the immutable boundary consumed by Runs. A Benchmark is therefore not a Snapshot, Benchmark Version, Run, or score.
## Fields, states, or lifecycle rules
- The Benchmark identity persists across changes to its current Dataset, coverage work, standards, and Harness candidates.
- Benchmark Datasets owns selected Cases and immutable Dataset Snapshots.
- A Benchmark Version fixes the evidence boundary used by a Run.
- Benchmark membership should be shaped by coverage work, not by whichever Cases are easiest to run.
- A Run result is weak if the Benchmark Version and saved Harness Version behind it are unclear.
- This page documents object semantics, not public execution, export, rate-limit, or API guarantees.
## Related objects
Benchmarks should be read with [Cases](/docs/assets/cases), [Policies](/docs/object-model/policies), [Rubrics](/docs/object-model/rubrics), [Coverage Engineering](/docs/coverage-engineering), and [Benchmark Evaluations](/docs/benchmark-evaluations). Use [Run an evaluation](/docs/benchmark-evaluations/run-evaluation) when the next step is execution.
{% example-demo title="Benchmarks boundary" %}
Raw case: A team refreshes coverage after finding unsupported compatibility claims.
Benchmark version: The refreshed version includes new unsupported-claim cases and the approved compatibility rubric.
Run: The candidate is evaluated against that version.
Interpretation: If the score drops, reviewers can see that the benchmark became harder instead of assuming the candidate behavior changed.
{% /example-demo %}
## Source confidence
Code-backed: the Benchmark type and workspace establish durable Benchmark identity; Benchmark Datasets → Snapshots establishes immutable Dataset boundaries; the evaluation-runs route consumes a specific Benchmark Version. The public object definition does not imply an execution or export API.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Benchmarks](/docs/coverage-engineering/benchmarks)
- [Benchmarks and versions](/docs/concepts/benchmarks-and-versions)
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Product quickstart](/docs/quickstart)
- [Task index](/docs/operating-manual/task-index)
{% /related-card-grid %}