# Product boundaries Generated: 2026-09-13T04:39:39.134Z Source build: local Canonical docs: https://teammately.ai/docs --- id: intro.product-boundaries title: Product boundaries summary: Understand what Teammately owns across correctness specification, benchmark development, evaluation, and improvement. kind: concept product_area: introduction status: stable updated: 2026-09-07 canonical: /docs/introduction/product-boundaries --- # Product boundaries Teammately owns the correctness system that connects specialist judgment to deliberate benchmark coverage, executable standards, constructed cases, evaluation evidence, and improvement history. This page distinguishes that system from adjacent inputs and downstream responsibilities. > Adjacent systems are inputs > > Logs, traces, source repositories, model endpoints, coding environments, and external evaluation results can supply material or receive work. Their presence does not change the Teammately ownership boundary: Teammately governs the connected correctness artifacts and the evidence produced from them. ## Definition The product boundary follows artifacts and authority. Teammately can index project knowledge, prepare an expert contribution, materialize an accepted policy or rubric, construct a case, execute an evaluation through a managed harness, and coordinate an Improvement Session. It preserves which inputs, versions, settings, and human decisions produced the resulting evidence. Customer teams own the AI system outside that evidence graph and the action taken afterward. Teammately can prepare a scoped package for an external coding worker, but it does not claim private work performed outside the product. It can show benchmark evidence, but it does not turn that evidence into an automatic downstream decision. ## Decision checkpoint | Area | Teammately owns | Boundary | | --- | --- | --- | | Project knowledge and agent context | Materials, Indexed Reference, Project Context, and Contribution-scoped agent behavior | Reference material is not automatically a governed policy or rubric | | Expert work | Contribution scope, tasks, checkpoints, attributable responses, and contributed artifacts | Agent preparation does not substitute for the expert's judgment | | Cases and worlds | Canonical case input, case materials, generated artifacts, and verified world references | Static materials and executable environments remain distinct | | Evaluation | Benchmark Versions, saved Harness Versions, settings, Runs, responses, Rubric results, and comparisons | A score alone does not explain correctness; execution traces are not currently exposed | | Improvement | Goal Contracts, candidates, evaluation receipts, frontiers, and chronology | External worker activity is represented only when returned through the defined contract | | Downstream action | Inspectable correctness evidence and review context | The customer decides what operational action follows | ## Human and agent authority AI agents scale preparation and exploration. They can organize source material, propose coverage structure, draft possible standards, generate cases, evaluate candidates, and suggest improvement directions. The owning surface determines when an artifact becomes durable or governed. An agent proposal does not silently acquire expert authority. Expert Contributions make this boundary explicit. The product can prepare focused questions and relevant evidence, while the domain specialist supplies the judgment. Correctness Governance records policies and rubrics as governed project assets. Improvement Sessions can branch candidate hypotheses, but retained candidates require observable evaluation evidence. ## Data and execution boundary Project Input Schema controls the accepted shape of case input and materials. Static context remains part of case content or case-material references. An executable or queryable environment uses a world reference and follows a separate runtime boundary. Public docs describe the behavior visible through stable product surfaces; they do not promote internal storage or service structures into customer-facing contracts. Similarly, the presence of Harness Assets and managed Runs does not imply that Teammately owns a customer's model registry, production telemetry, or deployment system. A harness is the executable candidate boundary used by a benchmark evaluation. {% example-demo title="External coding worker" %} An Improvement Session starts from failed grounding Cases and a confirmed Goal Contract. Teammately prepares a scoped package for a coding worker with the pinned target and evidence. The worker's private activity is outside the product boundary. A returned Harness Version and canonical evaluation request become observable candidates; their Rubric results enter the session record, while Improve may add a safe narrated trajectory of observable session activity. {% /example-demo %} ## Related workflows {% related-card-grid title="Related workflows" %} - [The correctness lifecycle](/docs/introduction/correctness-lifecycle) - [Start an Improvement Session](/docs/improve/start-improvement-session) - [Use Reference Materials](/docs/agent-setup/reference-materials) {% /related-card-grid %} ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Human Approval Boundaries](/docs/governance/human-approval-boundaries) - [What AI Features Can and Cannot Do](/docs/governance/what-ai-features-can-and-cannot-do) - [Project Input Schema](/docs/project-settings/input-schema) {% /related-card-grid %} ## Source confidence Doctrine-backed: this page states product ownership and authority boundaries. Exact UI and execution behavior is delegated to linked code-backed pages. --- id: intro.what-is-teammately title: What is Teammately? summary: Understand Teammately as correctness infrastructure for building trustworthy specialist AI with expert judgment and AI agents. kind: concept product_area: introduction status: stable updated: 2026-08-22 canonical: /docs/introduction/what-is-teammately --- # What is Teammately? Teammately is correctness infrastructure for teams building specialist AI. It turns in-house experts' judgment into an operating system for designing benchmark coverage, making correctness explicit, constructing challenging cases, evaluating candidate behavior, and deciding what to improve next. AI agents prepare and connect the work so scarce expert attention is spent on consequential judgment rather than manual organization. > Category boundary > > Teammately centers the definition and development of trustworthy AI behavior. Logs, traces, model endpoints, coding environments, and external data can enter the workflow, but the product's durable value is the connected correctness system built from expert judgment, cases, standards, evaluation evidence, and improvement history. ## Definition The system has five connected capabilities. [Coverage Engineering](/docs/coverage-engineering) defines the behavior space a benchmark must represent. [Correctness Elicitation](/docs/concepts/correctness-elicitation) turns tacit preferences, exceptions, and disagreements into policies, applicability conditions, and binary rubrics. [Weave](/docs/concepts/weave) constructs cases, response variants, case materials, and—where supported—worlds from that structure. [Trialground](/docs/concepts/trialground) evaluates Harnesses and weights against benchmark Cases and preserves responses and Rubric results. [Coevolve](/docs/concepts/coevolve) explores candidate improvements and keeps every retained direction tied to current benchmark evidence. These capability names explain how the system works. Procedures use the labels visible in the product, such as Correctness Governance, Agent Setup, Benchmark Datasets, Coverage Management, Expert Contributions, Benchmark Evaluations, and Improve. ## Decision checkpoint | If the team needs... | Capability | Product surfaces to open | | --- | --- | --- | | A deliberate map of important behavior | Coverage Engineering | Coverage Facets and Coverage Management | | Reusable standards grounded in specialist judgment | Correctness Elicitation | Correctness Governance and Expert Contributions | | Challenging cases and supporting materials | Weave | Assets, Benchmark Datasets, Case Construction Patterns, and Case Foundry | | Repeatable evidence about candidate behavior | Trialground | Harnesses and Benchmark Evaluations | | Evidence-backed candidate improvement | Coevolve | Improve and Improvement Sessions | ## Why teams use it A benchmark score cannot define correctness on its own. Specialist systems depend on domain rules, exceptions, source authority, interaction patterns, and consequences that generic criteria do not capture. Teammately gives experts and AI engineers a shared artifact graph: an expert contribution can inform a policy, a policy can produce a rubric, a coverage gap can motivate a case, a case can expose a harness weakness, and an evaluation can become the starting evidence for an Improvement Session. This reuse is the practical meaning of scaling expert judgment. Teammately prepares coverage structure, candidate cases, possible standards, and unresolved questions before asking an expert. The expert's response remains attributable and can be materialized into governed artifacts instead of disappearing into meeting notes. ## Product scope Project-level foundations hold reusable knowledge and assets: Correctness Governance, Coverage Facets, Assets, Agent Setup, and Project Settings. Benchmark workspaces bind those foundations to a concrete evaluation program through Benchmark Datasets, Coverage Management, Expert Contributions, Benchmark Evaluations, and Improve. Teammately preserves correctness evidence and makes the next engineering question inspectable. Customer teams remain responsible for downstream product, governance, deployment, and operational choices. > Human ownership > > AI agents can prepare, draft, classify, generate, evaluate, and propose. A suggestion is not a governed policy, accepted expert contribution, benchmark membership decision, or retained candidate merely because an agent produced it. Use the state shown by the owning product surface. {% example-demo title="Grounded enterprise search" %} Coverage Engineering identifies conflicting-current-source questions as an important behavior slice. Correctness Elicitation records the expert rule that material claims must cite the controlling source or state uncertainty. Weave creates cases with current and superseded documents. Trialground evaluates a retrieval harness and exposes unsupported blends of the two sources. Coevolve starts from those failures, tests a source-selection change, and retains only candidates supported by evaluation evidence. {% /example-demo %} ## Related workflows {% related-card-grid title="Related workflows" %} - [Product quickstart](/docs/quickstart) - [The correctness loop](/docs/product-loop) - [First correctness loop](/docs/operating-manual/first-correctness-loop) {% /related-card-grid %} ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Product map](/docs/getting-oriented/product-map) - [Key objects and relationships](/docs/getting-oriented/key-objects-and-relationships) - [Product boundaries](/docs/introduction/product-boundaries) {% /related-card-grid %} ## Source confidence Doctrine-backed: this page follows the current public top-page story and the approved product-to-UI mapping. Linked code-backed pages define exact routes, states, and controls. --- id: governance.ai-feature-boundaries title: What AI Features Can and Cannot Do summary: Explain the difference between AI-assisted suggestions and approved correctness infrastructure. kind: reference product_area: governance status: stable updated: 2026-09-07 canonical: /docs/governance/what-ai-features-can-and-cannot-do --- # What AI Features Can and Cannot Do ## Definition AI-assisted features can help prepare correctness work, but they do not own the approval boundary. They may draft, classify, summarize, propose, or organize artifacts; a human still needs to approve governed standards and review context before benchmark evidence depends on them. Use this page when a reader needs to separate preparation from authority. A fluent suggestion can be useful, but it is not an approved Policy, Rubric, Case-scoped reference output, Benchmark evidence, or downstream customer decision by itself. > Human approval boundary > > AI features can draft, summarize, classify, or suggest; they do not approve policies, rubrics, benchmark evidence, or customer-owned decisions by themselves. ## Fields, states, or lifecycle rules - AI-assisted output can be draft material, review support, classification help, or summarization. - AI-assisted output should not be treated as approved standards, benchmark evidence, or customer-owned decisions without human approval. - Suggested policies, rubrics, and classifications need visible approval or rejection state before they affect governed evidence. - AI-suggested Comparison Directions are different from suggested standards: they can become active directions immediately, but they still do not approve policies, rubrics, cases, benchmark membership, or review context. - AI-assistance language does not establish model-provider behavior, data retention, compliance posture, or autonomous approval. - Source-backed product pages decide exact supported behavior; this page defines the public boundary. ## Related objects Read this with [Human Approval Boundaries](/docs/governance/human-approval-boundaries), [Comparison Directions](/docs/assets/comparison-directions), [Using Expert Judgment](/docs/concepts/correctness-elicitation), [Using Checkpoints](/docs/expert-contributions/complete-contribution), and [Noisy AI Suggestions](/docs/troubleshooting/noisy-ai-suggestions). {% example-demo title="What AI Features Can and Cannot Do boundary" %} Draft suggestion: An AI-assisted workflow proposes a new compatibility policy after reading rejected cases. Allowed use: The suggestion can help a reviewer start the policy draft. Not allowed as governed evidence: A benchmark should not cite the policy until an accountable human approves the artifact and its applicability. Interpretation: The AI feature accelerated preparation; the human approval boundary still controls whether the standard can govern benchmark evidence. {% /example-demo %} {% example-demo title="Example: active AI-suggested direction" %} AI-suggested direction: Teammately suggests a stale-source conflict Comparison Direction after project cases and ontology values show that gap. Allowed use: The direction can become active immediately and can guide preview examples or future variant generation. Not allowed as governed evidence: The direction does not approve a policy, rubric, generated case, benchmark membership, or downstream decision. Users still review generated examples and manage the direction normally. {% /example-demo %} ## Source confidence Code-backed: generation surfaces can propose Dimension schemas, synthesize candidate Cases, and suggest Comparison Directions; Policy approval and contributed-artifact state establish separate human-governance boundaries. These sources support the product-state distinction here, not claims about model providers, retention, or autonomous authority. ## Related task pages {% related-card-grid title="Related task pages" %} - [Human Approval Boundaries](/docs/governance/human-approval-boundaries) - [Comparison Directions](/docs/assets/comparison-directions) - [Using Expert Judgment](/docs/concepts/correctness-elicitation) - [Using Checkpoints](/docs/expert-contributions/complete-contribution) - [Product quickstart](/docs/quickstart) - [Task index](/docs/operating-manual/task-index) {% /related-card-grid %} --- id: orientation.end-to-end title: Operating Teammately end to end summary: Operate the current product from project foundations through benchmark coverage, expert contribution, evaluation, and improvement. kind: task product_area: operating_manual status: stable updated: 2026-09-07 canonical: /docs/getting-oriented/operating-teammately-end-to-end --- # Operating Teammately end to end Use this workflow to coordinate the full correctness system while keeping project foundations, benchmark work, expert authority, evaluation evidence, and candidate improvement separate. ## Decision checkpoint | Phase | Owning scope | Exit condition | | --- | --- | --- | | Establish project understanding | Project | Project Agent Brief and Indexed Reference are usable | | Define content and reusable assets | Project | Project Input Schema, Cases, Harnesses, and Coverage Facets are explicit | | Establish benchmark evidence | Benchmark | Dataset snapshot and coverage state are reviewable | | Resolve specialist correctness | Benchmark Contribution and project governance | Attributable artifacts have explicit lifecycle state | | Evaluate candidates | Benchmark version | Exact Runs and case/rubric evidence are available | | Improve behavior | Benchmark version | Goal Contract, candidates, receipts, and frontier are durable | ## Prerequisites - A workspace and project. - An accountable operator, domain expert, and AI engineer or candidate owner. - Source knowledge, examples, and a candidate system appropriate to the intended benchmark. ## Before and after | Before | Operation | After | | --- | --- | --- | | Agents lack a shared project model | Configure Agent Setup | Project understanding is reusable and inspectable | | Examples have inconsistent shape | Save Project Input Schema and prepare Cases | Inputs and materials share a canonical contract | | Coverage and correctness are implicit | Define Coverage Facets and request Contributions | Benchmark intent and specialist standards are explicit | | Candidate claims depend on anecdotes | Run Benchmark Evaluations | Evidence is bound to versions, cases, and rubrics | | Engineering iterations lack chronology | Use Improve | Goals, candidates, receipts, and current frontier stay connected | ## Steps 1. Configure Project Context and Reference Materials in **Agent Setup**. Create or select Comparison Directions and Review Screens under **Assets** when a Contribution needs them. 2. Save Project Input Schema and establish reusable Coverage Facets. 3. Create or import Cases and save candidate Harness versions under Assets. 4. Create or select a benchmark, configure Coverage Management, select Cases in Benchmark Datasets, inspect Representation, and preserve a snapshot. 5. Request focused Expert Contributions for unresolved standards, cases, or coverage. Reconcile contributed artifacts in their owning project or benchmark surfaces. 6. Confirm governed policies and rubrics and the benchmark version that will use them. 7. Run exact saved Harness versions through Benchmark Evaluations. Inspect Dashboard, List, Compare, Arena, and Run detail according to the question. 8. Start an Improvement Session only from evidence that identifies measurable candidate work. 9. Return newly discovered correctness or coverage gaps to Expert Contributions, Correctness Governance, Coverage Management, or Cases. ## Object and state changes This workflow touches project context, reference indexes, input schema, facets, assets, benchmark datasets and snapshots, coverage state, Contributions and contributed artifacts, policies, rubrics, Harness versions, Runs, evaluation results, and Improvement Sessions. Each object remains in its owning scope and retains historical evidence. ## Success criteria - Project and benchmark scope is explicit at every operation. - Agent preparation, expert judgment, and governed artifacts remain distinguishable. - Dataset and candidate versions make evaluation reproducible. - Improvement begins with a measurable goal and pinned evidence. - New learning returns to one responsible upstream artifact. ## Common failure modes - Putting benchmark-specific instructions into permanent Project Context. - Treating connected sources as approved standards. - Selecting generated Cases without case review or schema conformance. - Comparing Runs after multiple evidence boundaries changed. - Treating external-worker activity as observable before an artifact returns. {% example-demo title="Example: full grounding loop" %} A team indexes source repositories, defines source-authority coverage, imports canonical cases, and requests a Contribution to resolve conflicting guidance. The governed rubric enters a benchmark version, two saved Harness versions are compared, and an Improvement Session tests retrieval changes. A missing-source pattern discovered during improvement returns to Coverage Management. {% /example-demo %} ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Product map](/docs/getting-oriented/product-map) - [Key objects and relationships](/docs/getting-oriented/key-objects-and-relationships) - [Product boundaries](/docs/introduction/product-boundaries) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Expert Contribution problems](/docs/troubleshooting/expert-contributions) - [Benchmark runs](/docs/troubleshooting/benchmark-runs) - [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage) {% /related-card-grid %} ## Source confidence Doctrine-backed: the sequence follows the current public capability model and active product topology. Linked pages provide code-backed operation details.