# Product quickstart
Generated: 2026-09-13T04:32:40.343Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: quickstart.product
title: Product quickstart
summary: Configure one project foundation, one benchmark slice, one expert contribution, one evaluation, and one evidence-backed improvement.
kind: quickstart
product_area: introduction
status: stable
updated: 2026-09-07
canonical: /docs/quickstart
---
# Product quickstart
Run one narrow correctness loop. The goal is not a large benchmark; it is a traceable chain from project context and deliberate coverage to expert-grounded standards, a versioned evaluation, and one justified next change.
## When to use it
Use this path for a new project or for an existing AI system whose correctness work is scattered across documents, examples, and informal expert feedback. Choose one behavior slice with a clear specialist owner.
## Decision checkpoint
| Starting point | First action | Ready to continue when... |
| --- | --- | --- |
| Agents do not understand the product or domain | Complete Agent Setup | Project Context and controlling Reference Materials are inspectable |
| Cases arrive in inconsistent shapes | Configure Project Input Schema | One input architecture and any required case materials are declared |
| Important behavior is not represented deliberately | Define Coverage Facets | Dimensions, Project Topics, and Case Construction Patterns name the slice |
| Correctness depends on tacit judgment | Request an Expert Contribution | The expert's scope, selected evidence, and required decisions are explicit |
| Cases and standards are ready | Create a benchmark snapshot and evaluate a saved Harness version | Exact cases, rubrics, candidate, and settings are bound to the Run |
## Prerequisites
- A Teammately project for the specialist AI behavior.
- An accountable project operator and at least one domain expert.
- A small number of representative examples or enough Reference Materials to construct them.
- A candidate system that can be represented by a saved Harness version before evaluation.
## Before and after
| Before | Action | After | Stop if... |
| --- | --- | --- | --- |
| Domain context is implicit | Write the Project Agent Brief and connect Reference Materials | Agents have explicit project understanding | Controlling sources are missing or contradictory without an owner |
| Inputs and supporting artifacts vary | Save Project Input Schema | Cases share one canonical content contract | Existing cases cannot satisfy the proposed schema |
| Expert knowledge is tacit | Request and complete a focused contribution | Policies, rubrics, cases, or coverage observations can be materialized | The request asks for a label without the evidence needed to explain it |
| Candidate behavior is anecdotal | Evaluate a saved Harness version | Results are traceable to cases and applicable rubrics | Dataset snapshot or candidate version is ambiguous |
| A weakness is confirmed | Start an Improvement Session from evidence | Candidate work follows a bounded Goal Contract | The requested outcome has no pinned measurement binding |
## Steps
1. Open or create the project and write the Project Agent Brief in **Agent Setup → Project Context**.
2. Add controlling knowledge through **Agent Setup → Reference Materials → Materials**, then inspect the published blocks in **Indexed Reference**.
3. Configure **Project Settings → Input Schema**. Select plain text, chat, or structured input and declare required case materials and accepted artifact families.
4. Create the smallest useful set of Coverage Facets: a Dimension, relevant Project Topics, and a Case Construction Pattern for the chosen behavior slice.
5. Add or construct cases in Assets, then select the intended cases in **Benchmark Datasets**. Confirm Representation and create or choose the appropriate snapshot.
6. In **Expert Contributions**, request one focused contribution. Select the expert, state the objective, attach or select the relevant cases, and include only the contribution components needed to resolve the question.
7. Inspect the completed contribution and materialize accepted policies, rubrics, cases, or coverage observations through their owning surfaces.
8. Save an exact Harness version. In **Benchmark Evaluations**, configure and run it against the selected benchmark version.
9. Inspect Dashboard and List results before using Compare or Arena. Trace important movement to case-level rubric evidence and run metadata.
10. If a candidate change is justified, open **Improve**, start from the relevant evidence, prepare and confirm the Goal Contract, and evaluate candidate work through the canonical Run path.
## Object and state changes
This path can create or update Project Context, Reference Materials items and indexed blocks, Project Input Schema, Coverage Facets, Cases, benchmark dataset membership and snapshots, Contributions, contributed artifacts, policies, rubrics, Harness drafts and saved versions, Runs, evaluation results, and Improvement Sessions. Each object keeps its own authority boundary; completing one step does not automatically approve or materialize every downstream artifact.
## Success criteria
- Another operator can identify the project context and source material used by agents.
- The case set conforms to the Project Input Schema and represents a named coverage slice.
- Expert judgment is attributable to a completed Contribution and its accepted artifacts.
- The evaluation binds an exact benchmark version to an exact saved Harness version.
- Any improvement work starts from pinned evidence and records its Goal Contract, candidate results, and current frontier.
## Common failure modes
- Treating Reference Materials as approved policies.
- Asking experts broad questions without selected cases or a concrete contribution objective.
- Evaluating an unsaved Harness draft or an unclear benchmark snapshot.
- Reading only an aggregate score and skipping failed case/rubric pairs.
- Starting improvement before the target and measurement evidence are resolved.
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Agent Setup](/docs/agent-setup)
- [Project Input Schema](/docs/project-settings/input-schema)
- [Expert Contributions](/docs/expert-contributions)
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Improve](/docs/improve)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Dataset upload](/docs/troubleshooting/dataset-upload)
- [Unclear cases](/docs/troubleshooting/unclear-cases)
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
{% /related-card-grid %}
## Source confidence
Doctrine-backed: this quickstart connects the current public story to code-backed product surfaces. Follow the linked pages for exact states and controls.
---
id: intro.what-is-teammately
title: What is Teammately?
summary: Understand Teammately as correctness infrastructure for building trustworthy specialist AI with expert judgment and AI agents.
kind: concept
product_area: introduction
status: stable
updated: 2026-08-22
canonical: /docs/introduction/what-is-teammately
---
# What is Teammately?
Teammately is correctness infrastructure for teams building specialist AI. It turns in-house experts' judgment into an operating system for designing benchmark coverage, making correctness explicit, constructing challenging cases, evaluating candidate behavior, and deciding what to improve next. AI agents prepare and connect the work so scarce expert attention is spent on consequential judgment rather than manual organization.
> Category boundary
>
> Teammately centers the definition and development of trustworthy AI behavior. Logs, traces, model endpoints, coding environments, and external data can enter the workflow, but the product's durable value is the connected correctness system built from expert judgment, cases, standards, evaluation evidence, and improvement history.
## Definition
The system has five connected capabilities. [Coverage Engineering](/docs/coverage-engineering) defines the behavior space a benchmark must represent. [Correctness Elicitation](/docs/concepts/correctness-elicitation) turns tacit preferences, exceptions, and disagreements into policies, applicability conditions, and binary rubrics. [Weave](/docs/concepts/weave) constructs cases, response variants, case materials, and—where supported—worlds from that structure. [Trialground](/docs/concepts/trialground) evaluates Harnesses and weights against benchmark Cases and preserves responses and Rubric results. [Coevolve](/docs/concepts/coevolve) explores candidate improvements and keeps every retained direction tied to current benchmark evidence.
These capability names explain how the system works. Procedures use the labels visible in the product, such as Correctness Governance, Agent Setup, Benchmark Datasets, Coverage Management, Expert Contributions, Benchmark Evaluations, and Improve.
## Decision checkpoint
| If the team needs... | Capability | Product surfaces to open |
| --- | --- | --- |
| A deliberate map of important behavior | Coverage Engineering | Coverage Facets and Coverage Management |
| Reusable standards grounded in specialist judgment | Correctness Elicitation | Correctness Governance and Expert Contributions |
| Challenging cases and supporting materials | Weave | Assets, Benchmark Datasets, Case Construction Patterns, and Case Foundry |
| Repeatable evidence about candidate behavior | Trialground | Harnesses and Benchmark Evaluations |
| Evidence-backed candidate improvement | Coevolve | Improve and Improvement Sessions |
## Why teams use it
A benchmark score cannot define correctness on its own. Specialist systems depend on domain rules, exceptions, source authority, interaction patterns, and consequences that generic criteria do not capture. Teammately gives experts and AI engineers a shared artifact graph: an expert contribution can inform a policy, a policy can produce a rubric, a coverage gap can motivate a case, a case can expose a harness weakness, and an evaluation can become the starting evidence for an Improvement Session.
This reuse is the practical meaning of scaling expert judgment. Teammately prepares coverage structure, candidate cases, possible standards, and unresolved questions before asking an expert. The expert's response remains attributable and can be materialized into governed artifacts instead of disappearing into meeting notes.
## Product scope
Project-level foundations hold reusable knowledge and assets: Correctness Governance, Coverage Facets, Assets, Agent Setup, and Project Settings. Benchmark workspaces bind those foundations to a concrete evaluation program through Benchmark Datasets, Coverage Management, Expert Contributions, Benchmark Evaluations, and Improve.
Teammately preserves correctness evidence and makes the next engineering question inspectable. Customer teams remain responsible for downstream product, governance, deployment, and operational choices.
> Human ownership
>
> AI agents can prepare, draft, classify, generate, evaluate, and propose. A suggestion is not a governed policy, accepted expert contribution, benchmark membership decision, or retained candidate merely because an agent produced it. Use the state shown by the owning product surface.
{% example-demo title="Grounded enterprise search" %}
Coverage Engineering identifies conflicting-current-source questions as an important behavior slice. Correctness Elicitation records the expert rule that material claims must cite the controlling source or state uncertainty. Weave creates cases with current and superseded documents. Trialground evaluates a retrieval harness and exposes unsupported blends of the two sources. Coevolve starts from those failures, tests a source-selection change, and retains only candidates supported by evaluation evidence.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Product quickstart](/docs/quickstart)
- [The correctness loop](/docs/product-loop)
- [First correctness loop](/docs/operating-manual/first-correctness-loop)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Product map](/docs/getting-oriented/product-map)
- [Key objects and relationships](/docs/getting-oriented/key-objects-and-relationships)
- [Product boundaries](/docs/introduction/product-boundaries)
{% /related-card-grid %}
## Source confidence
Doctrine-backed: this page follows the current public top-page story and the approved product-to-UI mapping. Linked code-backed pages define exact routes, states, and controls.
---
id: product-loop
title: The Teammately correctness loop
summary: See how coverage, elicitation, case construction, evaluation, and improvement reinforce one another.
kind: concept
product_area: introduction
status: stable
updated: 2026-08-22
canonical: /docs/product-loop
---
# The Teammately correctness loop
The correctness loop is how a team repeatedly turns domain knowledge into stronger AI behavior. It follows the five public capabilities while preserving a trace from every result back to the project context, expert contribution, case, policy, rubric, benchmark version, Harness version, and evaluation setting that made the result meaningful.
## Definition
1. **Design coverage.** Establish Dimensions, Project Topics, and Case Construction Patterns, then decide which combinations the benchmark must represent.
2. **Elicit correctness.** Use focused expert contributions to resolve policies, exceptions, applicability, disagreements, and binary rubric language.
3. **Construct the challenge set.** Create or import canonical cases, attach required materials, generate difficult variants, and curate benchmark dataset membership.
4. **Evaluate behavior.** Run an exact saved Harness Version against an exact Benchmark Version and inspect responses, Case-level Rubric evidence, comparisons, and rankings.
5. **Improve from evidence.** Start an Improvement Session with a bounded Goal Contract, explore candidates, evaluate them through the canonical path, and retain a current frontier.
6. **Return new learning.** Update coverage, correctness, cases, or the candidate according to what the evidence actually showed.
## Decision checkpoint
| Evidence says... | Responsible part of the loop | Change first |
| --- | --- | --- |
| Important behavior has no cases | Coverage Engineering or Weave | Coverage facet, construction pattern, or case set |
| Experts cannot apply the standard consistently | Correctness Elicitation | Policy scope, applicability, or rubric wording |
| A case cannot be interpreted or executed reliably | Weave and Project Input Schema | Input shape, case material, or world boundary |
| One saved candidate fails applicable rubrics | Trialground | Harness candidate or its runtime configuration |
| Several candidate branches improve different slices | Coevolve | Goal constraints, next experiment, or retained frontier |
| Result movement cannot be explained | Benchmark version and evaluation boundary | Versions, settings, mapping, or run metadata before any product change |
## How expert effort compounds
The loop should ask an expert only after agents have prepared the relevant structure and evidence. A Contribution can include selected Cases, source attachments, scoped statements, draft Policies, Rubric questions, or coverage uncertainty. Completed expert work can materialize as an attributable contributed Policy, Rubric, Case, or coverage observation through the owning workflow.
That same judgment can guide future case construction, determine which rubrics apply during evaluation, and identify missing correctness during improvement. Reuse across the loop is more valuable than maximizing the number of disconnected review actions.
## How product scope changes through the loop
Project foundations are reusable. Project Context, Reference Materials, policies, rubrics, Coverage Facets, Cases, and Harnesses do not belong to only one benchmark. A benchmark workspace selects and versions the relevant subset, manages coverage, coordinates contributions, evaluates candidates, and records improvement.
This scope distinction prevents accidental drift. Editing a project-level policy may affect several benchmarks. Changing dataset membership should create a new benchmark evidence boundary. Saving a Harness draft is different from selecting an exact saved Harness version for a Run.
## Before and after
| Before | Loop work | After |
| --- | --- | --- |
| Domain knowledge is distributed across people and files | Agent Setup and Correctness Elicitation organize it | Project context and governed correctness artifacts are inspectable |
| Examples are convenient rather than deliberate | Coverage Engineering and Weave shape the challenge set | Dataset representation and missing coverage are explicit |
| Candidate behavior is discussed from anecdotes | Trialground runs a versioned evaluation | Case-level rubric evidence and comparisons are available |
| Improvement is a sequence of untracked edits | Coevolve starts from pinned evidence | Candidate branches, receipts, chronology, and current frontier remain connected |
{% example-demo title="Changing a retrieval harness" %}
An evaluation shows failures only when current and superseded documents appear together. The team first confirms that the coverage slice and grounding rubric are valid. An Improvement Session pins those cases and the failing Harness version, then tests source-date filtering and citation-selection candidates. A stronger candidate becomes part of the current frontier only after a canonical evaluation produces the expected rubric evidence. If the work uncovers an unseen source-conflict pattern, that observation returns to Coverage Management.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Product quickstart](/docs/quickstart)
- [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation)
- [Start an Improvement Session](/docs/improve/start-improvement-session)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Product map](/docs/getting-oriented/product-map)
- [Project Input Schema](/docs/project-settings/input-schema)
- [Expert Contributions](/docs/expert-contributions)
{% /related-card-grid %}
## Source confidence
Doctrine-backed: this page explains the approved operating loop. Linked product pages are the authority for exact controls and lifecycle states.
---
id: agent-setup.overview
title: Agent Setup
summary: Configure Project Context and Reference Materials so Teammately agents have reusable project understanding before contribution work.
kind: concept
product_area: agent_setup
status: stable
updated: 2026-09-07
canonical: /docs/agent-setup
---
# Agent Setup
Agent Setup is the project-level workspace for configuring reusable project understanding before Teammately agents prepare or conduct expert contribution work. Expert-facing presentation and contribution-specific behavior are configured through Assets and the Contribution workflow.
## Definition
Agent Setup contains one project-understanding group:
- **Project Understanding:** Project Context and Reference Materials.
These settings are reusable project foundations rather than settings for one benchmark or one expert. Review Screens and Comparison Directions are project Assets, not Agent Setup tabs.
Project Context contains the Project Agent Brief. Reference Materials uses the Materials and Indexed Reference tabs to organize project knowledge for agents. A Contribution selects its benchmark-specific objective, components, and agent behavior; Review Screens control reusable expert-facing presentation from Assets.
## Decision checkpoint
| Need | Open | Keep distinct from... |
| --- | --- | --- |
| Explain the project, target behavior, and stable operating context | Project Context | Project name or memo in General settings |
| Supply manuals, sites, repositories, or files to agents | Reference Materials | Governed policies, rubrics, and case materials |
| Set benchmark-specific agent behavior | Expert Contribution | Project-wide context and screen configuration |
| Guide meaningful response variation | Assets → Comparison Directions | Coverage facets, generated cases, or approved standards |
| Configure what an expert sees while reviewing | Assets → Review Screens | Contribution objectives and selected cases |
## Project and benchmark scope
Agent Setup belongs to the project because the same project context may support many benchmarks. A benchmark-specific Contribution still selects its own objective, expert, cases, attachments, contribution components, and agent behavior. Review Screens and Comparison Directions are authored under Assets and selected when the Contribution needs them. Agent Setup provides the reusable understanding foundation; it does not create or schedule contribution work by itself.
Changes can affect future agent preparation. Before making broad edits, inspect active benchmark work and confirm whether the new context should apply across the project. A narrow contribution-specific request belongs in the Contribution rather than in permanent Agent Setup.
## Authority boundaries
Reference Materials can inform agents but does not automatically create policies or rubrics. Review Screens change presentation and requested inputs, not the meaning of the underlying case or standard. Contribution configuration guides agent behavior but cannot supply human approval.
These boundaries make contribution evidence interpretable. Another operator can distinguish what the project told the agent, what evidence the contribution supplied, what the agent proposed, and what the expert decided.
{% example-demo title="Policy-review preparation" %}
Project Context explains that the assistant must prioritize the current procurement agreement. Reference Materials indexes the agreement repository. A Review Screen shows the controlling document and relevant case-material fields. The benchmark Contribution then asks an expert to decide which behavior should become policy.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Product quickstart](/docs/quickstart)
- [Request an Expert Contribution](/docs/expert-contributions/request-contribution)
- [Configure Project Input Schema](/docs/project-settings/input-schema)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Project Context](/docs/agent-setup/project-context)
- [Reference Materials](/docs/agent-setup/reference-materials)
- [Comparison Directions](/docs/assets/comparison-directions)
- [Review Screens](/docs/assets/review-screens)
{% /related-card-grid %}
## Source confidence
Code-backed: the active Agent Setup layout and project navigation define these groups, labels, and routes.
---
id: expert-contributions.overview
title: Expert Contributions
summary: Coordinate benchmark-scoped expert work, attributable judgment, governed artifacts, and the decisions that move correctness forward.
kind: concept
product_area: expert_contributions
status: stable
updated: 2026-09-07
canonical: /docs/expert-contributions
---
# Expert Contributions
Expert Contributions is the benchmark-scoped workspace for requesting, conducting, and materializing specialist work. It coordinates the expert, objective, selected evidence, task sequence, checkpoints, attributable responses, and contributed artifacts needed to move a benchmark forward.
## Definition
The administrator workspace contains **Overview**, **Contributions**, **Contributed Artifacts**, and **Logs & Status**. **Request Contribution** opens the composer for a new contribution. The expert follows a contribution-specific experience that can contain form, chat, interview, and case-review tasks, along with checkpoints and completion states.
A Contribution is the unit of requested expert effort. It replaces broad workflow configuration with a bounded statement of what this benchmark needs from this expert now. The work can result in contributed policies, rubrics, cases, or coverage observations without flattening all expert activity into one generic approval record.
## Decision checkpoint
| Need | Contribution element | Result to inspect |
| --- | --- | --- |
| Resolve a specific benchmark question | Contribution statement and scoped objectives | The expert can explain the requested decision |
| Ground work in concrete behavior | Selected or designated cases | Case-level responses remain attributable |
| Supply supporting knowledge | Attachments and scoped statements | The expert sees the relevant source boundary |
| Choose the right interaction | Form, chat, interview, or case review task | Task output matches the kind of judgment needed |
| Confirm consequential learning | Checkpoint | Accepted, revised, or unresolved state is explicit |
| Reuse the result | Contributed Artifacts | Policies, rubrics, cases, and coverage observations retain provenance |
## Lifecycle and status
The durable Contribution statuses are `PREPARING_DIRECTION`, `AWAITING_DIRECTION_ALIGNMENT`, `MATERIALIZING_TASKS`, `READY`, `IN_PROGRESS`, `COMPLETED`, and `CANCELLED`. The interface presents these as planning direction, waiting for alignment, preparing tasks, ready, active, completed, or cancelled. The exact task sequence can vary by Contribution.
Realtime updates and durable transitions help the administrator and expert see current progress without inventing completion. A waiting state, checkpoint, or finalization step should be shown as such. Completing the expert experience does not imply that every proposed artifact has been accepted into its project-level owner.
## Contribution evidence
Logs & Status exposes operational and engagement records. Contributed Artifacts organizes materialized or contributed cases, policies, rubrics, and new coverage observations. Correctness Governance, Assets, or Coverage Management owns the resulting project or benchmark artifact after materialization.
This model improves return on expert effort. Agents prepare focused work from project context, indexed material, benchmark cases, and unresolved questions. The expert supplies the authority; the result can be reused across standards, coverage, evaluation, and improvement.
{% example-demo title="Resolve source authority" %}
A benchmark contains cases where an operational runbook conflicts with a newer policy page. The operator requests a Contribution from the policy owner, selects the conflicting cases, attaches both sources, and uses case review plus a checkpoint. The expert establishes which source controls, contributes a scoped policy and rubric, and records one coverage observation for an unrepresented exception.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Request an Expert Contribution](/docs/expert-contributions/request-contribution)
- [Complete an Expert Contribution](/docs/expert-contributions/complete-contribution)
- [Build policies and rubrics](/docs/operating-manual/build-policies-and-rubrics)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Contributed Artifacts](/docs/expert-contributions/contributed-artifacts)
- [Contribution lifecycle and status](/docs/expert-contributions/lifecycle-and-status)
- [Logs & Status](/docs/expert-contributions/logs-and-status)
- [Agent Setup](/docs/agent-setup)
- [Human Approval Boundaries](/docs/governance/human-approval-boundaries)
{% /related-card-grid %}
## Source confidence
Code-backed: the active benchmark workspace, Contribution dashboard, composer, administrator detail, and expert routes support the scope, task, status, and artifact model described here.
---
id: benchmark-evaluations.overview
title: Benchmark Evaluations
summary: Run and inspect exact Harness Versions against an immutable Benchmark Version through Dashboard, List, Arena, and Compare.
kind: concept
product_area: benchmark_evaluations
status: stable
updated: 2026-09-13
canonical: /docs/benchmark-evaluations
---
# Benchmark Evaluations
Benchmark Evaluations is the version-scoped workspace for executing and comparing candidate systems. The active top-level tabs are **Dashboard**, **List**, **Arena**, and **Compare**. Every managed Run binds an exact saved Harness Version to the immutable Benchmark Version shown in the route.
> Evaluation boundary
>
> Interpret evidence inside its recorded Benchmark Version, Harness Version, Run or Run Group, evaluator set, and metadata. Run counts belong to launches. Additional launches add evidence without rewriting earlier Runs.
## Surfaces and objects
Dashboard summarizes progress, leaderboards, rank progression across Runs, and available resource telemetry. List is segmented into **Runs**, **Evaluation results**, and **Traces / Spans**. The results segment summarizes Case outcomes and Policy or Rubric failures. Arena compares candidate pairs across governed metrics. Compare is a symmetric matrix of Harness Versions across selected evidence rows.
A Run Group can collect one standard attempt or repeated attempts. A Run records one candidate execution and its per-Case progress. Evaluation results record the admitted Policy and Rubric outcomes. Costs, tokens, and latency are telemetry only when the provider or execution path captured them.
> Traces / Spans capability fence
>
> The List navigation exposes Traces / Spans, but the current benchmark API does not expose evaluation execution traces. Do not claim that trajectories, spans, private reasoning, or tool traces can be inspected from Benchmark Evaluations today.
## Decision checkpoint
| Need | Open | Evidence to preserve |
| --- | --- | --- |
| Configure and launch managed Runs | Evaluation Settings and New evaluation run | Machine, saved Harness Versions, and per-Harness Run counts |
| Start candidate execution | Run modal | Exact Harness and Benchmark Versions |
| Inspect status and output summaries | List → Runs or Evaluation results | Run Group, attempt, Case counts, incomplete state |
| Compare candidate pairs | Arena | Metric family, pair count, only-A, only-B, shared failures |
| Compare many candidates by governed rows | Compare | Harness columns and chosen Case or facet row mode |
| Admit external reference outputs | Output mapping | Case mapping, attempt assignment, insert/update report |
## Rankings and repeated sampling
Dashboard aggregates compatible observed Runs for each saved Harness Version across launches. Average score weights Runs equally. Supported binary views report passed at least once or passed every time over the observed case outcomes. Counts and missing evidence are shown; unequal counts do not prevent comparison. Historical group metrics retain their recorded meanings.
Ranking is a routing signal. A candidate can lead overall while failing required Policy or high-impact Rubric evidence. Use Arena or Compare to locate the disagreement and List to confirm completeness before starting Improve work.
## External outputs
Uploaded or API-supplied reference outputs create output-only Runs that can be scored and inspected in List. They are not saved Harness Versions and therefore cannot be optimized in Improve or selected as Harness columns in Compare or Arena.
{% example-demo title="Example: repeated evaluation without evidence drift" %}
A team launches three Runs of Harness Version 8 and one Run of Version 11 against the same Benchmark Version. Both appear with their evidence counts. A later launch of Version 11 adds two Runs to its aggregate evidence without changing either launch group. The team can inspect individual Runs before deciding whether more evidence is useful.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Configure evaluation execution](/docs/benchmark-evaluations/execution-settings)
- [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation)
- [Inspect evaluation results](/docs/benchmark-evaluations/inspect-results)
- [Use Arena and rankings](/docs/benchmark-evaluations/arena-and-rankings)
- [Compare Harness Versions](/docs/benchmark-evaluations/compare)
- [Map external outputs](/docs/benchmark-evaluations/output-mapping)
{% /related-card-grid %}
## Source confidence
Code-backed: the active version-scoped workspace, settings, Run modal, List segments, Dashboard, Arena, and Compare routes define the current evaluation model and capability fences.