# Regime Settings Generated: 2026-09-13T04:36:26.937Z Source build: local Canonical docs: https://teammately.ai/docs --- id: project-settings.regime title: Regime Settings summary: Inspect and publish the project scoring Regime that governs future Benchmark Versions. kind: reference product_area: project_settings status: stable updated: 2026-09-07 canonical: /docs/project-settings/regime --- # Regime Settings ## Definition Regime Settings define how approved and applicable Rubrics contribute to Case and Benchmark results. The Regime is part of the project contract, but its published version is pinned into later Benchmark Versions so historical evaluation evidence remains interpretable. The settings surface shows the Project's locked framework, its decision rule, what the framework is suited for, and which scoring controls administrators can configure. A framework is not a Policy or Rubric; it determines how those governed artifacts are interpreted and aggregated for evaluation. ## Fields, states, or lifecycle rules - The framework is locked for an existing Project. Other framework choices can be available when creating a new Project. - Depending on the framework contract, administrators can configure Case aggregation, score range, penalty policy, Rubric role assignment, safe custom formulas, and an Importance scale. - A custom formula is validated and stored as typed configuration; it is not executable code. - Editing creates unpublished scoring changes. **Publish new Version** creates an immutable Regime Version after validation and concurrency checks. - New drafts and finalized Benchmark Versions retain the Regime Version they were created or finalized with. Publishing a later version does not rewrite existing plans, Runs, or results. - Published version history shows prior configurations and can show which open drafts or Benchmark Versions use a version. - A concurrent publication requires the administrator to review the latest version before publishing retained edits. ## Configure safely Review the framework decision rule before changing scoring behavior. Confirm that the desired aggregation, score range, penalties, and Importance semantics match the Policies and Rubrics that will be evaluated. Publish only after the resulting version can be explained to someone reading a future Benchmark result. Do not use Regime Settings to change the meaning of a Policy or Rubric. Change those governed artifacts through Correctness Governance, then create or evaluate the appropriate versioned benchmark evidence. {% example-demo title="Example: preserving historical scoring" %} An administrator changes the score range and publishes Regime Version 4. New Benchmark Versions pin Version 4, while an existing finalized Benchmark Version continues to report under Version 3. The administrator can inspect both configurations and their usage without treating the new scoring choice as a rewrite of old results. {% /example-demo %} ## Source confidence Code-backed: the Project Settings Regime panel exposes locked framework guidance, scoring configuration, validation, immutable publication, concurrent-publication recovery, and version history. The project API exposes framework discovery, current Regime reads, publication, history, and version usage. ## Related task pages {% related-card-grid title="Related task pages" %} - [Configure Project Settings](/docs/project-settings) - [Build Policies and Rubrics](/docs/operating-manual/build-policies-and-rubrics) - [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation) {% /related-card-grid %} --- id: project-settings.overview title: Project Settings summary: Configure the project identity, evaluation Regime, case input contract, and project members. kind: concept product_area: project_settings status: stable updated: 2026-09-07 canonical: /docs/project-settings --- # Project Settings Project Settings is a floating project-level surface rather than a benchmark workspace. Its current tabs are **General**, **Regime**, **Input Schema**, and **Project Members**. | Setting | Governs | Does not replace | | --- | --- | --- | | General | Project name and Project Memo | Agent Setup context or instructions | | Regime | How approved and applicable Rubrics contribute to future Benchmark Versions | Policies, Rubrics, or already-published Benchmark evidence | | Input Schema | Canonical case input and material contract | A benchmark dataset snapshot | | Project Members | User and group access to this project | Contribution task assignment | Changes are project-scoped and may affect future work across multiple benchmarks. Treat input-contract and access changes as governance decisions, and preserve exact versions and snapshots wherever historical evidence depends on them. ## Change boundaries Project Settings is intentionally separate from Agent Setup and benchmark workspaces. A settings change can influence what future work accepts or displays, but it does not silently rewrite a saved Harness Version, Dataset Snapshot, Contribution, or Run. When a project-wide contract changes, inspect downstream readiness and create new versioned evidence where the product workflow requires it. Input Schema deserves the most caution because future Case validation follows it. Before tightening a required material or changing architecture, identify existing Cases that may no longer conform. Regime changes create a new immutable Regime Version for future Benchmark Versions; existing Benchmark Versions, Runs, and results retain their published Regime Version. Project Members affects access, not authorship or task history. ## Operating sequence 1. Set a clear project name and human-facing memo. 2. Review the locked Regime and publish a new version only when the scoring contract should change for future Benchmark Versions. 3. Define the Input Schema before importing or constructing substantial Case evidence. 4. Grant users and groups the project access needed for their role. 5. Revisit settings when the project contract changes, then check Assets, benchmarks, and active Contributions for downstream impact. {% example-demo title="Example: adding a required document" %} A project decides every future Case must include a controlling policy document. The operator updates Input Schema only after auditing current Cases. Existing Dataset Snapshots remain historical evidence; corrected live Cases enter a new Snapshot. The Project Memo may explain the ownership decision, but it does not enforce the material requirement. {% /example-demo %} ## Relationship to governance Workspace administration controls the wider account boundary. Correctness Governance owns Policies and Rubrics. Project Settings should therefore express project contracts and access, not become a catch-all place for evaluator rules, secret values, or informal candidate configuration. {% related-card-grid title="Project settings" %} - [General settings](/docs/project-settings/general) - [Regime settings](/docs/project-settings/regime) - [Project Input Schema](/docs/project-settings/input-schema) - [Project Members](/docs/project-settings/project-members) {% /related-card-grid %} {% related-card-grid title="Connected workspaces" %} - [Agent Setup](/docs/agent-setup) - [Assets](/docs/assets) - [Benchmark Datasets](/docs/benchmark-datasets) {% /related-card-grid %} ## Source confidence Code-backed: the active Project Settings surface defines the General, Regime, Input Schema, and Project Members tabs. Reference Materials is documented under Agent Setup because it supplies governed project knowledge rather than these four settings contracts. --- id: correctness.policies-rubrics title: Policies and Rubrics summary: Understand the governed relationship between behavior policies, applicability, binary rubrics, linked cases, and expert provenance. kind: reference product_area: correctness_governance status: stable updated: 2026-08-23 canonical: /docs/correctness-governance/policies-and-rubrics --- # Policies and Rubrics ## Definition A **Policy** is a reusable statement of expected specialist AI behavior. Its applicability explains the situations in which the rule controls. A **Rubric** is an evaluation criterion that turns the policy into observable evidence for a case and candidate response. Correctness Governance owns both artifact types. Expert Contributions can supply proposed or accepted policy and rubric material, while the governance surfaces preserve the artifact's current state, links, activity, and provenance. ## Fields, states, or lifecycle rules - Policies have identity, descriptive rule content, scope or applicability, linked cases, linked rubrics, activity, and approval context. - Rubrics have identity, criterion wording, policy or case relationships, evaluation relevance, and lifecycle context. - A policy can connect to several rubrics when its behavior requirements need separate checks. - A rubric should express one inspectable criterion wherever independent diagnosis matters. - Linked cases demonstrate applicability or behavior; benchmark dataset membership remains a separate benchmark-scoped decision. - Proposed applications and agent suggestions remain proposals until the owning workflow records acceptance. - Expert contribution provenance should remain visible when contributed material becomes a governed artifact. - Editing a project-level standard does not retroactively change the standard boundary used by an already recorded Run. ![Correctness Governance rows showing Policy titles, lifecycle state, required behavior force, linked-count columns, and accountable owners.](/docs-assets/assets/screenshots/policies-rubrics-neutral-rows.png) Read the rule, state, links, and owner together; a plausible title alone does not establish governed authority. ## Reading the pair Begin with the policy when deciding what should happen and why. Inspect applicability before assuming the policy governs a case. Then read the linked rubric as the testable question applied to candidate behavior. If the rubric cannot be answered from the response and visible case evidence, revise the criterion or the case rather than relying on reviewer intuition. When standards overlap, distinguish complementary criteria from contradictory authority. Preserve unresolved conflict until an accountable expert contribution or governance action settles the intended rule. {% example-demo title="Example: escalation policy and rubrics" %} A policy states that unresolved eligibility exceptions must be escalated. One rubric checks that the response does not promise the exception; another checks that it gives the correct escalation path. Separating the checks lets an evaluation show whether a candidate avoided the unsupported promise but still failed to guide the user correctly. {% /example-demo %} ## Source confidence Code-backed: active policy and rubric detail routes expose linked cases, linked rubrics, approval and activity context, and evaluation relationships. Exact editable fields can vary by artifact state. ## Related task pages {% related-card-grid title="Related task pages" %} - [Build policies and rubrics](/docs/operating-manual/build-policies-and-rubrics) - [Write binary rubrics](/docs/correctness-governance/binary-rubrics) - [Request an Expert Contribution](/docs/expert-contributions/request-contribution) {% /related-card-grid %} --- id: benchmark-evaluations.overview title: Benchmark Evaluations summary: Run and inspect exact Harness Versions against an immutable Benchmark Version through Dashboard, List, Arena, and Compare. kind: concept product_area: benchmark_evaluations status: stable updated: 2026-09-13 canonical: /docs/benchmark-evaluations --- # Benchmark Evaluations Benchmark Evaluations is the version-scoped workspace for executing and comparing candidate systems. The active top-level tabs are **Dashboard**, **List**, **Arena**, and **Compare**. Every managed Run binds an exact saved Harness Version to the immutable Benchmark Version shown in the route. > Evaluation boundary > > Interpret evidence inside its recorded Benchmark Version, Harness Version, Run or Run Group, evaluator set, and metadata. Run counts belong to launches. Additional launches add evidence without rewriting earlier Runs. ## Surfaces and objects Dashboard summarizes progress, leaderboards, rank progression across Runs, and available resource telemetry. List is segmented into **Runs**, **Evaluation results**, and **Traces / Spans**. The results segment summarizes Case outcomes and Policy or Rubric failures. Arena compares candidate pairs across governed metrics. Compare is a symmetric matrix of Harness Versions across selected evidence rows. A Run Group can collect one standard attempt or repeated attempts. A Run records one candidate execution and its per-Case progress. Evaluation results record the admitted Policy and Rubric outcomes. Costs, tokens, and latency are telemetry only when the provider or execution path captured them. > Traces / Spans capability fence > > The List navigation exposes Traces / Spans, but the current benchmark API does not expose evaluation execution traces. Do not claim that trajectories, spans, private reasoning, or tool traces can be inspected from Benchmark Evaluations today. ## Decision checkpoint | Need | Open | Evidence to preserve | | --- | --- | --- | | Configure and launch managed Runs | Evaluation Settings and New evaluation run | Machine, saved Harness Versions, and per-Harness Run counts | | Start candidate execution | Run modal | Exact Harness and Benchmark Versions | | Inspect status and output summaries | List → Runs or Evaluation results | Run Group, attempt, Case counts, incomplete state | | Compare candidate pairs | Arena | Metric family, pair count, only-A, only-B, shared failures | | Compare many candidates by governed rows | Compare | Harness columns and chosen Case or facet row mode | | Admit external reference outputs | Output mapping | Case mapping, attempt assignment, insert/update report | ## Rankings and repeated sampling Dashboard aggregates compatible observed Runs for each saved Harness Version across launches. Average score weights Runs equally. Supported binary views report passed at least once or passed every time over the observed case outcomes. Counts and missing evidence are shown; unequal counts do not prevent comparison. Historical group metrics retain their recorded meanings. Ranking is a routing signal. A candidate can lead overall while failing required Policy or high-impact Rubric evidence. Use Arena or Compare to locate the disagreement and List to confirm completeness before starting Improve work. ## External outputs Uploaded or API-supplied reference outputs create output-only Runs that can be scored and inspected in List. They are not saved Harness Versions and therefore cannot be optimized in Improve or selected as Harness columns in Compare or Arena. {% example-demo title="Example: repeated evaluation without evidence drift" %} A team launches three Runs of Harness Version 8 and one Run of Version 11 against the same Benchmark Version. Both appear with their evidence counts. A later launch of Version 11 adds two Runs to its aggregate evidence without changing either launch group. The team can inspect individual Runs before deciding whether more evidence is useful. {% /example-demo %} ## Related workflows {% related-card-grid title="Related workflows" %} - [Configure evaluation execution](/docs/benchmark-evaluations/execution-settings) - [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation) - [Inspect evaluation results](/docs/benchmark-evaluations/inspect-results) - [Use Arena and rankings](/docs/benchmark-evaluations/arena-and-rankings) - [Compare Harness Versions](/docs/benchmark-evaluations/compare) - [Map external outputs](/docs/benchmark-evaluations/output-mapping) {% /related-card-grid %} ## Source confidence Code-backed: the active version-scoped workspace, settings, Run modal, List segments, Dashboard, Arena, and Compare routes define the current evaluation model and capability fences. --- id: governance.benchmark-versioning title: Benchmark Versioning summary: Preserve benchmark snapshots so evidence can be compared across target and standard changes. kind: reference product_area: governance status: stable updated: 2026-08-23 canonical: /docs/governance/benchmark-versioning --- # Benchmark Versioning ## Definition A Benchmark Version is the immutable evidence boundary used by Runs. It identifies the frozen dataset state and admitted evaluator relationships that make a result interpretable. The Benchmark remains a durable program; its versions preserve successive evidence boundaries as Cases, materials, coverage, Policies, or Rubrics change. ## Fields, states, or lifecycle rules - Create a new Snapshot and resulting Benchmark Version when changed evidence would alter what a Run claims to test. - Existing Runs remain attached to their original Benchmark Version. - Current Dataset edits do not mutate a historical version. - A new Harness Version alone does not require a new Benchmark Version; candidate and evidence versions move independently. - Comparisons within one Benchmark Version isolate candidate differences more cleanly. - Cross-version comparisons must name the changed Cases, evaluators, or representation boundary. - Version identity does not prove that coverage is sufficient or that every admitted Rubric is correct. ## Related objects Use [Dataset Snapshots](/docs/benchmark-datasets/snapshots) to create the frozen dataset boundary. Use [Benchmark Evaluations](/docs/benchmark-evaluations) to inspect Runs inside one exact version, and [Compare Harness Versions](/docs/benchmark-evaluations/compare) to interpret candidate movement without hiding version changes. {% example-demo title="Separating candidate change from standard change" %} Harness Version 12 improves retrieval and is evaluated against Benchmark Version 4, the same boundary used for Version 11. That comparison isolates candidate behavior. Later, experts approve a stricter source-authority Rubric and the dataset gains conflict Cases. The team creates Benchmark Version 5 and reports subsequent Runs under that new boundary instead of presenting the lower score as a regression against unchanged evidence. {% /example-demo %} ## Source confidence Code-backed: Benchmark Datasets → Snapshots preserves immutable Dataset boundaries, and the version-scoped Evaluation route binds Runs to one selected Benchmark Version. Coverage quality and downstream decisions remain outside version identity itself. ## Related task pages {% related-card-grid title="Related task pages" %} - [Benchmark snapshots](/docs/coverage-engineering/benchmark-snapshots) - [Benchmarks](/docs/object-model/benchmarks) - [Compare Harness Versions](/docs/benchmark-evaluations/compare) - [Product quickstart](/docs/quickstart) - [Task index](/docs/operating-manual/task-index) {% /related-card-grid %}