# Policies and Rubrics Generated: 2026-09-13T04:43:22.040Z Source build: local Canonical docs: https://teammately.ai/docs --- id: correctness.policies-rubrics title: Policies and Rubrics summary: Understand the governed relationship between behavior policies, applicability, binary rubrics, linked cases, and expert provenance. kind: reference product_area: correctness_governance status: stable updated: 2026-08-23 canonical: /docs/correctness-governance/policies-and-rubrics --- # Policies and Rubrics ## Definition A **Policy** is a reusable statement of expected specialist AI behavior. Its applicability explains the situations in which the rule controls. A **Rubric** is an evaluation criterion that turns the policy into observable evidence for a case and candidate response. Correctness Governance owns both artifact types. Expert Contributions can supply proposed or accepted policy and rubric material, while the governance surfaces preserve the artifact's current state, links, activity, and provenance. ## Fields, states, or lifecycle rules - Policies have identity, descriptive rule content, scope or applicability, linked cases, linked rubrics, activity, and approval context. - Rubrics have identity, criterion wording, policy or case relationships, evaluation relevance, and lifecycle context. - A policy can connect to several rubrics when its behavior requirements need separate checks. - A rubric should express one inspectable criterion wherever independent diagnosis matters. - Linked cases demonstrate applicability or behavior; benchmark dataset membership remains a separate benchmark-scoped decision. - Proposed applications and agent suggestions remain proposals until the owning workflow records acceptance. - Expert contribution provenance should remain visible when contributed material becomes a governed artifact. - Editing a project-level standard does not retroactively change the standard boundary used by an already recorded Run. ![Correctness Governance rows showing Policy titles, lifecycle state, required behavior force, linked-count columns, and accountable owners.](/docs-assets/assets/screenshots/policies-rubrics-neutral-rows.png) Read the rule, state, links, and owner together; a plausible title alone does not establish governed authority. ## Reading the pair Begin with the policy when deciding what should happen and why. Inspect applicability before assuming the policy governs a case. Then read the linked rubric as the testable question applied to candidate behavior. If the rubric cannot be answered from the response and visible case evidence, revise the criterion or the case rather than relying on reviewer intuition. When standards overlap, distinguish complementary criteria from contradictory authority. Preserve unresolved conflict until an accountable expert contribution or governance action settles the intended rule. {% example-demo title="Example: escalation policy and rubrics" %} A policy states that unresolved eligibility exceptions must be escalated. One rubric checks that the response does not promise the exception; another checks that it gives the correct escalation path. Separating the checks lets an evaluation show whether a candidate avoided the unsupported promise but still failed to guide the user correctly. {% /example-demo %} ## Source confidence Code-backed: active policy and rubric detail routes expose linked cases, linked rubrics, approval and activity context, and evaluation relationships. Exact editable fields can vary by artifact state. ## Related task pages {% related-card-grid title="Related task pages" %} - [Build policies and rubrics](/docs/operating-manual/build-policies-and-rubrics) - [Write binary rubrics](/docs/correctness-governance/binary-rubrics) - [Request an Expert Contribution](/docs/expert-contributions/request-contribution) {% /related-card-grid %} --- id: correctness.overview title: Correctness Governance summary: Govern project policies and rubrics, their applicability, linked cases, approval state, and contribution provenance. kind: concept product_area: correctness_governance status: stable updated: 2026-09-07 canonical: /docs/correctness-governance --- # Correctness Governance Correctness Governance is the project-level surface for policies and rubrics. It makes the specialist standards used by expert work and benchmark evaluation inspectable, attributable, and reusable across benchmark workspaces. ## Definition The surface has **Policies** and **Rubrics**. Policy pages describe a behavior rule, its scope, linked examples, approval state, activity, and connected rubrics. Rubric pages define testable evaluation criteria and show how they relate to cases, policies, and evaluation use. Correctness Elicitation is the broader capability that discovers and resolves specialist judgment. Correctness Governance is the product surface that owns the resulting governed project artifacts. Expert Contributions may propose or contribute policies and rubrics, but those artifacts retain their own lifecycle and provenance. ## Decision checkpoint | Need | Inspect or change | Boundary | | --- | --- | --- | | State a reusable behavior rule | Policy | Keep source and expert rationale visible | | Decide where a rule applies | Policy scope and applicability | Do not encode broad intent only in rubric wording | | Make the rule testable | Binary rubric | Define observable pass and fail evidence | | Connect standards to examples | Linked cases | A linked case does not automatically belong to every benchmark | | Review expert-originated material | Contribution provenance and approval | Contribution output is not silently governed | | Explain evaluation movement | Policy, rubric, case, and version context | Do not rely on an aggregate score alone | ## Policies and rubrics A policy explains what behavior should occur and why. Applicability determines the situations in which the rule controls. A rubric turns that rule into an evaluation question whose result can be traced to observable behavior. Several rubrics may operationalize different parts of one policy, and cases can help demonstrate where each rubric applies. Rubrics should remain atomic enough to interpret. If one rubric simultaneously checks grounding, tone, escalation, and completeness, a failure does not identify the responsible behavior. Split criteria where independent failure evidence matters. ## Human and agent roles Agents can draft possible wording, surface linked cases, identify proposed applications, and prepare follow-up questions. Experts contribute domain authority through benchmark-scoped work. Project operators inspect and maintain the governed artifacts. Approval history and activity should make the transition between proposal, contribution, and governed state visible. Changing a project-level policy or rubric may affect several benchmarks. Before interpreting a later Run, confirm which benchmark version and standard boundary it used. {% example-demo title="Unsupported compatibility" %} A specialist confirms that compatibility may be claimed only when an authoritative source explicitly supports the exact equipment combination. Correctness Governance records the policy, scopes it to recommendation and validation responses, links representative cases, and defines a binary rubric that passes only when the response cites support or clearly states uncertainty. {% /example-demo %} ## Related workflows {% related-card-grid title="Related workflows" %} - [Build policies and rubrics](/docs/operating-manual/build-policies-and-rubrics) - [Write binary rubrics](/docs/correctness-governance/binary-rubrics) - [Request an Expert Contribution](/docs/expert-contributions/request-contribution) {% /related-card-grid %} ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Policies](/docs/object-model/policies) - [Applicability Logic](/docs/object-model/applicability-logic) - [Rubrics](/docs/object-model/rubrics) {% /related-card-grid %} ## Source confidence Code-backed: the active Correctness Governance layout, policy list, rubric list, and detail surfaces establish the current ownership and relationships described here. --- id: correctness.binary-rubrics title: Write Binary Rubrics summary: Write atomic pass-or-fail criteria grounded in governed policies, applicable cases, and observable candidate behavior. kind: task product_area: correctness_governance status: stable updated: 2026-08-22 canonical: /docs/correctness-governance/binary-rubrics --- # Write Binary Rubrics Write a rubric when a governed policy needs an observable pass-or-fail check for benchmark evaluation. A strong rubric identifies one behavior, the cases where it applies, and the evidence that distinguishes pass from fail. ## Prerequisites - A policy or expert-grounded correctness statement. - Representative passing, failing, and boundary cases. - Clear applicability for the behavior being checked. - Access to Correctness Governance → Rubrics. ## Steps 1. State one behavior that can be inspected in the candidate response and visible case evidence. 2. Name the policy or specialist judgment that authorizes the criterion. 3. Define applicability before writing exceptions into the pass condition. 4. Write explicit pass evidence and fail evidence. Avoid “good,” “appropriate,” or “high quality” without observable conditions. 5. Link representative cases and test whether two informed reviewers would reach the same binary result. 6. Split independent requirements into separate rubrics when each failure should be diagnosed separately. 7. Inspect contribution provenance and approval state before relying on the rubric in benchmark interpretation. ## Object and state changes This task creates or updates a project-level rubric and can change its wording, policy relationship, linked cases, evaluation use, activity, and approval context. Linking a case does not add it to a benchmark dataset. Editing a rubric does not alter historical Run evidence that used an earlier benchmark boundary. ## Success criteria - The rubric tests one behavior and can be answered from visible evidence. - Applicability excludes irrelevant cases without hidden reviewer judgment. - Pass and fail conditions are explicit. - Linked cases include at least one meaningful boundary. - Policy authority and expert provenance are inspectable. ## Common failure modes - Combining several behaviors into one criterion. - Restating the policy without defining observable evidence. - Encoding applicability only as exceptions inside the rubric. - Using a suggested or contributed draft as if it were already governed. - Changing rubric wording and comparing Runs without checking the benchmark version boundary. {% example-demo title="Example: grounding rubric" %} Policy: material claims must use the controlling source or state uncertainty. Rubric: pass only when every material claim is supported by the current controlling source, or the response explicitly says the available sources do not resolve the claim. Unsupported blending of current and superseded sources fails. {% /example-demo %} ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Policies and Rubrics](/docs/correctness-governance/policies-and-rubrics) - [Applicability Logic](/docs/object-model/applicability-logic) - [Rubrics](/docs/object-model/rubrics) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Overlapping rubrics](/docs/troubleshooting/overlapping-rubrics) - [Weak applicability logic](/docs/troubleshooting/weak-applicability-logic) - [Low expert agreement](/docs/troubleshooting/low-expert-agreement) {% /related-card-grid %} ## Source confidence Code-backed: the active Correctness Governance rubric list and detail surfaces support rubric inspection, relationships, and lifecycle context. The drafting guidance is constrained to those verified artifact boundaries. --- id: expert-contributions.artifacts title: Contributed Artifacts summary: Inspect policies, rubrics, cases, and coverage observations produced through attributable expert contribution work. kind: reference product_area: expert_contributions status: stable updated: 2026-09-07 canonical: /docs/expert-contributions/contributed-artifacts --- # Contributed Artifacts ## Definition Contributed Artifacts is the benchmark workspace for inspecting durable material produced through Expert Contributions. It organizes contributed **Policies**, **Rubrics**, **Cases**, and **new coverage observations** while preserving their relationship to the Contribution and expert work that produced them. The view is a provenance and reconciliation surface. The final owner of a materialized artifact remains Correctness Governance, Assets, or Coverage Management according to artifact type. ## Fields, states, or lifecycle rules - Policy contributions represent expert-grounded behavior rules or revisions. - Rubric contributions represent proposed or accepted evaluation criteria tied to specialist judgment. - Case contributions represent situations supplied or corrected through expert work. - Coverage observations identify missing, thin, conflicting, or newly important benchmark behavior. - A coverage observation preserves its source, proposed facet applications, and application status so an operator can distinguish a recorded observation from one incorporated into coverage structure. - Each artifact should remain traceable to the Contribution, expert, selected evidence, task responses, and checkpoints that support it. - Contribution completion and artifact governance are separate transitions. Inspect the artifact's owning surface before treating it as active policy, active rubric, benchmark dataset membership, or resolved coverage. - Reconciliation can accept, revise, route, or leave material unresolved according to the active workflow. ## Interpreting contributed material Use the artifact type to choose the next surface. A contributed policy or rubric belongs in Correctness Governance. A contributed case belongs in the project Assets pool before benchmark selection. A coverage observation belongs in Coverage Management and may motivate a Coverage Story, case construction, or another focused Contribution. Preserve disagreements. Two experts can contribute conflicting Policy interpretations, and the artifact view should help an operator trace each interpretation rather than merge them into an invented consensus. Materialization should keep the Contribution, activity, checkpoint, expert, and scoped evidence links needed to explain why the artifact exists. {% example-demo title="Example: contribution provenance" %} An expert contributes a policy limiting compatibility claims, a rubric for explicit uncertainty, and a new case involving an unsupported adapter. The policy and rubric move to Correctness Governance for their lifecycle. The case enters Assets and is later selected into a benchmark dataset snapshot. All three retain the Contribution as their provenance. {% /example-demo %} ## Source confidence Code-backed: the active Contributed Artifacts workspace exposes policy, rubric, case, and new-coverage groupings. This page preserves the separation between contribution provenance and the lifecycle of each owning artifact. ## Related task pages {% related-card-grid title="Related task pages" %} - [Request an Expert Contribution](/docs/expert-contributions/request-contribution) - [Complete an Expert Contribution](/docs/expert-contributions/complete-contribution) - [Build policies and rubrics](/docs/operating-manual/build-policies-and-rubrics) {% /related-card-grid %}