# Policies and rubrics
Generated: 2026-09-13T04:43:12.640Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: concepts.policies-rubrics
title: Policies and rubrics
summary: Learn how Teammately turns product judgment into reusable policies and scoring rubrics.
kind: concept
product_area: object_model
status: stable
updated: 2026-08-23
canonical: /docs/concepts/policies-and-rubrics
---
# Policies and rubrics
## Definition
Policies state governed expectations for behavior and the situations to which those expectations apply. Rubrics are evaluator definitions used to judge Case responses. A Policy can link relevant Cases and Rubrics, but the objects remain separately versioned and reviewable.
## Why it matters
This separation lets a team correct the right layer. A mistaken rule belongs in the Policy; an overbroad scope belongs in applicability; an unreliable check belongs in the Rubric. Evaluation evidence should show which applicable Rubric produced each outcome rather than treating an aggregate score as the standard itself.
## Standard pair check
| The pair is healthy when... | Rework it when... |
| --- | --- |
| The policy states the product behavior rule. | The policy is only tone, preference, or broad quality advice. |
| Applicability names the cases where the rule belongs. | The same rubric could apply to nearly everything. |
| The Rubric tests one observable requirement. | The Rubric combines several decisions into one unclear result. |
## Where it appears in the product
Create and inspect Policies and Rubrics under **Correctness Governance**. Expert Contributions can supply attributable candidate artifacts, but contributed content is not automatically approved. Benchmark Evaluations reports applicable Rubric outcomes for the frozen Benchmark Version.
## Artifacts it affects
Policies and Rubrics affect Case links, applicability, Benchmark Versions, evaluator coverage, Run results, failure clusters, and staleness. Changing either governed object requires a new version boundary before the revised standard is treated as current evaluation evidence.
{% example-demo title="Eligibility policy to must-level rubric" %}
A support Policy requires entitlement answers to use the controlling contract or state uncertainty. Its applicability is limited to plan limits, contract exceptions, and admin-controlled access. A linked binary Rubric checks whether the response identifies that source or explicitly withholds an unsupported eligibility claim. Evaluation results can then show the failed Rubric on the affected Cases without broadening the rule to unrelated setup questions.
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Work with Policies and Rubrics](/docs/correctness-governance/policies-and-rubrics)
- [Design binary Rubrics](/docs/correctness-governance/binary-rubrics)
- [Request an Expert Contribution](/docs/expert-contributions/request-contribution)
- [Inspect evaluation results](/docs/benchmark-evaluations/inspect-results)
{% /related-card-grid %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Applicability logic](/docs/object-model/applicability-logic)
- [Cases](/docs/object-model/cases)
- [Policies](/docs/object-model/policies)
- [Rubrics](/docs/object-model/rubrics)
- [Human Approval Boundaries](/docs/governance/human-approval-boundaries)
{% /related-card-grid %}
## Source confidence
Code-backed: the current Policy and Rubric list and detail routes define their separate identities, editable fields, links, versions, and approval state. Expert Contribution and Benchmark Evaluation pages define how attributable input and evaluator outcomes enter those objects' wider lifecycle.
---
id: correctness.policies-rubrics
title: Policies and Rubrics
summary: Understand the governed relationship between behavior policies, applicability, binary rubrics, linked cases, and expert provenance.
kind: reference
product_area: correctness_governance
status: stable
updated: 2026-08-23
canonical: /docs/correctness-governance/policies-and-rubrics
---
# Policies and Rubrics
## Definition
A **Policy** is a reusable statement of expected specialist AI behavior. Its applicability explains the situations in which the rule controls. A **Rubric** is an evaluation criterion that turns the policy into observable evidence for a case and candidate response.
Correctness Governance owns both artifact types. Expert Contributions can supply proposed or accepted policy and rubric material, while the governance surfaces preserve the artifact's current state, links, activity, and provenance.
## Fields, states, or lifecycle rules
- Policies have identity, descriptive rule content, scope or applicability, linked cases, linked rubrics, activity, and approval context.
- Rubrics have identity, criterion wording, policy or case relationships, evaluation relevance, and lifecycle context.
- A policy can connect to several rubrics when its behavior requirements need separate checks.
- A rubric should express one inspectable criterion wherever independent diagnosis matters.
- Linked cases demonstrate applicability or behavior; benchmark dataset membership remains a separate benchmark-scoped decision.
- Proposed applications and agent suggestions remain proposals until the owning workflow records acceptance.
- Expert contribution provenance should remain visible when contributed material becomes a governed artifact.
- Editing a project-level standard does not retroactively change the standard boundary used by an already recorded Run.

Read the rule, state, links, and owner together; a plausible title alone does not establish governed authority.
## Reading the pair
Begin with the policy when deciding what should happen and why. Inspect applicability before assuming the policy governs a case. Then read the linked rubric as the testable question applied to candidate behavior. If the rubric cannot be answered from the response and visible case evidence, revise the criterion or the case rather than relying on reviewer intuition.
When standards overlap, distinguish complementary criteria from contradictory authority. Preserve unresolved conflict until an accountable expert contribution or governance action settles the intended rule.
{% example-demo title="Example: escalation policy and rubrics" %}
A policy states that unresolved eligibility exceptions must be escalated. One rubric checks that the response does not promise the exception; another checks that it gives the correct escalation path. Separating the checks lets an evaluation show whether a candidate avoided the unsupported promise but still failed to guide the user correctly.
{% /example-demo %}
## Source confidence
Code-backed: active policy and rubric detail routes expose linked cases, linked rubrics, approval and activity context, and evaluation relationships. Exact editable fields can vary by artifact state.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Build policies and rubrics](/docs/operating-manual/build-policies-and-rubrics)
- [Write binary rubrics](/docs/correctness-governance/binary-rubrics)
- [Request an Expert Contribution](/docs/expert-contributions/request-contribution)
{% /related-card-grid %}
---
id: correctness.binary-rubrics
title: Write Binary Rubrics
summary: Write atomic pass-or-fail criteria grounded in governed policies, applicable cases, and observable candidate behavior.
kind: task
product_area: correctness_governance
status: stable
updated: 2026-08-22
canonical: /docs/correctness-governance/binary-rubrics
---
# Write Binary Rubrics
Write a rubric when a governed policy needs an observable pass-or-fail check for benchmark evaluation. A strong rubric identifies one behavior, the cases where it applies, and the evidence that distinguishes pass from fail.
## Prerequisites
- A policy or expert-grounded correctness statement.
- Representative passing, failing, and boundary cases.
- Clear applicability for the behavior being checked.
- Access to Correctness Governance → Rubrics.
## Steps
1. State one behavior that can be inspected in the candidate response and visible case evidence.
2. Name the policy or specialist judgment that authorizes the criterion.
3. Define applicability before writing exceptions into the pass condition.
4. Write explicit pass evidence and fail evidence. Avoid “good,” “appropriate,” or “high quality” without observable conditions.
5. Link representative cases and test whether two informed reviewers would reach the same binary result.
6. Split independent requirements into separate rubrics when each failure should be diagnosed separately.
7. Inspect contribution provenance and approval state before relying on the rubric in benchmark interpretation.
## Object and state changes
This task creates or updates a project-level rubric and can change its wording, policy relationship, linked cases, evaluation use, activity, and approval context. Linking a case does not add it to a benchmark dataset. Editing a rubric does not alter historical Run evidence that used an earlier benchmark boundary.
## Success criteria
- The rubric tests one behavior and can be answered from visible evidence.
- Applicability excludes irrelevant cases without hidden reviewer judgment.
- Pass and fail conditions are explicit.
- Linked cases include at least one meaningful boundary.
- Policy authority and expert provenance are inspectable.
## Common failure modes
- Combining several behaviors into one criterion.
- Restating the policy without defining observable evidence.
- Encoding applicability only as exceptions inside the rubric.
- Using a suggested or contributed draft as if it were already governed.
- Changing rubric wording and comparing Runs without checking the benchmark version boundary.
{% example-demo title="Example: grounding rubric" %}
Policy: material claims must use the controlling source or state uncertainty. Rubric: pass only when every material claim is supported by the current controlling source, or the response explicitly says the available sources do not resolve the claim. Unsupported blending of current and superseded sources fails.
{% /example-demo %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Policies and Rubrics](/docs/correctness-governance/policies-and-rubrics)
- [Applicability Logic](/docs/object-model/applicability-logic)
- [Rubrics](/docs/object-model/rubrics)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Overlapping rubrics](/docs/troubleshooting/overlapping-rubrics)
- [Weak applicability logic](/docs/troubleshooting/weak-applicability-logic)
- [Low expert agreement](/docs/troubleshooting/low-expert-agreement)
{% /related-card-grid %}
## Source confidence
Code-backed: the active Correctness Governance rubric list and detail surfaces support rubric inspection, relationships, and lifecycle context. The drafting guidance is constrained to those verified artifact boundaries.
---
id: object-model.applicability-logic
title: Applicability logic
summary: Explain when a policy or rubric should be used for a case, output, or coverage segment.
kind: reference
product_area: object_model
status: stable
updated: 2026-08-23
canonical: /docs/object-model/applicability-logic
---
# Applicability logic
## Definition
Applicability logic is the boundary that decides whether a policy or rubric should be considered for a specific case, output, or coverage segment. It prevents a good standard from being applied to the wrong behavior.
Use this reference when a benchmark result is hard to explain because a standard seems relevant in some cases but not others. The question is not only whether the policy is correct; it is whether the policy was eligible to judge that output.
## Fields, states, or lifecycle rules
- Applicability sits between the case/output and the policy/rubric that may judge it.
- Weak applicability makes benchmark failures noisy: a candidate can fail a good rubric on a case where the rule should not have applied.
- Strong applicability names the behavior condition, source context, or case segment that brings the standard into scope.
- Applicability changes can make old benchmark evidence stale because the same output may be judged by a different standard boundary.
- This page explains the public object relationship, not a public rule language, API schema, or export contract.
## Related objects
Applicability logic should be read with [Policies](/docs/object-model/policies), [Rubrics](/docs/object-model/rubrics), [Cases](/docs/object-model/cases), and [Coverage Dimensions](/docs/object-model/coverage-dimensions). Use the Correctness Governance workflow to write the boundary and troubleshooting when the observed Case set is wrong.
{% example-demo title="Applicability logic boundary" %}
Raw case: A user asks whether a product works with equipment they already own.
Policy: Compatibility claims require explicit source support.
Applicability logic: The policy applies only when the answer recommends, validates, or compares a product for a concrete use context.
Benchmark interpretation: If the case only asks for a product description, the compatibility rubric should not judge it. If the answer claims the product will work with the user's equipment, the rule is in scope.
{% /example-demo %}
## Source confidence
Code-backed: Policy and Rubric types and their detail routes expose applicability fields and the links among governed standards, Cases, and evaluation checks. This page describes those product relationships; it does not define a portable rule language.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Work with Policies and Rubrics](/docs/correctness-governance/policies-and-rubrics)
- [Fix weak applicability logic](/docs/troubleshooting/weak-applicability-logic)
- [Resolve conflicting correctness evidence](/docs/governance/conflict-resolution)
- [Product quickstart](/docs/quickstart)
- [Task index](/docs/operating-manual/task-index)
{% /related-card-grid %}
---
id: benchmark-evaluations.inspect
title: Inspect Evaluation Results
summary: Trace Dashboard and List signals to Run, Case, Policy, Rubric, completeness, and telemetry evidence.
kind: task
product_area: benchmark_evaluations
status: stable
updated: 2026-09-13
canonical: /docs/benchmark-evaluations/inspect-results
---
# Inspect Evaluation Results
## Prerequisites
- A visible Run or Run Group with output or evaluation progress.
- Access to the exact Benchmark and Harness Version evidence.
Start with completeness and identity, then move from aggregate signals to the Cases and evaluator failures that support them.
## Steps
1. Open Dashboard and confirm the Benchmark Version, candidate Harness Version, Run Group type, attempt count, and evaluation progress.
2. Read rankings with their metric family and uncertainty. Distinguish average score, passed at least once, and passed every time over observed Runs. Inspect Run counts and coverage; historical group-specific pass@n and pass^n retain their original meanings.
3. Open **List → Runs** to inspect group and Run status, output progress, evaluation progress, metadata, and available resource telemetry.
4. Open **List → Evaluation results** for the Case summary, outcome, failed Policies, failed Rubrics, and evaluated count.
5. Use Arena for pairwise disagreement or Compare for a Harness matrix across Cases, evaluator facts, or Coverage Facets.
6. Classify the next action as candidate work, evaluator clarification, Case correction, coverage work, external-output remapping, or no action.
The List results surface is intentionally compact. Do not claim that it exposes full execution trajectories. The **Traces / Spans** segment currently reports a capability fence because the benchmark API does not provide evaluation execution traces.
## Reading incomplete and repeated evidence
An aggregate calculated over fewer evaluable Cases can look better while covering less evidence. Record evaluated, incomplete, and missing counts before comparing candidates. For repeated groups, inspect whether the configured number of attempts exists for every candidate and whether one failed attempt changes the metric interpretation.
Cost, tokens, and latency help route operational work but are nullable telemetry. Missing capture means unknown, not free or instantaneous execution.
> Evaluator authority
>
> Policy and Rubric results are the correctness evidence admitted by the Benchmark Version. Rankings and telemetry summarize that evidence; they do not create a new standard.
{% example-demo title="Example: apparent gain from incomplete evidence" %}
Harness B leads the overall table, but List shows that twelve difficult Cases are still unevaluated for B. Arena also reports incomplete pairs. The operator waits for terminal evidence instead of starting Improve from a ranking that covers a smaller Case population.
{% /example-demo %}
## Object and state changes
Inspection, filtering, and navigation are read-only. Starting Improve, a Contribution, coverage work, or a later Run creates separate durable work while preserving the inspected evidence.
## Success criteria
- Identity, completeness, metric family, and uncertainty are explicit.
- Important signals resolve to Cases and admitted evaluator outcomes.
- The next action targets the responsible artifact or candidate boundary.
## Common failure modes
- Reporting rank without the evaluated population.
- Inventing execution traces from the unavailable segment.
- Starting candidate work when the Case or Rubric is wrong.
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Arena and Rankings](/docs/benchmark-evaluations/arena-and-rankings)
- [Dataset Snapshots](/docs/benchmark-datasets/snapshots)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
- [Benchmark runs](/docs/troubleshooting/benchmark-runs)
- [Missing outputs](/docs/troubleshooting/missing-outputs)
{% /related-card-grid %}
## Source confidence
Code-backed: Dashboard, List, Run detail, and workspace types establish result summaries, completion, rankings, repeated metrics, telemetry, and the current trace capability fence.