Teammately Docs
Docs menu

task

Write Binary Rubrics

Write atomic pass-or-fail criteria grounded in governed policies, applicable cases, and observable candidate behavior.

Write Binary Rubrics

Write a rubric when a governed policy needs an observable pass-or-fail check for benchmark evaluation. A strong rubric identifies one behavior, the cases where it applies, and the evidence that distinguishes pass from fail.

Prerequisites

  • A policy or expert-grounded correctness statement.
  • Representative passing, failing, and boundary cases.
  • Clear applicability for the behavior being checked.
  • Access to Correctness Governance → Rubrics.

Steps

  1. State one behavior that can be inspected in the candidate response and visible case evidence.
  2. Name the policy or specialist judgment that authorizes the criterion.
  3. Define applicability before writing exceptions into the pass condition.
  4. Write explicit pass evidence and fail evidence. Avoid “good,” “appropriate,” or “high quality” without observable conditions.
  5. Link representative cases and test whether two informed reviewers would reach the same binary result.
  6. Split independent requirements into separate rubrics when each failure should be diagnosed separately.
  7. Inspect contribution provenance and approval state before relying on the rubric in benchmark interpretation.

Object and state changes

This task creates or updates a project-level rubric and can change its wording, policy relationship, linked cases, evaluation use, activity, and approval context. Linking a case does not add it to a benchmark dataset. Editing a rubric does not alter historical Run evidence that used an earlier benchmark boundary.

Success criteria

  • The rubric tests one behavior and can be answered from visible evidence.
  • Applicability excludes irrelevant cases without hidden reviewer judgment.
  • Pass and fail conditions are explicit.
  • Linked cases include at least one meaningful boundary.
  • Policy authority and expert provenance are inspectable.

Common failure modes

  • Combining several behaviors into one criterion.
  • Restating the policy without defining observable evidence.
  • Encoding applicability only as exceptions inside the rubric.
  • Using a suggested or contributed draft as if it were already governed.
  • Changing rubric wording and comparing Runs without checking the benchmark version boundary.

Worked example

Example: grounding rubric

01

Middle

Judgment into standard

Policy
material claims must use the controlling source or state uncertainty. Rubric: pass only when every material claim is supported by the current controlling source, or the response explicitly says the available sources do not resolve the claim. Unsupported blending of current and superseded sources fails.

Source confidence

Code-backed: the active Correctness Governance rubric list and detail surfaces support rubric inspection, relationships, and lifecycle context. The drafting guidance is constrained to those verified artifact boundaries.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context