Teammately Docs
Docs menu

recipe

Check candidate correctness

Use Teammately benchmark evidence to compare a candidate behavior change against a baseline before internal human review.

Check candidate correctness

This page belongs with Benchmark Evaluations because it uses benchmark runs, comparisons, failures, and coverage gaps to prepare evidence for internal human review.

When to use this check

Use this when a candidate behavior change already has comparable benchmark runs and the team needs human-readable context before an internal reviewer decides what to do next.

Operating pattern

  1. Confirm the baseline and candidate runs use the intended benchmark version and run metadata.
  2. Inspect comparison results, must-level rubric failures, incomplete evidence, and coverage gaps.
  3. Separate improvements from regressions by policy, rubric, case segment, or benchmark dimension.
  4. Identify failures that require expert judgment before the candidate can be trusted.
  5. Prepare review context from Benchmark Evaluations results without claiming Teammately records the final approval.
  6. Link reviewers to the relevant run results, comparison view, and failing case details.

Worked example

Candidate review

01

Start

Behavior input

Raw case
A team compares the current assistant with a candidate assistant before human review.
02

Middle

Judgment into standard

Expert judgment
Downstream action should depend on approved standards, representative coverage, and explainable failures.
Policy
Must-level policies are hard gates; prefer-level policies are quality signals for tradeoff review.
Applicability
Applies to benchmark cases included in the approved benchmark version.
Binary rubric
Each applicable rubric produces pass, fail, or incomplete evidence for the candidate output.
03

Result

Benchmark output

Benchmark result
The candidate improves grounding but fails several must-level compatibility rubrics.
Human review context
If the customer assembles a review packet, it names the exact Benchmark Version and saved Harness Version and separates compatibility failures from unresolved coverage gaps.

Evidence to collect

  • Baseline run, candidate run, benchmark version, and run metadata.
  • Comparison results for must-level rubrics, incomplete evidence, regressions, and improvements.
  • Failing cases, relevant policies, applicability logic, and rubric results.
  • Coverage gaps or missing-output states that limit confidence.
  • Review notes that explain what Teammately evidence covers and what remains a human decision.

Source confidence

Code-backed: this page is grounded in benchmark evaluations routes and types. It does not claim a code-backed downstream approval workflow.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context