Check candidate correctness
This page belongs with Benchmark Evaluations because it uses benchmark runs, comparisons, failures, and coverage gaps to prepare evidence for internal human review.
When to use this check
Use this when a candidate behavior change already has comparable benchmark runs and the team needs human-readable context before an internal reviewer decides what to do next.
Operating pattern
- Confirm the baseline and candidate runs use the intended benchmark version and run metadata.
- Inspect comparison results, must-level rubric failures, incomplete evidence, and coverage gaps.
- Separate improvements from regressions by policy, rubric, case segment, or benchmark dimension.
- Identify failures that require expert judgment before the candidate can be trusted.
- Prepare review context from Benchmark Evaluations results without claiming Teammately records the final approval.
- Link reviewers to the relevant run results, comparison view, and failing case details.
Worked example
Candidate review
Start
Behavior input
- Raw case
- A team compares the current assistant with a candidate assistant before human review.
Middle
Judgment into standard
- Expert judgment
- Downstream action should depend on approved standards, representative coverage, and explainable failures.
- Policy
- Must-level policies are hard gates; prefer-level policies are quality signals for tradeoff review.
- Applicability
- Applies to benchmark cases included in the approved benchmark version.
- Binary rubric
- Each applicable rubric produces pass, fail, or incomplete evidence for the candidate output.
Result
Benchmark output
- Benchmark result
- The candidate improves grounding but fails several must-level compatibility rubrics.
- Human review context
- If the customer assembles a review packet, it names the exact Benchmark Version and saved Harness Version and separates compatibility failures from unresolved coverage gaps.
Evidence to collect
- Baseline run, candidate run, benchmark version, and run metadata.
- Comparison results for must-level rubrics, incomplete evidence, regressions, and improvements.
- Failing cases, relevant policies, applicability logic, and rubric results.
- Coverage gaps or missing-output states that limit confidence.
- Review notes that explain what Teammately evidence covers and what remains a human decision.
Related docs
Source confidence
Code-backed: this page is grounded in benchmark evaluations routes and types. It does not claim a code-backed downstream approval workflow.