{"query":"Check candidate correctness","corpusVersion":"local","generatedAt":"2026-09-14T05:50:41.839Z","results":[{"blockId":"integrations.check-candidate-correctness#check-candidate-correctness","pageId":"integrations.check-candidate-correctness","title":"Check candidate correctness","pageTitle":"Check candidate correctness","url":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","humanUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness#check-candidate-correctness","markdownUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","sectionId":"check-candidate-correctness","kind":"recipe","productArea":"benchmark_evaluations","score":2065.470148645045,"reasons":["search_match","title_match","display_title_match","term_match"],"markdown":"# Check candidate correctness\n\nThis page belongs with Benchmark Evaluations because it uses benchmark runs, comparisons, failures, and coverage gaps to prepare evidence for internal human review."},{"blockId":"integrations.check-candidate-correctness#when-to-use-this-check","pageId":"integrations.check-candidate-correctness","title":"When to use this check","pageTitle":"Check candidate correctness","url":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","humanUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness#when-to-use-this-check","markdownUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","sectionId":"when-to-use-this-check","kind":"recipe","productArea":"benchmark_evaluations","score":1043.5763858703478,"reasons":["search_match","page_title_match","term_match"],"markdown":"## When to use this check\n\nUse this when a candidate behavior change already has comparable benchmark runs and the team needs human-readable context before an internal reviewer decides what to do next."},{"blockId":"integrations.check-candidate-correctness#operating-pattern","pageId":"integrations.check-candidate-correctness","title":"Operating pattern","pageTitle":"Check candidate correctness","url":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","humanUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness#operating-pattern","markdownUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","sectionId":"operating-pattern","kind":"recipe","productArea":"benchmark_evaluations","score":698.315985133204,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Operating pattern\n\n1. Confirm the baseline and candidate runs use the intended benchmark version and run metadata.\n2. Inspect comparison results, must-level rubric failures, incomplete evidence, and coverage gaps.\n3. Separate improvements from regressions by policy, rubric, case segment, or benchmark dimension.\n4. Identify failures that require expert judgment before the candidate can be trusted.\n5. Prepare review context from Benchmark Evaluations results without claiming Teammately records the final approval.\n6. Link reviewers to the relevant run results, comparison view, and failing case details.\n\n{% example-demo title=\"Candidate review\" %}\nRaw case: A team compares the current assistant with a candidate assistant before human review.\n\nExpert judgment: Downstream action should depend on approved standards, representative coverage, and explainable failures.\n\nPolicy: Must-level policies are hard gates; prefer-level policies are quality signals for tradeoff review.\n\nApplicability: Applies to benchmark cases included in the approved benchmark version.\n\nBinary rubric: Each applicable rubric produces pass, fail, or incomplete evidence for the candidate output.\n\nBenchmark result: The candidate improves grounding but fails several must-level compatibility rubrics.\n\nHuman review context: If the customer assembles a review packet, it names the exact Benchmark Version and saved Harness Version and separates compatibility failures from unresolved coverage gaps.\n{% /example-demo %}"},{"blockId":"integrations.check-candidate-correctness#evidence-to-collect","pageId":"integrations.check-candidate-correctness","title":"Evidence to collect","pageTitle":"Check candidate correctness","url":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","humanUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness#evidence-to-collect","markdownUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","sectionId":"evidence-to-collect","kind":"recipe","productArea":"benchmark_evaluations","score":695.1402314365465,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Evidence to collect\n\n- Baseline run, candidate run, benchmark version, and run metadata.\n- Comparison results for must-level rubrics, incomplete evidence, regressions, and improvements.\n- Failing cases, relevant policies, applicability logic, and rubric results.\n- Coverage gaps or missing-output states that limit confidence.\n- Review notes that explain what Teammately evidence covers and what remains a human decision."},{"blockId":"integrations.check-candidate-correctness#related-docs","pageId":"integrations.check-candidate-correctness","title":"Related docs","pageTitle":"Check candidate correctness","url":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","humanUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness#related-docs","markdownUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","sectionId":"related-docs","kind":"recipe","productArea":"benchmark_evaluations","score":685.5238409939274,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Related docs\n\n{% related-card-grid title=\"Related docs\" %}\n- [Read run results](/docs/benchmark-evaluations/inspect-results)\n- [Compare Harness Versions](/docs/benchmark-evaluations/compare)\n- [Run a benchmark](/docs/benchmark-evaluations/run-evaluation)\n{% /related-card-grid %}"},{"blockId":"integrations.check-candidate-correctness#source-confidence","pageId":"integrations.check-candidate-correctness","title":"Source confidence","pageTitle":"Check candidate correctness","url":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","humanUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness#source-confidence","markdownUrl":"https://teammately.ai/docs/integrations/check-candidate-correctness.md","sectionId":"source-confidence","kind":"recipe","productArea":"benchmark_evaluations","score":685.5238409939274,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Source confidence\n\nCode-backed: this page is grounded in benchmark evaluations routes and types. It does not claim a code-backed downstream approval workflow."},{"blockId":"governance.overview#governance-readiness-check","pageId":"governance.overview","title":"Governance readiness check","pageTitle":"Governance Overview","url":"https://teammately.ai/docs/governance.md","humanUrl":"https://teammately.ai/docs/governance#governance-readiness-check","markdownUrl":"https://teammately.ai/docs/governance.md","sectionId":"governance-readiness-check","kind":"concept","productArea":"governance","score":493.6177378953593,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Governance readiness check\n\n| Evidence is ready for accountable review when... | Hold the decision when... |\n| --- | --- |\n| Approved policies and rubrics are separated from suggestions and drafts. | Any governing standard is unapproved, stale, or ambiguous. |\n| Benchmark evidence names case, standard, benchmark, run, and candidate versions. | A score is detached from version boundaries. |\n| Access and reviewer roles are checked against source-backed permission pages. | A role label is used as a permission guarantee. |\n| Unsupported enterprise claims are left out or routed to source-backed docs. | The page implies compliance, retention, security, billing, or deployment behavior without evidence. |"},{"blockId":"coverage.benchmarks#operational-check","pageId":"coverage.benchmarks","title":"Operational check","pageTitle":"Benchmarks","url":"https://teammately.ai/docs/coverage-engineering/benchmarks.md","humanUrl":"https://teammately.ai/docs/coverage-engineering/benchmarks#operational-check","markdownUrl":"https://teammately.ai/docs/coverage-engineering/benchmarks.md","sectionId":"operational-check","kind":"concept","productArea":"coverage_engineering","score":468.45560583713257,"reasons":["search_match","term_match"],"markdown":"## Operational check\n\nBefore interpreting a Benchmark result, confirm the Benchmark purpose, exact Benchmark Version, selected Case population, evaluator boundary, saved Harness Version, and Run completeness. When the benchmark's intended behavior changes, update its coverage and dataset deliberately and create a new evidence boundary instead of treating current mutable state as historical truth.\n\n{% example-demo title=\"One benchmark, two evidence boundaries\" %}\nA support-assistant Benchmark initially covers ordinary return requests. After specialists document an exception for opened safety equipment, Coverage Management identifies the missing boundary and the current dataset gains reviewed Cases and a new Rubric relationship. The Benchmark remains the same program, but the team creates a new Snapshot and Benchmark Version. Comparisons name the version so readers can separate candidate improvement from the expanded correctness boundary.\n{% /example-demo %}"}]}