# Building a Correctness Benchmark for a RAG System Generated: 2026-09-14T05:50:32.742Z Source build: local Canonical docs: https://teammately.ai/docs --- id: playbooks.rag-benchmark title: Building a Correctness Benchmark for a RAG System summary: Represent retrieval-grounded behavior through cases, context, policies, rubrics, and benchmark evidence. kind: recipe product_area: playbooks status: stable updated: 2026-08-23 canonical: /docs/playbooks/building-correctness-benchmark-rag --- # Building a Correctness Benchmark for a RAG System Use this playbook when correctness depends on retrieved context, source authority, and whether the answer should cite, abstain, or explain uncertainty. ## Entry conditions Use this when you can preserve the query, retrieved material, candidate response, and source identity for representative RAG behavior. If you have only aggregate retrieval metrics, first collect Case-level evidence; Teammately cannot infer source authority from a score. ## Route through Teammately 1. In **Agent Setup**, make the Project Agent Brief describe the retrieval architecture and connect the Reference Materials needed to interpret sources. 2. Configure Project Input Schema fields for the query, retrieved passages, source identifiers, and freshness or authority metadata actually available to the Harness. 3. Import representative Cases under **Assets → Cases**. Keep missing-source and conflicting-source Cases instead of filtering them out as bad data. 4. In **Coverage Facets**, model the slices that change grounding behavior: authority, freshness, answerability, retrieval completeness, and question type. 5. Request an **Expert Contribution** for Cases where the controlling source, required caveat, or abstention boundary is unclear. 6. Materialize and approve the resulting Policies and binary Rubrics in **Correctness Governance**. 7. Use **Coverage Management** to expose missing combinations, review new Cases, and select the intended set in **Benchmark Datasets**. 8. Run saved Harness Versions in **Benchmark Evaluations**. Read response and Rubric evidence; execution traces are not currently exposed. ## Decision gates - If the correct source was never retrieved, route the finding to retrieval or coverage work. - If the source was present but the response blended, ignored, or contradicted it, route the finding to Harness behavior. - If specialists disagree about which source controls, resolve correctness before expanding the Dataset. - If an important source condition has too few Cases, hold aggregate interpretation until representation improves. {% example-demo title="Benefits policy retrieval" %} An employee asks whether caregiver leave applies to contractors. The Case contains an obsolete handbook page and the current controlling policy, which does not state contractor eligibility. Experts approve a Policy requiring the controlling source and a Rubric that passes only when the answer cites it and withholds the unsupported eligibility claim. Results show one Harness succeeds when both passages are retrieved but still fails when the current policy is absent, separating answer behavior from retrieval coverage. {% /example-demo %} ## Evidence to collect - Canonical Case input containing the query and material actually available at execution. - Source identifiers, authority, and freshness facts that reviewers can verify. - Approved grounding, citation, contradiction, and abstention standards. - Dataset representation across answerable, conflicting, stale, missing, and multi-source conditions. - Saved Harness Version, Benchmark Version, execution settings, Run Metadata, response, and Rubric outcomes. ## Related docs {% related-card-grid title="Related docs" %} - [Configure Coverage Management](/docs/coverage-management/get-started) - [Represent conversations in Cases](/docs/object-model/represent-conversations-in-cases) - [Compare Harness Versions](/docs/benchmark-evaluations/compare) - [Read run results](/docs/benchmark-evaluations/inspect-results) - [Run a benchmark](/docs/benchmark-evaluations/run-evaluation) - [Importing cases](/docs/operating-manual/import-and-prepare-cases) {% /related-card-grid %} ## Source confidence Doctrine-backed: the approved five-capability model establishes the RAG correctness loop. Linked code-backed pages define the current Agent Setup, Case, Coverage Management, Expert Contribution, Dataset, and Evaluation surfaces and their capability fences. --- id: coverage.plan-benchmark-coverage title: Plan Benchmark Coverage summary: Apply project Coverage Facets to one benchmark, inspect representation, and turn important gaps into concrete case or contribution work. kind: task product_area: coverage_engineering status: stable updated: 2026-09-07 canonical: /docs/coverage-engineering/plan-benchmark-coverage --- # Plan Benchmark Coverage Plan coverage by applying reusable project facets to one benchmark and comparing the intended behavior space with the selected dataset representation. ## Prerequisites - A selected project and benchmark. - A clear benchmark purpose. - Relevant Dimensions, Project Topics, and Case Construction Patterns, or enough project knowledge to create them. - Existing Cases or a plan for sourcing and constructing them. ## Steps 1. Review **Coverage Facets** at project scope. Confirm that Dimensions, Project Topics, and Case Construction Patterns describe reusable behavior structure rather than one benchmark's current case count. 2. Open the benchmark and select **Coverage Management → Get Started**. 3. Define the benchmark-specific coverage guidance and confirm setup readiness. 4. Open Coverage Management and inspect current dataset representation across the relevant facets and tuples. 5. Name important thin or absent combinations as Coverage Stories. Explain why each slice matters and what evidence would make it usable. 6. Route the gap according to its cause: Case Foundry or case sourcing for missing situations, Expert Contributions for missing judgment, Correctness Governance for missing standards, or Benchmark Datasets for missing selection. 7. Review generated or contributed cases in Case Review before relying on them. 8. Update dataset selection and create a new snapshot when the represented evidence changes materially. ## Object and state changes This task can update benchmark coverage setup, representation guidance, Coverage Stories, Case Foundry work, case-review state, contribution requests, dataset selection, and snapshots. Project Coverage Facets may also change when the work discovers a reusable missing axis or construction pattern. ## Success criteria - The benchmark purpose maps to explicit project Coverage Facets. - Important combinations have selected evidence or a named gap. - Each gap is routed to a responsible artifact or workstream. - Constructed cases pass case review and Project Input Schema checks. - Dataset snapshots make material coverage changes explicit. ## Common failure modes - Using case count as the coverage goal. - Creating benchmark-only tags where a reusable Dimension or Topic is needed. - Treating response-variation guidance as coverage structure. - Generating cases before defining which gap they should close. - Trusting representation after selection changes without a new snapshot boundary. {% example-demo title="Example: plan high-impact exception coverage" %} The team maps exception type, source authority, and customer impact. Representation shows many low-impact ordinary cases but no high-impact cases with conflicting authority. A Coverage Story names the gap, an expert Contribution clarifies the controlling rule, and Case Foundry prepares cases for the missing tuple before a new snapshot is created. {% /example-demo %} ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Coverage Engineering](/docs/coverage-engineering) - [Coverage Management](/docs/coverage-management) - [Benchmark Datasets](/docs/benchmark-datasets) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage) - [Stale dimensions](/docs/troubleshooting/stale-dimensions) - [Unrealistic synthetic cases](/docs/troubleshooting/unrealistic-synthetic-cases) {% /related-card-grid %} ## Source confidence Code-backed: current setup, overview, representation, Coverage Story, Case Foundry, and Case Review routes support this workflow. --- id: agent-setup.project-context title: Project Context summary: Maintain the Project Agent Brief that gives Teammately agents stable, project-wide understanding. kind: reference product_area: agent_setup status: stable updated: 2026-08-22 canonical: /docs/agent-setup/project-context --- # Project Context ## Definition Project Context is the Agent Setup surface that edits the **Project Agent Brief**. The brief gives Teammately agents stable project-wide understanding: what the specialist AI is for, which behavior matters, important constraints, terminology, and other context that should carry across coverage, contribution preparation, case construction, and improvement work. The Project Agent Brief is different from the Project Memo in General settings. The memo is administrative project text and explicitly is not used as prompt or agent context. Put agent-relevant project understanding in Project Context. ## Fields, states, or lifecycle rules - The brief is scoped to the project and reused across benchmark workspaces. - Editing the brief changes future agent context; it does not rewrite completed Contributions, Runs, or Improvement Session history. - The brief provides orientation and constraints, not governed correctness authority. Policies and rubrics remain in Correctness Governance. - Controlling source material belongs in Reference Materials. Summarize stable project intent in the brief and keep source-backed detail in the indexed material. - The brief should state product-specific meaning directly. Avoid copying transient benchmark goals, one expert's unconfirmed opinion, or a temporary candidate hypothesis into permanent project context. ## Writing a useful brief Describe the specialist AI's purpose, users, important domain vocabulary, expected interaction shape, and constraints that affect many workflows. Include explicit boundaries where agents might otherwise make unsafe assumptions. Name controlling authorities without duplicating entire manuals. Review the brief when the product purpose, domain, input architecture, or correctness boundary changes materially. If only one benchmark needs a special objective, put it in that benchmark's setup or Contribution. If only one Improvement Session needs a constraint, put it in the Goal Contract. {% example-demo title="Example: project context boundary" %} The brief states that a procurement assistant supports internal buyers, must distinguish current agreements from expired ones, and should expose uncertainty rather than invent an exception. The current agreements themselves remain indexed Reference Materials. The exact evaluation target for expired-agreement cases belongs to the benchmark and Improvement Session, not the brief. {% /example-demo %} ## Source confidence Code-backed: the active Agent Setup Project Context route renders the Project Agent Brief editor. The distinction from General settings is supported by the current project settings UI. ## Related task pages {% related-card-grid title="Related task pages" %} - [Product quickstart](/docs/quickstart) - [Use Reference Materials](/docs/agent-setup/reference-materials) - [Request an Expert Contribution](/docs/expert-contributions/request-contribution) {% /related-card-grid %} --- id: benchmark-evaluations.inspect title: Inspect Evaluation Results summary: Trace Dashboard and List signals to Run, Case, Policy, Rubric, completeness, and telemetry evidence. kind: task product_area: benchmark_evaluations status: stable updated: 2026-09-13 canonical: /docs/benchmark-evaluations/inspect-results --- # Inspect Evaluation Results ## Prerequisites - A visible Run or Run Group with output or evaluation progress. - Access to the exact Benchmark and Harness Version evidence. Start with completeness and identity, then move from aggregate signals to the Cases and evaluator failures that support them. ## Steps 1. Open Dashboard and confirm the Benchmark Version, candidate Harness Version, Run Group type, attempt count, and evaluation progress. 2. Read rankings with their metric family and uncertainty. Distinguish average score, passed at least once, and passed every time over observed Runs. Inspect Run counts and coverage; historical group-specific pass@n and pass^n retain their original meanings. 3. Open **List → Runs** to inspect group and Run status, output progress, evaluation progress, metadata, and available resource telemetry. 4. Open **List → Evaluation results** for the Case summary, outcome, failed Policies, failed Rubrics, and evaluated count. 5. Use Arena for pairwise disagreement or Compare for a Harness matrix across Cases, evaluator facts, or Coverage Facets. 6. Classify the next action as candidate work, evaluator clarification, Case correction, coverage work, external-output remapping, or no action. The List results surface is intentionally compact. Do not claim that it exposes full execution trajectories. The **Traces / Spans** segment currently reports a capability fence because the benchmark API does not provide evaluation execution traces. ## Reading incomplete and repeated evidence An aggregate calculated over fewer evaluable Cases can look better while covering less evidence. Record evaluated, incomplete, and missing counts before comparing candidates. For repeated groups, inspect whether the configured number of attempts exists for every candidate and whether one failed attempt changes the metric interpretation. Cost, tokens, and latency help route operational work but are nullable telemetry. Missing capture means unknown, not free or instantaneous execution. > Evaluator authority > > Policy and Rubric results are the correctness evidence admitted by the Benchmark Version. Rankings and telemetry summarize that evidence; they do not create a new standard. {% example-demo title="Example: apparent gain from incomplete evidence" %} Harness B leads the overall table, but List shows that twelve difficult Cases are still unevaluated for B. Arena also reports incomplete pairs. The operator waits for terminal evidence instead of starting Improve from a ranking that covers a smaller Case population. {% /example-demo %} ## Object and state changes Inspection, filtering, and navigation are read-only. Starting Improve, a Contribution, coverage work, or a later Run creates separate durable work while preserving the inspected evidence. ## Success criteria - Identity, completeness, metric family, and uncertainty are explicit. - Important signals resolve to Cases and admitted evaluator outcomes. - The next action targets the responsible artifact or candidate boundary. ## Common failure modes - Reporting rank without the evaluated population. - Inventing execution traces from the unavailable segment. - Starting candidate work when the Case or Rubric is wrong. ## Related reference pages {% related-card-grid title="Related reference pages" %} - [Benchmark Evaluations](/docs/benchmark-evaluations) - [Arena and Rankings](/docs/benchmark-evaluations/arena-and-rankings) - [Dataset Snapshots](/docs/benchmark-datasets/snapshots) {% /related-card-grid %} ## Related troubleshooting pages {% related-card-grid title="Related troubleshooting pages" %} - [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly) - [Benchmark runs](/docs/troubleshooting/benchmark-runs) - [Missing outputs](/docs/troubleshooting/missing-outputs) {% /related-card-grid %} ## Source confidence Code-backed: Dashboard, List, Run detail, and workspace types establish result summaries, completion, rankings, repeated metrics, telemetry, and the current trace capability fence.