---
id: operating.prepare-review-packet
title: Prepare human review context
summary: Assemble customer-owned review context from exact evaluation, contribution, coverage, and improvement evidence.
kind: task
product_area: operating_manual
status: stable
updated: 2026-08-22
canonical: /docs/operating-manual/prepare-review-packet
---

# Prepare human review context

Assemble review context when an accountable customer team needs to inspect what the benchmark evidence says, why it says it, and which uncertainty or follow-up remains. This is a customer-owned packet or process, not a separate Teammately product object.

## Prerequisites

- Completed or clearly bounded Benchmark Evaluation evidence.
- Exact benchmark, dataset snapshot, Harness, Run, settings, and metadata identities.
- Relevant Expert Contributions and governed policies or rubrics.
- Coverage and Improvement Session context where it affects interpretation.

## Steps

1. State the review question and the downstream owner without implying that Teammately makes the final decision.
2. Identify the exact benchmark version, dataset snapshot, candidate Harness version, Runs, settings, and Run Metadata.
3. Summarize overall movement, then list material case-level gains, regressions, and uncertainty.
4. Link each important conclusion to applicable policies, rubrics, cases, and expert provenance.
5. Include relevant coverage gaps or representation limits.
6. Describe Improvement Session candidates and frontier evidence without claiming unobserved external-worker activity.
7. Separate confirmed findings, unresolved correctness, missing evidence, and recommended next investigation.
8. Preserve the source links or identifiers another reviewer needs to reproduce the interpretation.

## Object and state changes

Preparing context should read existing Teammately artifacts rather than mutate them. Follow-up work may create a Contribution, policy or rubric revision, Coverage Story, Case, dataset snapshot, Run, or Improvement Session. Keep the reviewed evidence unchanged so the reason for follow-up remains available.

## Success criteria

- Every conclusion is traceable to exact product evidence.
- Aggregate results are supported by case and rubric detail.
- Coverage limitations and unresolved expert disagreement are explicit.
- Historical and current candidate boundaries are distinguishable.
- The customer-owned downstream decision is not represented as a Teammately state.

## Common failure modes

- Copying a score without versions and settings.
- Omitting must-level regressions because the average improved.
- Treating an AI summary or trajectory as expert authority.
- Hiding missing coverage or unresolved source conflict.
- Describing a downstream choice as if Teammately automatically made it.

{% example-demo title="Example: review context for a retrieval change" %}
The packet names the two saved Harness versions, benchmark snapshot, grounding and uncertainty rubrics, and compared Runs. It highlights improved current-source cases, regressed missing-source cases, one unresolved expert contribution, and the Improvement Session frontier. The accountable team can inspect the evidence and decide its own next action.
{% /example-demo %}

## Related reference pages

{% related-card-grid title="Related reference pages" %}
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Expert Contributions](/docs/expert-contributions)
- [Improve](/docs/improve)
{% /related-card-grid %}

## Related troubleshooting pages

{% related-card-grid title="Related troubleshooting pages" %}
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
- [Low expert agreement](/docs/troubleshooting/low-expert-agreement)
{% /related-card-grid %}

## Source confidence

Doctrine-backed: this page defines the customer-owned human review boundary using current product artifacts without inventing a dedicated review-packet object or downstream-decision workflow.
