Teammately Docs
Docs menu

concept

The Teammately correctness loop

See how coverage, elicitation, case construction, evaluation, and improvement reinforce one another.

The Teammately correctness loop

The correctness loop is how a team repeatedly turns domain knowledge into stronger AI behavior. It follows the five public capabilities while preserving a trace from every result back to the project context, expert contribution, case, policy, rubric, benchmark version, Harness version, and evaluation setting that made the result meaningful.

Definition

  1. Design coverage. Establish Dimensions, Project Topics, and Case Construction Patterns, then decide which combinations the benchmark must represent.
  2. Elicit correctness. Use focused expert contributions to resolve policies, exceptions, applicability, disagreements, and binary rubric language.
  3. Construct the challenge set. Create or import canonical cases, attach required materials, generate difficult variants, and curate benchmark dataset membership.
  4. Evaluate behavior. Run an exact saved Harness Version against an exact Benchmark Version and inspect responses, Case-level Rubric evidence, comparisons, and rankings.
  5. Improve from evidence. Start an Improvement Session with a bounded Goal Contract, explore candidates, evaluate them through the canonical path, and retain a current frontier.
  6. Return new learning. Update coverage, correctness, cases, or the candidate according to what the evidence actually showed.

Decision checkpoint

Evidence says...Responsible part of the loopChange first
Important behavior has no casesCoverage Engineering or WeaveCoverage facet, construction pattern, or case set
Experts cannot apply the standard consistentlyCorrectness ElicitationPolicy scope, applicability, or rubric wording
A case cannot be interpreted or executed reliablyWeave and Project Input SchemaInput shape, case material, or world boundary
One saved candidate fails applicable rubricsTrialgroundHarness candidate or its runtime configuration
Several candidate branches improve different slicesCoevolveGoal constraints, next experiment, or retained frontier
Result movement cannot be explainedBenchmark version and evaluation boundaryVersions, settings, mapping, or run metadata before any product change

How expert effort compounds

The loop should ask an expert only after agents have prepared the relevant structure and evidence. A Contribution can include selected Cases, source attachments, scoped statements, draft Policies, Rubric questions, or coverage uncertainty. Completed expert work can materialize as an attributable contributed Policy, Rubric, Case, or coverage observation through the owning workflow.

That same judgment can guide future case construction, determine which rubrics apply during evaluation, and identify missing correctness during improvement. Reuse across the loop is more valuable than maximizing the number of disconnected review actions.

How product scope changes through the loop

Project foundations are reusable. Project Context, Reference Materials, policies, rubrics, Coverage Facets, Cases, and Harnesses do not belong to only one benchmark. A benchmark workspace selects and versions the relevant subset, manages coverage, coordinates contributions, evaluates candidates, and records improvement.

This scope distinction prevents accidental drift. Editing a project-level policy may affect several benchmarks. Changing dataset membership should create a new benchmark evidence boundary. Saving a Harness draft is different from selecting an exact saved Harness version for a Run.

Before and after

BeforeLoop workAfter
Domain knowledge is distributed across people and filesAgent Setup and Correctness Elicitation organize itProject context and governed correctness artifacts are inspectable
Examples are convenient rather than deliberateCoverage Engineering and Weave shape the challenge setDataset representation and missing coverage are explicit
Candidate behavior is discussed from anecdotesTrialground runs a versioned evaluationCase-level rubric evidence and comparisons are available
Improvement is a sequence of untracked editsCoevolve starts from pinned evidenceCandidate branches, receipts, chronology, and current frontier remain connected

Worked example

Changing a retrieval harness

An evaluation shows failures only when current and superseded documents appear together. The team first confirms that the coverage slice and grounding rubric are valid. An Improvement Session pins those cases and the failing Harness version, then tests source-date filtering and citation-selection candidates. A stronger candidate becomes part of the current frontier only after a canonical evaluation produces the expected rubric evidence. If the work uncovers an unseen source-conflict pattern, that observation returns to Coverage Management.

Source confidence

Doctrine-backed: this page explains the approved operating loop. Linked product pages are the authority for exact controls and lifecycle states.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context