Teammately Docs
Docs menu

task

First correctness loop

Complete one traceable path from project context and benchmark coverage to expert judgment, evaluation evidence, and improvement.

First correctness loop

Complete one narrow loop that another operator can reconstruct. Choose one specialist behavior slice and preserve the path from project knowledge through coverage, expert contribution, governed standards, benchmark evidence, and any candidate change.

Decision checkpoint

StateNext actionDo not continue when...
Project intent or sources are implicitComplete Agent SetupAgents cannot find the controlling context
Case shape variesConfigure Project Input SchemaExisting and planned cases do not share a valid contract
Important behavior is unnamedDefine Coverage Facets and benchmark guidanceThe selected cases are merely convenient examples
Correctness remains tacitRequest a focused Expert ContributionThe expert lacks cases or source evidence
Cases and standards are readySnapshot the dataset and run an evaluationCandidate, benchmark, mapping, or settings are ambiguous
Candidate weakness is confirmedStart an Improvement SessionThe target cannot be measured from pinned evidence

Prerequisites

  • One project, one benchmark, and one narrow specialist behavior.
  • An accountable operator and domain expert.
  • Representative examples or enough Reference Materials to construct them.
  • A candidate that can be saved as a Harness version.

Before and after

BeforeWorkAfter
Knowledge is distributed across people and sourcesProject Context and Indexed ReferenceAgents have inspectable project understanding
Benchmark examples lack deliberate structureCoverage Facets, Coverage Management, and dataset selectionThe behavior slice and snapshot are explicit
Judgment is tacitExpert Contribution and Correctness GovernancePolicies and rubrics preserve authority and applicability
Candidate quality is anecdotalBenchmark EvaluationResponses and rubric results bind to exact versions
Improvement is an informal editImprovement SessionGoal, candidate, receipt, and frontier remain connected

Steps

  1. Write a concise Project Agent Brief and connect the controlling Reference Materials.
  2. Configure Project Input Schema for the input architecture and required case materials.
  3. Define the relevant Dimensions, Project Topics, and Case Construction Pattern.
  4. Add or construct a small case set, inspect its representation, and record any known gap.
  5. Request an Expert Contribution with selected cases and a concrete correctness objective.
  6. Reconcile the resulting policy, rubric, case, or coverage observation in its owning surface.
  7. Select the benchmark dataset cases and create or choose the intended snapshot.
  8. Save the candidate Harness version and run a Benchmark Evaluation.
  9. Inspect failures at case and rubric level; compare only after confirming evidence boundaries.
  10. Start an Improvement Session if candidate work is justified, or return upstream to the specific coverage, correctness, or case artifact that needs change.

Object and state changes

The loop can create or update project context, Reference Materials items and indexed blocks, Project Input Schema, Coverage Facets, Cases, benchmark coverage guidance, Contributions, contributed artifacts, policies, rubrics, dataset selection and snapshots, Harness versions, Runs, evaluation results, and Improvement Sessions. Each transition retains its own authority and scope.

Success criteria

  • The selected behavior slice has a named coverage reason.
  • Expert judgment is attributable and materialized only through an explicit lifecycle.
  • Case content follows the Project Input Schema.
  • Evaluation evidence identifies exact candidate and benchmark versions.
  • The next action names one responsible artifact or candidate boundary.

Common failure modes

  • Beginning with a broad benchmark and vague expert request.
  • Treating Reference Materials as governed standards.
  • Adding generated cases without a named coverage gap.
  • Running an editable Harness Draft.
  • Starting improvement from an aggregate result without pinned measurement evidence.

Worked example

Example: one exception slice

The first loop targets exception requests with conflicting sources. The project indexes both sources, defines the source-authority facet, asks an expert to establish the controlling rule, creates the corresponding rubric, snapshots ten reviewed cases, evaluates one saved Harness version, and starts improvement from the three exact grounding failures.

Source confidence

Doctrine-backed: this workflow applies the current five-capability model and links to code-backed pages for every exact product operation.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context