Teammately Docs
Docs menu

task

Operating Teammately end to end

Operate the current product from project foundations through benchmark coverage, expert contribution, evaluation, and improvement.

Operating Teammately end to end

Use this workflow to coordinate the full correctness system while keeping project foundations, benchmark work, expert authority, evaluation evidence, and candidate improvement separate.

Decision checkpoint

PhaseOwning scopeExit condition
Establish project understandingProjectProject Agent Brief and Indexed Reference are usable
Define content and reusable assetsProjectProject Input Schema, Cases, Harnesses, and Coverage Facets are explicit
Establish benchmark evidenceBenchmarkDataset snapshot and coverage state are reviewable
Resolve specialist correctnessBenchmark Contribution and project governanceAttributable artifacts have explicit lifecycle state
Evaluate candidatesBenchmark versionExact Runs and case/rubric evidence are available
Improve behaviorBenchmark versionGoal Contract, candidates, receipts, and frontier are durable

Prerequisites

  • A workspace and project.
  • An accountable operator, domain expert, and AI engineer or candidate owner.
  • Source knowledge, examples, and a candidate system appropriate to the intended benchmark.

Before and after

BeforeOperationAfter
Agents lack a shared project modelConfigure Agent SetupProject understanding is reusable and inspectable
Examples have inconsistent shapeSave Project Input Schema and prepare CasesInputs and materials share a canonical contract
Coverage and correctness are implicitDefine Coverage Facets and request ContributionsBenchmark intent and specialist standards are explicit
Candidate claims depend on anecdotesRun Benchmark EvaluationsEvidence is bound to versions, cases, and rubrics
Engineering iterations lack chronologyUse ImproveGoals, candidates, receipts, and current frontier stay connected

Steps

  1. Configure Project Context and Reference Materials in Agent Setup. Create or select Comparison Directions and Review Screens under Assets when a Contribution needs them.
  2. Save Project Input Schema and establish reusable Coverage Facets.
  3. Create or import Cases and save candidate Harness versions under Assets.
  4. Create or select a benchmark, configure Coverage Management, select Cases in Benchmark Datasets, inspect Representation, and preserve a snapshot.
  5. Request focused Expert Contributions for unresolved standards, cases, or coverage. Reconcile contributed artifacts in their owning project or benchmark surfaces.
  6. Confirm governed policies and rubrics and the benchmark version that will use them.
  7. Run exact saved Harness versions through Benchmark Evaluations. Inspect Dashboard, List, Compare, Arena, and Run detail according to the question.
  8. Start an Improvement Session only from evidence that identifies measurable candidate work.
  9. Return newly discovered correctness or coverage gaps to Expert Contributions, Correctness Governance, Coverage Management, or Cases.

Object and state changes

This workflow touches project context, reference indexes, input schema, facets, assets, benchmark datasets and snapshots, coverage state, Contributions and contributed artifacts, policies, rubrics, Harness versions, Runs, evaluation results, and Improvement Sessions. Each object remains in its owning scope and retains historical evidence.

Success criteria

  • Project and benchmark scope is explicit at every operation.
  • Agent preparation, expert judgment, and governed artifacts remain distinguishable.
  • Dataset and candidate versions make evaluation reproducible.
  • Improvement begins with a measurable goal and pinned evidence.
  • New learning returns to one responsible upstream artifact.

Common failure modes

  • Putting benchmark-specific instructions into permanent Project Context.
  • Treating connected sources as approved standards.
  • Selecting generated Cases without case review or schema conformance.
  • Comparing Runs after multiple evidence boundaries changed.
  • Treating external-worker activity as observable before an artifact returns.

Worked example

Example: full grounding loop

A team indexes source repositories, defines source-authority coverage, imports canonical cases, and requests a Contribution to resolve conflicting guidance. The governed rubric enters a benchmark version, two saved Harness versions are compared, and an Improvement Session tests retrieval changes. A missing-source pattern discovered during improvement returns to Coverage Management.

Source confidence

Doctrine-backed: the sequence follows the current public capability model and active product topology. Linked pages provide code-backed operation details.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context