Product quickstart
Run one narrow correctness loop. The goal is not a large benchmark; it is a traceable chain from project context and deliberate coverage to expert-grounded standards, a versioned evaluation, and one justified next change.
When to use it
Use this path for a new project or for an existing AI system whose correctness work is scattered across documents, examples, and informal expert feedback. Choose one behavior slice with a clear specialist owner.
Decision checkpoint
| Starting point | First action | Ready to continue when... |
|---|---|---|
| Agents do not understand the product or domain | Complete Agent Setup | Project Context and controlling Reference Materials are inspectable |
| Cases arrive in inconsistent shapes | Configure Project Input Schema | One input architecture and any required case materials are declared |
| Important behavior is not represented deliberately | Define Coverage Facets | Dimensions, Project Topics, and Case Construction Patterns name the slice |
| Correctness depends on tacit judgment | Request an Expert Contribution | The expert's scope, selected evidence, and required decisions are explicit |
| Cases and standards are ready | Create a benchmark snapshot and evaluate a saved Harness version | Exact cases, rubrics, candidate, and settings are bound to the Run |
Prerequisites
- A Teammately project for the specialist AI behavior.
- An accountable project operator and at least one domain expert.
- A small number of representative examples or enough Reference Materials to construct them.
- A candidate system that can be represented by a saved Harness version before evaluation.
Before and after
| Before | Action | After | Stop if... |
|---|---|---|---|
| Domain context is implicit | Write the Project Agent Brief and connect Reference Materials | Agents have explicit project understanding | Controlling sources are missing or contradictory without an owner |
| Inputs and supporting artifacts vary | Save Project Input Schema | Cases share one canonical content contract | Existing cases cannot satisfy the proposed schema |
| Expert knowledge is tacit | Request and complete a focused contribution | Policies, rubrics, cases, or coverage observations can be materialized | The request asks for a label without the evidence needed to explain it |
| Candidate behavior is anecdotal | Evaluate a saved Harness version | Results are traceable to cases and applicable rubrics | Dataset snapshot or candidate version is ambiguous |
| A weakness is confirmed | Start an Improvement Session from evidence | Candidate work follows a bounded Goal Contract | The requested outcome has no pinned measurement binding |
Steps
- Open or create the project and write the Project Agent Brief in Agent Setup → Project Context.
- Add controlling knowledge through Agent Setup → Reference Materials → Materials, then inspect the published blocks in Indexed Reference.
- Configure Project Settings → Input Schema. Select plain text, chat, or structured input and declare required case materials and accepted artifact families.
- Create the smallest useful set of Coverage Facets: a Dimension, relevant Project Topics, and a Case Construction Pattern for the chosen behavior slice.
- Add or construct cases in Assets, then select the intended cases in Benchmark Datasets. Confirm Representation and create or choose the appropriate snapshot.
- In Expert Contributions, request one focused contribution. Select the expert, state the objective, attach or select the relevant cases, and include only the contribution components needed to resolve the question.
- Inspect the completed contribution and materialize accepted policies, rubrics, cases, or coverage observations through their owning surfaces.
- Save an exact Harness version. In Benchmark Evaluations, configure and run it against the selected benchmark version.
- Inspect Dashboard and List results before using Compare or Arena. Trace important movement to case-level rubric evidence and run metadata.
- If a candidate change is justified, open Improve, start from the relevant evidence, prepare and confirm the Goal Contract, and evaluate candidate work through the canonical Run path.
Object and state changes
This path can create or update Project Context, Reference Materials items and indexed blocks, Project Input Schema, Coverage Facets, Cases, benchmark dataset membership and snapshots, Contributions, contributed artifacts, policies, rubrics, Harness drafts and saved versions, Runs, evaluation results, and Improvement Sessions. Each object keeps its own authority boundary; completing one step does not automatically approve or materialize every downstream artifact.
Success criteria
- Another operator can identify the project context and source material used by agents.
- The case set conforms to the Project Input Schema and represents a named coverage slice.
- Expert judgment is attributable to a completed Contribution and its accepted artifacts.
- The evaluation binds an exact benchmark version to an exact saved Harness version.
- Any improvement work starts from pinned evidence and records its Goal Contract, candidate results, and current frontier.
Common failure modes
- Treating Reference Materials as approved policies.
- Asking experts broad questions without selected cases or a concrete contribution objective.
- Evaluating an unsaved Harness draft or an unclear benchmark snapshot.
- Reading only an aggregate score and skipping failed case/rubric pairs.
- Starting improvement before the target and measurement evidence are resolved.
Related reference pages
Related troubleshooting pages
Source confidence
Doctrine-backed: this quickstart connects the current public story to code-backed product surfaces. Follow the linked pages for exact states and controls.