{"query":"First correctness loop","corpusVersion":"local","generatedAt":"2026-09-13T04:37:10.525Z","results":[{"blockId":"operating.first-correctness-loop#first-correctness-loop","pageId":"operating.first-correctness-loop","title":"First correctness loop","pageTitle":"First correctness loop","url":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","humanUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop#first-correctness-loop","markdownUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","sectionId":"first-correctness-loop","kind":"task","productArea":"operating_manual","score":2233.472399091324,"reasons":["search_match","title_match","display_title_match","term_match"],"markdown":"# First correctness loop\n\nComplete one narrow loop that another operator can reconstruct. Choose one specialist behavior slice and preserve the path from project knowledge through coverage, expert contribution, governed standards, benchmark evidence, and any candidate change."},{"blockId":"operating.first-correctness-loop#common-failure-modes","pageId":"operating.first-correctness-loop","title":"Common failure modes","pageTitle":"First correctness loop","url":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","humanUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop#common-failure-modes","markdownUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","sectionId":"common-failure-modes","kind":"task","productArea":"operating_manual","score":616.4129276121598,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Common failure modes\n\n- Beginning with a broad benchmark and vague expert request.\n- Treating Reference Materials as governed standards.\n- Adding generated cases without a named coverage gap.\n- Running an editable Harness Draft.\n- Starting improvement from an aggregate result without pinned measurement evidence.\n\n{% example-demo title=\"Example: one exception slice\" %}\nThe first loop targets exception requests with conflicting sources. The project indexes both sources, defines the source-authority facet, asks an expert to establish the controlling rule, creates the corresponding rubric, snapshots ten reviewed cases, evaluates one saved Harness version, and starts improvement from the three exact grounding failures.\n{% /example-demo %}"},{"blockId":"operating.first-correctness-loop#object-and-state-changes","pageId":"operating.first-correctness-loop","title":"Object and state changes","pageTitle":"First correctness loop","url":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","humanUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop#object-and-state-changes","markdownUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","sectionId":"object-and-state-changes","kind":"task","productArea":"operating_manual","score":600.0228788089621,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Object and state changes\n\nThe loop can create or update project context, Reference Materials items and indexed blocks, Project Input Schema, Coverage Facets, Cases, benchmark coverage guidance, Contributions, contributed artifacts, policies, rubrics, dataset selection and snapshots, Harness versions, Runs, evaluation results, and Improvement Sessions. Each transition retains its own authority and scope."},{"blockId":"operating.first-correctness-loop#steps","pageId":"operating.first-correctness-loop","title":"Steps","pageTitle":"First correctness loop","url":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","humanUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop#steps","markdownUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","sectionId":"steps","kind":"task","productArea":"operating_manual","score":587.6420839866031,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Steps\n\n1. Write a concise Project Agent Brief and connect the controlling Reference Materials.\n2. Configure Project Input Schema for the input architecture and required case materials.\n3. Define the relevant Dimensions, Project Topics, and Case Construction Pattern.\n4. Add or construct a small case set, inspect its representation, and record any known gap.\n5. Request an Expert Contribution with selected cases and a concrete correctness objective.\n6. Reconcile the resulting policy, rubric, case, or coverage observation in its owning surface.\n7. Select the benchmark dataset cases and create or choose the intended snapshot.\n8. Save the candidate Harness version and run a Benchmark Evaluation.\n9. Inspect failures at case and rubric level; compare only after confirming evidence boundaries.\n10. Start an Improvement Session if candidate work is justified, or return upstream to the specific coverage, correctness, or case artifact that needs change."},{"blockId":"operating.first-correctness-loop#before-and-after","pageId":"operating.first-correctness-loop","title":"Before and after","pageTitle":"First correctness loop","url":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","humanUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop#before-and-after","markdownUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","sectionId":"before-and-after","kind":"task","productArea":"operating_manual","score":586.2319877427936,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Before and after\n\n| Before | Work | After |\n| --- | --- | --- |\n| Knowledge is distributed across people and sources | Project Context and Indexed Reference | Agents have inspectable project understanding |\n| Benchmark examples lack deliberate structure | Coverage Facets, Coverage Management, and dataset selection | The behavior slice and snapshot are explicit |\n| Judgment is tacit | Expert Contribution and Correctness Governance | Policies and rubrics preserve authority and applicability |\n| Candidate quality is anecdotal | Benchmark Evaluation | Responses and rubric results bind to exact versions |\n| Improvement is an informal edit | Improvement Session | Goal, candidate, receipt, and frontier remain connected |"},{"blockId":"operating.first-correctness-loop#decision-checkpoint","pageId":"operating.first-correctness-loop","title":"Decision checkpoint","pageTitle":"First correctness loop","url":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","humanUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop#decision-checkpoint","markdownUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","sectionId":"decision-checkpoint","kind":"task","productArea":"operating_manual","score":585.5610371234912,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Decision checkpoint\n\n| State | Next action | Do not continue when... |\n| --- | --- | --- |\n| Project intent or sources are implicit | Complete Agent Setup | Agents cannot find the controlling context |\n| Case shape varies | Configure Project Input Schema | Existing and planned cases do not share a valid contract |\n| Important behavior is unnamed | Define Coverage Facets and benchmark guidance | The selected cases are merely convenient examples |\n| Correctness remains tacit | Request a focused Expert Contribution | The expert lacks cases or source evidence |\n| Cases and standards are ready | Snapshot the dataset and run an evaluation | Candidate, benchmark, mapping, or settings are ambiguous |\n| Candidate weakness is confirmed | Start an Improvement Session | The target cannot be measured from pinned evidence |"},{"blockId":"operating.first-correctness-loop#prerequisites","pageId":"operating.first-correctness-loop","title":"Prerequisites","pageTitle":"First correctness loop","url":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","humanUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop#prerequisites","markdownUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","sectionId":"prerequisites","kind":"task","productArea":"operating_manual","score":575.7820527724837,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Prerequisites\n\n- One project, one benchmark, and one narrow specialist behavior.\n- An accountable operator and domain expert.\n- Representative examples or enough Reference Materials to construct them.\n- A candidate that can be saved as a Harness version."},{"blockId":"operating.first-correctness-loop#success-criteria","pageId":"operating.first-correctness-loop","title":"Success criteria","pageTitle":"First correctness loop","url":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","humanUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop#success-criteria","markdownUrl":"https://teammately.ai/docs/operating-manual/first-correctness-loop.md","sectionId":"success-criteria","kind":"task","productArea":"operating_manual","score":575.7820527724837,"reasons":["search_match","page_title_match","term_match"],"markdown":"## Success criteria\n\n- The selected behavior slice has a named coverage reason.\n- Expert judgment is attributable and materialized only through an explicit lifecycle.\n- Case content follows the Project Input Schema.\n- Evaluation evidence identifies exact candidate and benchmark versions.\n- The next action names one responsible artifact or candidate boundary."}]}