{"query":"Refreshing a Benchmark from New Signals","corpusVersion":"local","generatedAt":"2026-09-13T04:40:52.234Z","results":[{"blockId":"playbooks.benchmark-refresh#refreshing-a-benchmark-from-new-signals","pageId":"playbooks.benchmark-refresh","title":"Refreshing a Benchmark from New Signals","pageTitle":"Refreshing a Benchmark from New Signals","url":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","humanUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals#refreshing-a-benchmark-from-new-signals","markdownUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","sectionId":"refreshing-a-benchmark-from-new-signals","kind":"recipe","productArea":"playbooks","score":7063.435238905628,"reasons":["search_match","title_match","display_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Refreshing a Benchmark from New Signals\n\nUse this playbook when customer-provided production examples, Contribution findings, source changes, or target-system changes show that the current Benchmark no longer represents the intended behavior."},{"blockId":"playbooks.benchmark-refresh#classify-the-signal-before-editing","pageId":"playbooks.benchmark-refresh","title":"Classify the signal before editing","pageTitle":"Refreshing a Benchmark from New Signals","url":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","humanUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals#classify-the-signal-before-editing","markdownUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","sectionId":"classify-the-signal-before-editing","kind":"recipe","productArea":"playbooks","score":2673.120634082345,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Classify the signal before editing\n\nIdentify whether the signal changes the correctness standard, coverage vocabulary, available Cases, selected Dataset membership, or candidate behavior. These changes have different owners and version consequences."},{"blockId":"playbooks.benchmark-refresh#source-confidence","pageId":"playbooks.benchmark-refresh","title":"Source confidence","pageTitle":"Refreshing a Benchmark from New Signals","url":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","humanUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals#source-confidence","markdownUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","sectionId":"source-confidence","kind":"recipe","productArea":"playbooks","score":2430.959817528683,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Source confidence\n\nDoctrine-backed: the approved lifecycle routes new signals to the artifact that owns the change and preserves historical evidence. Linked code-backed pages define the current Coverage, Dataset, Snapshot, governance, and Evaluation operations."},{"blockId":"playbooks.benchmark-refresh#refresh-path","pageId":"playbooks.benchmark-refresh","title":"Refresh path","pageTitle":"Refreshing a Benchmark from New Signals","url":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","humanUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals#refresh-path","markdownUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","sectionId":"refresh-path","kind":"recipe","productArea":"playbooks","score":2407.1901048485324,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Refresh path\n\n1. Record the signal, source, effective date, and latest Benchmark Version it may affect.\n2. Route a source or rule change to Reference Materials, Expert Contributions, and Correctness Governance.\n3. Route a changed behavior axis to Coverage Facets and review the impact on existing classifications.\n4. Route missing examples to Assets Cases, Synthesis, or Coverage Management Case Foundry; review candidates in Case Review.\n5. Change selected membership under Benchmark Datasets and create a new Dataset Snapshot.\n6. Create the Benchmark Version that represents the revised Dataset and governed evaluator boundary.\n7. Run the same saved Harness Version against the new boundary when you need to isolate the benchmark change. Run a new Harness Version separately when candidate behavior also changed.\n8. Explain results using the two named boundaries; Compare does not implicitly normalize evidence across different Benchmark Versions."},{"blockId":"playbooks.benchmark-refresh#evidence-to-collect","pageId":"playbooks.benchmark-refresh","title":"Evidence to collect","pageTitle":"Refreshing a Benchmark from New Signals","url":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","humanUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals#evidence-to-collect","markdownUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","sectionId":"evidence-to-collect","kind":"recipe","productArea":"playbooks","score":2395.8218407922354,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Evidence to collect\n\n- The new signal, its source, and the benchmark version it affects.\n- Cases, reference responses where supported, Policies, Rubrics, or Coverage Facets changed by the signal.\n- Human approval for whether the change updates standards, coverage, or both.\n- Previous and refreshed Snapshot and Benchmark Version IDs, saved Harness Version, settings, and Run Metadata.\n- Case-level evidence that explains movement caused by the revised boundary rather than candidate behavior."},{"blockId":"playbooks.benchmark-refresh#preserve-history","pageId":"playbooks.benchmark-refresh","title":"Preserve history","pageTitle":"Refreshing a Benchmark from New Signals","url":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","humanUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals#preserve-history","markdownUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","sectionId":"preserve-history","kind":"recipe","productArea":"playbooks","score":2394.0295625552567,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Preserve history\n\nDo not edit an older Snapshot or Run to resemble current reality. Staleness means the evidence may no longer answer the current question; it does not erase what the old version measured.\n\n{% example-demo title=\"Refund policy change\" %}\nA support team adds a new outage-credit exception. The existing Benchmark remains valid for the former rule but has no Cases for the new exception. Experts approve revised applicability and Rubrics; Case Foundry prepares eligible, ineligible, and missing-account-evidence Cases; Case Review admits them into the current Dataset. A new Snapshot and Benchmark Version show that the unchanged Harness still passes ordinary refunds but fails the new outage boundary.\n{% /example-demo %}"},{"blockId":"playbooks.benchmark-refresh#related-docs","pageId":"playbooks.benchmark-refresh","title":"Related docs","pageTitle":"Refreshing a Benchmark from New Signals","url":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","humanUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals#related-docs","markdownUrl":"https://teammately.ai/docs/playbooks/refreshing-a-benchmark-from-new-signals.md","sectionId":"related-docs","kind":"recipe","productArea":"playbooks","score":2350.212723297513,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Related docs\n\n{% related-card-grid title=\"Related docs\" %}\n- [Coverage Refresh](/docs/coverage-engineering/coverage-refresh)\n- [Staleness Detection](/docs/governance/staleness-detection)\n- [Work with Dataset Snapshots](/docs/benchmark-datasets/snapshots)\n- [Compare Harness Versions](/docs/benchmark-evaluations/compare)\n- [Read run results](/docs/benchmark-evaluations/inspect-results)\n{% /related-card-grid %}"},{"blockId":"playbooks.overview#choose-a-playbook","pageId":"playbooks.overview","title":"Choose a playbook","pageTitle":"Enterprise Playbooks","url":"https://teammately.ai/docs/playbooks.md","humanUrl":"https://teammately.ai/docs/playbooks#choose-a-playbook","markdownUrl":"https://teammately.ai/docs/playbooks.md","sectionId":"choose-a-playbook","kind":"concept","productArea":"playbooks","score":981.4928604403548,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Choose a playbook\n\n| Starting problem | Playbook |\n| --- | --- |\n| Retrieved sources, citation, abstention, or answer grounding | [RAG correctness benchmark](/docs/playbooks/building-correctness-benchmark-rag) |\n| Search intent, source authority, document conflicts, or freshness | [Enterprise search](/docs/playbooks/enterprise-search) |\n| Refunds, commitments, account context, or escalation | [Customer support AI](/docs/playbooks/customer-support-ai) |\n| Many rules, exceptions, or controlled source hierarchies | [Policy-heavy AI systems](/docs/playbooks/policy-heavy-ai-systems) |\n| A bounded question requires accountable specialist judgment | [Run Expert Contributions](/docs/playbooks/running-expert-contributions-enterprise-assistant) |\n| Contribution evidence needs to become reusable standards | [Turn judgment into Policies and Rubrics](/docs/playbooks/turning-expert-judgment-into-policies-and-rubrics) |\n| Qualified experts disagree | [Handle conflicting opinions](/docs/playbooks/handling-conflicting-expert-opinions) |\n| Important behavior may be absent from the selected Dataset | [Find coverage gaps](/docs/playbooks/finding-coverage-gaps-before-review) |\n| New evidence or a changed rule makes the current boundary stale | [Refresh a Benchmark](/docs/playbooks/refreshing-a-benchmark-from-new-signals) |\n| CI or another evaluation system already owns execution facts | [Use existing evaluation infrastructure](/docs/playbooks/using-teammately-alongside-existing-evaluation-infrastructure) |"}]}