{"query":"Create a benchmark","corpusVersion":"local","generatedAt":"2026-09-13T04:32:08.627Z","results":[{"blockId":"coverage.create-benchmark#create-a-benchmark","pageId":"coverage.create-benchmark","title":"Create a benchmark","pageTitle":"Create a benchmark","url":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","humanUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark#create-a-benchmark","markdownUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","sectionId":"create-a-benchmark","kind":"task","productArea":"coverage_engineering","score":1934.8896678254603,"reasons":["search_match","title_match","display_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Create a benchmark"},{"blockId":"coverage.create-benchmark#task-steps-create-and-prepare-a-benchmark","pageId":"coverage.create-benchmark","title":"Task steps: Create and prepare a benchmark","pageTitle":"Create a benchmark","url":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","humanUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark#task-steps-create-and-prepare-a-benchmark","markdownUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","sectionId":"task-steps-create-and-prepare-a-benchmark","kind":"task","productArea":"coverage_engineering","score":1461.355487929229,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"### Task steps: Create and prepare a benchmark\n\n1. Open the Benchmark selector in the project navigation.\n2. Select **Create New Benchmark**.\n3. Enter a name that identifies the target behavior or decision boundary, then select **Create Benchmark**.\n4. Open the new Benchmark and describe its purpose before curating evidence.\n5. Open **Coverage Management → Get Started** to define the benchmark denominator and coverage guidance.\n6. Use **Coverage Management** and **Benchmark Datasets** to prepare, review, and select Cases. Keep candidate material distinct from the selected Dataset.\n7. Create a Dataset Snapshot only when the selected Cases are ready to become an immutable evidence boundary."},{"blockId":"benchmark-datasets.snapshots#create-a-snapshot","pageId":"benchmark-datasets.snapshots","title":"Create a Snapshot","pageTitle":"Dataset Snapshots","url":"https://teammately.ai/docs/benchmark-datasets/snapshots.md","humanUrl":"https://teammately.ai/docs/benchmark-datasets/snapshots#create-a-snapshot","markdownUrl":"https://teammately.ai/docs/benchmark-datasets/snapshots.md","sectionId":"create-a-snapshot","kind":"task","productArea":"benchmark_datasets","score":999.5702320973692,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Create a Snapshot\n\nThe readiness check reports Case count, approved eligible Policy and Rubric counts, and blockers. Resolve every blocker before creation. Record a meaningful Snapshot label, then verify the displayed version, content hash, creation time, and Case count.\n\nCreation does not make weak input trustworthy. Review Case clarity, coverage, materials, and evaluator applicability first. After creation, do not describe later mutable classifications or links as if they were part of the frozen state."},{"blockId":"coverage.create-benchmark#name-benchmarks-for-durable-interpretation","pageId":"coverage.create-benchmark","title":"Name benchmarks for durable interpretation","pageTitle":"Create a benchmark","url":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","humanUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark#name-benchmarks-for-durable-interpretation","markdownUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","sectionId":"name-benchmarks-for-durable-interpretation","kind":"task","productArea":"coverage_engineering","score":915.769382806382,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Name benchmarks for durable interpretation\n\nPrefer a name such as **Support assistant — refund eligibility** over **August test**. Dates and change markers belong in Snapshots, Benchmark Versions, or version notes. The Benchmark name should remain meaningful as the Case set improves."},{"blockId":"coverage.dimensions-ontology#create-or-generate-a-schema","pageId":"coverage.dimensions-ontology","title":"Create or generate a schema","pageTitle":"Dimensions and Ontology","url":"https://teammately.ai/docs/coverage-engineering/dimensions-ontology.md","humanUrl":"https://teammately.ai/docs/coverage-engineering/dimensions-ontology#create-or-generate-a-schema","markdownUrl":"https://teammately.ai/docs/coverage-engineering/dimensions-ontology.md","sectionId":"create-or-generate-a-schema","kind":"reference","productArea":"coverage_engineering","score":847.065914212153,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"### Create or generate a schema\n\nCreate a Dimension manually when the axis and vocabulary are already understood. Use the dimension-schema generator when project context or source material should produce a reviewable proposal. Generated proposals can include a definition, why the Dimension matters, proposed ontology members, and warnings.\n\nA proposal is not the active schema. Review each proposed Dimension and value before accepting it. Avoid accepting near-duplicates simply because they use different wording.\n\n> Classification boundary\n>\n> Creating or editing a Dimension does not instantly classify every existing case. Missing or stale classifications can be queued and monitored separately. Treat unclassified cases as missing evidence, not as an implicit ontology value."},{"blockId":"benchmark-evaluations.run#evidence-created","pageId":"benchmark-evaluations.run","title":"Evidence created","pageTitle":"Run a Benchmark Evaluation","url":"https://teammately.ai/docs/benchmark-evaluations/run-evaluation.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/run-evaluation#evidence-created","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/run-evaluation.md","sectionId":"evidence-created","kind":"task","productArea":"benchmark_evaluations","score":653.690180398065,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Evidence created\n\nThe launch creates Run Groups and Runs bound to exact Harness and Benchmark Versions. Per-Case outputs and evaluator outcomes accrue separately, so output completion can precede evaluation completion. Provider telemetry can include tokens, cost, and latency when captured; absence of telemetry is not zero usage.\n\nDashboard aggregates compatible observed Runs across launches. Choose average score, passed at least once, or passed every time where supported. Each Run retains its own outputs and status; inspect the group and individual Runs when work is incomplete.\n\n> No Draft execution\n>\n> A managed benchmark Run does not evaluate the mutable Harness Draft. Save the candidate and select its exact saved Version when launching.\n\n{% example-demo title=\"Example: two candidates, three attempts\" %}\nHarness Versions 6 and 9 are active with `n=3`. One launch creates two Run Groups and six independent Runs against the same Benchmark Version. If one attempt fails preparation, the group reports incomplete evidence instead of silently treating the remaining two as the configured cohort.\n{% /example-demo %}"},{"blockId":"coverage.create-benchmark#common-failure-modes","pageId":"coverage.create-benchmark","title":"Common failure modes","pageTitle":"Create a benchmark","url":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","humanUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark#common-failure-modes","markdownUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","sectionId":"common-failure-modes","kind":"task","productArea":"coverage_engineering","score":651.8384460192758,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Common failure modes\n\n- Naming the Benchmark after a date or experiment rather than its durable behavior claim.\n- Treating creation as though a complete Benchmark Version or Snapshot now exists.\n- Selecting convenient Cases before defining the intended coverage boundary.\n- Starting Runs before selected Cases and governed standards are ready.\n\n{% example-demo title=\"Support escalation benchmark\" %}\nAn AI engineer creates **Support assistant — escalation decisions**. The description says the Benchmark tests whether the assistant escalates high-risk cases while resolving routine ones. The team plans risk, account tier, and source-authority coverage, then curates Cases from the Case Pool. Only after review does the team create its first Snapshot.\n{% /example-demo %}"},{"blockId":"coverage.create-benchmark#source-confidence","pageId":"coverage.create-benchmark","title":"Source confidence","pageTitle":"Create a benchmark","url":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","humanUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark#source-confidence","markdownUrl":"https://teammately.ai/docs/coverage-engineering/create-a-benchmark.md","sectionId":"source-confidence","kind":"task","productArea":"coverage_engineering","score":646.9763257681193,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Source confidence\n\nCode-backed: the Benchmark selector and creation modal define the current creation path; Coverage Management Get Started and Benchmark Datasets define the immediate next work. A newly created Benchmark is a durable workspace, not ready evaluation evidence."}]}