{"query":"Evaluation Execution Settings","corpusVersion":"local","generatedAt":"2026-09-13T04:33:45.141Z","results":[{"blockId":"benchmark-evaluations.execution-settings#evaluation-execution-settings","pageId":"benchmark-evaluations.execution-settings","title":"Evaluation Execution Settings","pageTitle":"Evaluation Execution Settings","url":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings#evaluation-execution-settings","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","sectionId":"evaluation-execution-settings","kind":"reference","productArea":"benchmark_evaluations","score":2219.830868222977,"reasons":["search_match","title_match","display_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Evaluation Execution Settings"},{"blockId":"benchmark-evaluations.execution-settings#definition","pageId":"benchmark-evaluations.execution-settings","title":"Definition","pageTitle":"Evaluation Execution Settings","url":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings#definition","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","sectionId":"definition","kind":"reference","productArea":"benchmark_evaluations","score":670.154920092424,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Definition\n\nEvaluation Settings contains machine settings for managed execution. Harness selection and repetition are choices made when launching an evaluation."},{"blockId":"benchmark-evaluations.execution-settings#source-confidence","pageId":"benchmark-evaluations.execution-settings","title":"Source confidence","pageTitle":"Evaluation Execution Settings","url":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings#source-confidence","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","sectionId":"source-confidence","kind":"reference","productArea":"benchmark_evaluations","score":643.1868467089499,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Source confidence\n\nCode-backed: Evaluation Settings defines machine configuration; the shared Run modal defines saved-Version selection and per-launch Run counts."},{"blockId":"benchmark-evaluations.execution-settings#related-task-pages","pageId":"benchmark-evaluations.execution-settings","title":"Related task pages","pageTitle":"Evaluation Execution Settings","url":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings#related-task-pages","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","sectionId":"related-task-pages","kind":"reference","productArea":"benchmark_evaluations","score":638.9767792940075,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Related task pages\n\n{% related-card-grid title=\"Related task pages\" %}\n- [Manage Harnesses](/docs/assets/harnesses)\n- [Run a benchmark evaluation](/docs/benchmark-evaluations/run-evaluation)\n{% /related-card-grid %}"},{"blockId":"benchmark-evaluations.execution-settings#before-launching","pageId":"benchmark-evaluations.execution-settings","title":"Before launching","pageTitle":"Evaluation Execution Settings","url":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings#before-launching","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","sectionId":"before-launching","kind":"reference","productArea":"benchmark_evaluations","score":633.7104441589864,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Before launching\n\nCheck the exact saved Version, machine selection, and requested execution volume. Runtime preparation and required access still apply. Imported external outputs remain a separate flow because they have no executable Harness Version."},{"blockId":"benchmark-evaluations.execution-settings#reliability-choices","pageId":"benchmark-evaluations.execution-settings","title":"Reliability choices","pageTitle":"Evaluation Execution Settings","url":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings#reliability-choices","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","sectionId":"reliability-choices","kind":"reference","productArea":"benchmark_evaluations","score":623.0174486940417,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Reliability choices\n\nDashboard, Compare, and Arena aggregate compatible observed Runs for each saved Harness Version across launch groups. Average score weights individual Runs equally. Passed at least once and passed every time summarize observed binary case outcomes where the evaluation framework supports them.\n\nRun counts can differ. The product displays the counts and coverage and warns about unequal evidence without requiring another launch. Missing or infrastructure-failed observations are not numerical successes or failures."},{"blockId":"benchmark-evaluations.execution-settings#fields-states-or-lifecycle-rules","pageId":"benchmark-evaluations.execution-settings","title":"Fields, states, or lifecycle rules","pageTitle":"Evaluation Execution Settings","url":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","humanUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings#fields-states-or-lifecycle-rules","markdownUrl":"https://teammately.ai/docs/benchmark-evaluations/execution-settings.md","sectionId":"fields-states-or-lifecycle-rules","kind":"reference","productArea":"benchmark_evaluations","score":613.9665969446235,"reasons":["search_match","page_title_match","term_match","prefix_or_fuzzy_match"],"markdown":"## Fields, states, or lifecycle rules\n\n- Choose an exact saved project Harness Version when launching. Benchmark activation is not required.\n- Choose a whole number of Runs from 1 through 50 for each selected Harness. The default is one, and different Harnesses may have different counts.\n- Each managed Harness launch creates one Run Group containing the requested Runs, including a one-Run group.\n- Existing Runs keep their immutable configuration and evidence. Launching more Runs creates another group rather than rewriting previous membership."},{"blockId":"project-settings.overview#project-settings","pageId":"project-settings.overview","title":"Project Settings","pageTitle":"Project Settings","url":"https://teammately.ai/docs/project-settings.md","humanUrl":"https://teammately.ai/docs/project-settings#project-settings","markdownUrl":"https://teammately.ai/docs/project-settings.md","sectionId":"project-settings","kind":"concept","productArea":"project_settings","score":438.49864029476925,"reasons":["search_match","term_match","prefix_or_fuzzy_match"],"markdown":"# Project Settings\n\nProject Settings is a floating project-level surface rather than a benchmark workspace. Its current tabs are **General**, **Regime**, **Input Schema**, and **Project Members**.\n\n| Setting | Governs | Does not replace |\n| --- | --- | --- |\n| General | Project name and Project Memo | Agent Setup context or instructions |\n| Regime | How approved and applicable Rubrics contribute to future Benchmark Versions | Policies, Rubrics, or already-published Benchmark evidence |\n| Input Schema | Canonical case input and material contract | A benchmark dataset snapshot |\n| Project Members | User and group access to this project | Contribution task assignment |\n\nChanges are project-scoped and may affect future work across multiple benchmarks. Treat input-contract and access changes as governance decisions, and preserve exact versions and snapshots wherever historical evidence depends on them."}]}