Task index
Use this index when you know the work that must happen but need the current product surface. First decide whether the object is a reusable project foundation or belongs to one benchmark workspace.
Decision checkpoint
| Need | Open | Task |
|---|---|---|
| Give agents stable project understanding | Agent Setup | Maintain Project Context |
| Connect and verify project knowledge | Agent Setup → Reference Materials | Use Reference Materials |
| Define case input and material shape | Project Settings → Input Schema | Configure Project Input Schema |
| Govern policies and rubrics | Correctness Governance | Build policies and rubrics |
| Create or inspect reusable cases | Assets → Cases | Import and prepare cases |
| Edit a candidate implementation | Assets → Harnesses | Harnesses |
| Define reusable coverage structure | Coverage Facets | Coverage Engineering |
| Select benchmark cases and snapshots | Benchmark Datasets | Benchmark Datasets |
| Find and close coverage gaps | Coverage Management | Plan benchmark coverage |
| Ask a specialist for judgment | Expert Contributions | Request an Expert Contribution |
| Evaluate a saved candidate | Benchmark Evaluations | Run a Benchmark Evaluation |
| Diagnose candidate behavior | Dashboard, List, Compare, or Arena | Inspect evaluation results |
| Coordinate a justified candidate change | Improve | Start an Improvement Session |
Route by scope
Project foundations are reusable across benchmarks. Project Context, Reference Materials, policies, rubrics, Coverage Facets, Cases, Harnesses, and Project Input Schema belong at project scope. Changing one can affect future work in several benchmarks.
Benchmark work is deliberately scoped. Dataset selection and snapshots, Coverage Management, Expert Contributions, Benchmark Evaluations, and Improvement Sessions belong to the selected benchmark or benchmark version. Confirm the benchmark selector before making changes or interpreting evidence.
Route by evidence problem
If an evaluation fails, do not assume the Harness is responsible. An unclear Case belongs in Assets or case preparation. Missing behavior belongs in Coverage Management. Ambiguous correctness belongs in an Expert Contribution or Correctness Governance. A changed snapshot, setting, mapping, or metadata value belongs in evaluation diagnosis. Use Improve only when candidate work is justified by pinned evidence.
If agents lack source authority, update Reference Materials or Project Context before asking experts or generating more cases. If experts see the wrong fields or interaction, update Review Screen or the scoped Contribution rather than changing benchmark correctness.
Worked example
Route a grounding regression
A Run regresses on conflicting-source cases. The operator opens List and confirms that the cases, rubric, and settings are valid. Because the candidate selects a superseded document, the work belongs in Improve. If the expert could not determine which source controls, the same evidence would instead route to an Expert Contribution and Correctness Governance.
Related workflows
Related reference pages
Source confidence
Code-backed: the task routing follows current project and benchmark navigation and the active owning routes for each workflow.