# Work and Evolve
Generated: 2026-09-13T04:36:27.341Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: improve.work-evolve
title: Work and Evolve
summary: Choose bounded implementation work or multi-branch evolutionary search with explicit epoch and provider authorization.
kind: task
product_area: improve
status: stable
updated: 2026-09-07
canonical: /docs/improve/work-and-evolve
---
# Work and Evolve
## Prerequisites
- A confirmed Goal Contract with exact starting evidence.
- A saved target Harness Version and canonical measurement bindings.
- Authorization for the selected worker source, epochs, and provider use.
Within the Coevolve experience, choose **Work** for one bounded implementation or investigation and **Evolve** when the confirmed Goal supports systematic exploration across several candidate branches and epochs. External Agents and External fine-tuning use different execution boundaries and should be selected only when their return or provider contracts are clear.
## Work
Work can run through Teammately or an offered external execution source such as Codex, Claude, or another worker. The source receives the confirmed Goal and scoped evidence. It must return an observable immutable Harness Version or evaluation request; a handoff package alone is not proof that private work occurred.
Use Work when the likely intervention is known, the change is narrow, or operator review should follow one candidate at a time. Review returned source, candidate identity, evaluation receipt, constraint result, and chronology before treating the work as complete.
## Evolve
Evolve uses Teammately and explores three Patch-to-Eval branches per epoch against the full pinned cohort. Choose **user-gated** to approve each epoch or **autonomous** to authorize a bounded number of epochs from 1 through 100. Set the target Case pass percentage and review provider usage by model.
An epoch does not merely generate text. Each viable branch must materialize an exact saved Harness Version and obtain canonical evaluation evidence before retention. Failed, incomparable, or constraint-violating branches remain visible rather than being presented as improvement.
## Select the mode
Use Work when evidence points to a specific retrieval filter, prompt rule, tool call, or output mapping change. Use Evolve when multiple independent interventions could satisfy the Goal and the benchmark can distinguish them. Do not use autonomous epochs when the Goal has unresolved authority, the evaluation boundary is unstable, or provider usage is not authorized.
> Epoch authorization
>
> User-gated and autonomous execution change how future epochs are scheduled, not the acceptance criteria. Every retained candidate still needs exact identity, canonical evaluation, and Goal-constraint compliance.
## Object and state changes
Starting Work or Evolve advances the Session and can create handoff packages, epoch branches, saved Harness Versions, evaluations, receipts, usage, and frontier decisions. Pause and cancellation preserve recorded evidence.
## Success criteria
- Mode and execution source match the problem.
- Every retained branch has exact candidate identity and canonical evidence.
- Epoch limits, pass target, and provider usage remain within authorization.
## Common failure modes
- Using Evolve before the benchmark can distinguish hypotheses.
- Treating a handoff package as returned implementation.
- Retaining a focused-only or constraint-violating candidate.
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Goal Contracts](/docs/improve/goal-contracts)
- [Candidates and the Current Frontier](/docs/improve/candidates-and-frontier)
- [Chronology, Trajectories, and Receipts](/docs/improve/chronology-and-trajectories)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Benchmark runs](/docs/troubleshooting/benchmark-runs)
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
{% /related-card-grid %}
## Source confidence
Code-backed: the active Improve setup, session contract, and work-review model define sources, Work and Evolve modes, authorization styles, epoch bounds, three branches, pinned cohort, pass target, and provider usage.
---
id: improve.start-session
title: Start an Improvement Session
summary: Start from benchmark evidence, prepare a measurable Goal Contract, and choose bounded Work or Evolve behavior.
kind: task
product_area: improve
status: stable
updated: 2026-09-07
canonical: /docs/improve/start-improvement-session
---
# Start an Improvement Session
Start an Improvement Session when evaluation evidence justifies a candidate change or bounded investigation. The setup should turn a free-form intention into a measurable Goal Contract before work begins.
Choose the saved Harness Version and a specific baseline launch. A baseline may contain one Run. Set the number of Runs for future candidate evaluations independently; a difference from the baseline count is informational. The work forecast and authorized budget use the chosen candidate count.
## Prerequisites
- A selected benchmark and evidence that identifies the relevant Benchmark Version or Run.
- An existing target Harness and saved starting version.
- Case, rubric, Run, comparison, or frontier evidence that explains the need.
- A measurable outcome and constraints that should remain protected.
- An operator authorized to start and control the session.
## Steps
1. Open the selected benchmark and choose **Improve**.
2. Choose the available Improve experience, then create a new Improvement Session and select **Work** or **Evolve** when using Coevolve.
3. Select the target Harness, exact starting Harness Version, starting Run, and execution source.
4. For Work, choose the available Coevolve or External Agents path. For Evolve, use Coevolve and choose user-gated or autonomous execution. External fine-tuning has its own provider and return boundary when available.
5. State the desired behavior change and important non-regression constraints.
6. Prepare the Goal Contract. Resolve canonical target identities, objectives, measurement bindings, intervention constraints, and unresolved items.
7. Inspect the proposed revision and confirm it only when the evidence can measure the requested outcome.
8. For Evolve, configure epoch authorization, Case pass target, and provider usage bounds before starting.
9. If using an external worker, verify the scoped package and return contract after the Goal is confirmed.
10. Start the session and use chronology, trajectories, candidates, receipts, and current frontier to follow observable progress.
## Object and state changes
This task creates a benchmark-scoped Improvement Session, records its mode, experience, source, target, and pinned starting evidence, and establishes a Goal Contract revision. Session lifecycle states are draft, active, paused, completing, completed, cancelling, cancelled, or failed, with attention states when operator action is needed. Starting work can create worker packages, candidate Harness Versions, canonical evaluation requests and receipts, frontier changes, chronology events, and usage records.
## Success criteria
- The target and starting evidence use canonical identities.
- Every objective has an observable measurement binding.
- Constraints protect important behavior from hidden regression.
- Work or Evolve is chosen deliberately.
- The exact starting Harness Version and Run are visible.
- Candidate progress is supported by returned artifacts and evaluation receipts.
- The current frontier is explainable from the Goal Contract and evidence.
## Common failure modes
- Starting from an aggregate score without selected case or rubric evidence.
- Confirming a Goal Contract whose outcome cannot be measured.
- Allowing Evolve without bounded authorization.
- Treating a worker package as proof that private work occurred.
- Retaining the newest candidate without checking constraints and regressions.
- Changing the benchmark boundary during the session without making the new evidence explicit.
{% example-demo title="Example: bounded Work session" %}
A Run fails three cases because the Harness uses a superseded source. The operator pins those cases and the grounding rubric, targets the exact saved Harness version, and writes a Goal Contract requiring current-source selection without reducing missing-source uncertainty performance. Work begins only after both objectives have measurement bindings.
{% /example-demo %}
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Improve](/docs/improve)
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Harnesses](/docs/assets/harnesses)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Benchmark runs](/docs/troubleshooting/benchmark-runs)
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
{% /related-card-grid %}
## Source confidence
Code-backed: the current Improve setup, session command, and Goal Contract behavior support this workflow. External workers remain bounded by observable return artifacts and requests.
---
id: improve.goal-contracts
title: Goal Contracts
summary: Bind an Improvement Session to exact targets, measurable objectives, constraints, and permitted intervention scope.
kind: reference
product_area: improve
status: stable
updated: 2026-09-07
canonical: /docs/improve/goal-contracts
---
# Goal Contracts
## Definition
A Goal Contract is the authority and measurement boundary for one Improvement Session. It turns a natural-language intent into canonical targets, prioritized objectives, protected constraints, and evaluable bindings.
## Fields, states, or lifecycle rules
### Contract contents
The contract identifies exact target IDs and records intent. Objectives carry priorities. Constraints are hard or soft. Measurement bindings name whether evidence should **improve**, **preserve**, or **reduce** a canonical Case, Policy, Rubric, metric, or other admitted reference. Intervention constraints state what work may change. Unresolved items remain explicit until the contract can be confirmed safely.
Every requested outcome needs an observable binding. “Make answers better” is not measurable; “improve the selected grounding Rubric while preserving the selected uncertainty Cases” is. A hard constraint disqualifies a candidate when violated. A soft constraint records a tradeoff that still requires review.
### Revisions and confirmation
Goal Contract revisions can be proposed, confirmed, superseded, or rejected. Keep superseded and rejected revisions as history so candidate activity can be explained against the contract that authorized it. Do not silently edit the meaning of a session after work has begun.
Evolve locks the Goal Contract after the first epoch begins. If new evidence reveals a fundamentally different goal, stop or complete the current session and create an explicit new boundary rather than retrofitting prior epochs.
### Confirmation checklist
- Target Harness, Harness Version, Benchmark Version, and starting Run resolve to exact identities.
- Each objective has a direction and canonical measurement binding.
- Non-regression behavior is represented by preserve bindings or constraints.
- Hard and soft constraints are distinguishable.
- Intervention scope permits the intended code, prompt, retrieval, or configuration work.
- No unresolved item makes evaluation or authorization ambiguous.
> Authority boundary
>
> The Goal Contract authorizes session work; it does not change the Benchmark Version, approve a new correctness standard, or waive human governance of upstream artifacts.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Start an Improvement Session](/docs/improve/start-improvement-session)
- [Choose Work or Evolve](/docs/improve/work-and-evolve)
{% /related-card-grid %}
## Source confidence
Code-backed: the active session contract and commands define canonical targets, objectives, constraints, measurement bindings, intervention scope, revision states, confirmation, and Evolve locking.
---
id: improve.chronology-trajectories
title: Chronology, Trajectories, and Receipts
summary: Read durable session events, safe narrated work segments, evaluation receipts, external handoffs, and usage evidence.
kind: reference
product_area: improve
status: stable
updated: 2026-09-07
canonical: /docs/improve/chronology-and-trajectories
---
# Chronology, Trajectories, and Receipts
## Definition
Chronology is the durable ordered record of an Improvement Session. Event blocks can represent user or assistant messages, activity, Harness output, usage, Goal changes, input requests, external handoffs, and cancellation. Pagination preserves long sessions without implying that the initially loaded page is complete history.
## Fields, states, or lifecycle rules
- Chronology event types retain their identity and ordering.
- Trajectory segments link only to observable activity.
- Evaluation ledger rows preserve Case counts, completion, receipts, and hashes.
- Worker packages and sessions have independent lifecycle states.
## Safe trajectories
Trajectory segments narrate observable work and link to durable activity references. They may explain that a candidate was prepared, evaluated, retained, or rejected. They exclude hidden reasoning and private worker state. Do not rewrite silence between returned artifacts as a detailed external-worker trajectory.
Benchmark Evaluations does not currently expose Traces / Spans; Improve trajectories are a separate safe session narration surface. They should not be described as raw model reasoning, execution spans, or evaluator authority.
## Evaluation ledger and receipts
The evaluation ledger records exact Case counts, evaluable and incomplete populations, scores, constraint results, and receipt or content hashes. A receipt links a candidate claim to the canonical Benchmark Evaluation that supports it. When the cohort is incomplete, preserve that state in frontier and completion decisions.
Usage entries can identify provider use by model. Missing usage evidence is unknown, not zero. Harness output records candidate-visible results without granting access to private runtime reasoning.
## External handoffs
An external worker package is prepared only after a Goal Contract is confirmed. Its status can be prepared, submitted, expired, or closed. The return contract accepts an immutable Harness Version or a supported evaluation request. Expiration or closure describes the package lifecycle, not whether private work occurred.
## Session lifecycle
Sessions can be draft, active, paused, completing, completed, cancelling, cancelled, or failed. Attention states identify an input or review need. Pausing stops new scheduling while retaining chronology and evidence. Cancelling records a terminal path; it does not erase candidates, receipts, or usage already recorded.
> Observable-state rule
>
> Chronology may report only durable messages, events, returned artifacts, requests, and receipts. Never infer hidden chain of thought or fabricate progress for an external worker.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Start an Improvement Session](/docs/improve/start-improvement-session)
- [Choose Work or Evolve](/docs/improve/work-and-evolve)
{% /related-card-grid %}
## Source confidence
Code-backed: chronology rendering, the session contract, and archive model define event types, safe trajectories, ledger and receipt fields, handoff states, usage, pagination, and lifecycle controls.