# Versioning and Staleness
Generated: 2026-09-13T04:42:55.326Z
Source build: local
Canonical docs: https://teammately.ai/docs
---
id: governance.versioning-staleness
title: Versioning and Staleness
summary: Know when correctness objects changed and when old evidence may need review.
kind: concept
product_area: governance
status: stable
updated: 2026-09-07
canonical: /docs/governance/versioning-and-staleness
---
# Versioning and Staleness
## Definition
Versioning preserves the exact artifact state used by an earlier decision or evaluation. Staleness is the signal that current evidence, configuration, or interpretation may no longer support the same claim after a related object changes. A stale signal routes review; it does not automatically delete an artifact, invalidate every historical result, or approve a replacement.
## Why it matters
Cases, Policies, Rubrics, Benchmark Versions, Harness Versions, and Improvement Session goals can change independently. Named versions keep old evidence interpretable. Staleness helps teams decide which current Datasets, evaluator links, Runs, or customer-owned human review context need attention before being treated as current.
## Where it appears in the product
Use the owning object page to inspect its current version and activity. Use Dataset Snapshots and Benchmark Versioning for immutable evaluation boundaries. Use Staleness Detection to identify downstream artifacts affected by a change. Use Conflict Resolution when expert or source evidence disagrees about what the new governed state should be.
## Artifacts it affects
Common triggers include changed Case input or materials, revised Policy scope, revised Rubric criteria, changed dataset membership, changed Harness configuration, and superseded source material. The responsible next action depends on the owner: correct a Case, approve a new standard version, refresh coverage, create a new Snapshot, run a new evaluation, or preserve an old result as historical context.
Comparison Directions use a narrower stale signal. **Potentially stale** is an advisory label for untouched AI-suggested directions, not a versioned approval state and not automatic removal. User-created and user-edited directions remain user-owned even when Teammately considers them while avoiding duplicate suggestions.
## Operational check
Name the changed artifact and version, identify which downstream claim depended on it, and decide whether the old evidence remains historically valid, requires qualification, or needs replacement through a new canonical workflow. Never “resolve” staleness by editing a label while leaving the evidence boundary ambiguous.
When the object is a Comparison Direction, also check whether a **Potentially stale** label is only advisory. Dismiss the label if the team decides the direction still represents a useful boundary.
{% example-demo title="Revised applicability after evaluation" %}
Experts revise a Policy so it applies only when the customer explicitly requests a recommendation. Runs against the old Benchmark Version remain valid evidence under the former applicability rule. The current Dataset and linked Rubrics are reviewed, a new Snapshot and Benchmark Version establish the revised boundary, and new Runs use it. Any customer-owned human review context names both boundaries instead of marking every old result simply “wrong.”
{% /example-demo %}
## Related workflows
{% related-card-grid title="Related workflows" %}
- [Versions, staleness, and resolution](/docs/object-model/versions-staleness-and-resolution)
- [Staleness Detection](/docs/governance/staleness-detection)
- [Compare Harness Versions](/docs/benchmark-evaluations/compare)
- [Product quickstart](/docs/quickstart)
- [Task index](/docs/operating-manual/task-index)
{% /related-card-grid %}
## Source confidence
Code-backed: current Policy and Rubric types preserve versioned governance facts, Snapshot routes preserve immutable benchmark evidence, and Comparison Directions expose a deliberately narrower advisory stale label. Cross-object staleness remains a review and routing decision, not an automatic global state transition.
---
id: object-model.versions-staleness-resolution
title: Versions, staleness, and resolution
summary: Track how correctness objects evolve and how teams resolve conflicting evidence.
kind: reference
product_area: object_model
status: stable
updated: 2026-08-23
canonical: /docs/object-model/versions-staleness-and-resolution
---
# Versions, staleness, and resolution
## Definition
Versions, staleness, and resolution describe how correctness artifacts evolve without making old evidence ambiguous. Cases, Policies, Rubrics, Benchmark Versions, supported Case-scoped reference outputs, and customer-owned review context can change at different times; version boundaries explain which evidence belongs to which state.
Use this reference when a result changed unexpectedly, a policy was revised, a case was refreshed, or reviewers need to know whether older benchmark evidence still applies.
## Fields, states, or lifecycle rules
- Versions preserve what changed and what evidence was produced before the change.
- Staleness means older evidence may no longer reflect the current case, standard, coverage, or candidate boundary.
- Resolution work should name whether the fix belongs to a case, output, policy, rubric, coverage plan, benchmark version, or run metadata.
- Comparisons are weak when artifact versions are hidden.
- This page describes public object semantics, not retention, audit-log completeness, or compliance guarantees.
## Related objects
Versions, staleness, and resolution should be read with [Versioning and Staleness](/docs/governance/versioning-and-staleness), [Benchmark versioning](/docs/governance/benchmark-versioning), [Case versioning](/docs/governance/case-versioning), and [Policy Conflicts and Revisions](/docs/governance/conflict-resolution).
{% example-demo title="Versions, staleness, and resolution boundary" %}
State change: Reviewers revise a compatibility policy after finding unsupported-claim failures.
Benchmark evidence: Runs against the old policy remain interpretable, but they should not be summarized as current evidence without naming the old policy version.
Interpretation: The resolution note explains whether to rerun, revise the benchmark version, or preserve the old result as historical context.
{% /example-demo %}
## Source confidence
Code-backed: Benchmark, Policy, and Rubric types carry version facts; Benchmark Datasets → Snapshots and Policy activity preserve named historical boundaries. Cross-object staleness and conflict resolution are explicit review decisions rather than a universal automatic state.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Versioning and Staleness](/docs/governance/versioning-and-staleness)
- [Conflict Resolution](/docs/governance/conflict-resolution)
- [Product quickstart](/docs/quickstart)
- [Task index](/docs/operating-manual/task-index)
{% /related-card-grid %}
---
id: governance.staleness-detection
title: Detect and route stale correctness evidence
summary: Identify which current claims need review after Cases, standards, coverage, sources, or target behavior change.
kind: task
product_area: governance
status: stable
updated: 2026-09-07
canonical: /docs/governance/staleness-detection
---
# Detect and route stale correctness evidence
## What staleness means
Staleness means a downstream claim may no longer be supported by the current upstream state. It is a review signal, not automatic deletion and not a claim that historical evidence was invalid when produced.
Teammately objects change independently. A revised Policy may affect linked Rubrics and current benchmark interpretation without changing the exact evidence captured in an earlier Snapshot. A changed Harness may require a new Run while leaving the Benchmark Version unchanged.
## Common triggers
- A Case input, context, source attachment, or supported reference output changes.
- A Policy rule, scope, approval, or applicable boundary changes.
- A Rubric criterion or link changes.
- Coverage facets or selected Case membership change.
- A source becomes superseded or the target product changes.
- A saved Harness version changes before candidate comparison.
## Prerequisites
- The changed object and its earlier and current versions can be identified.
- The team can trace current downstream artifacts that rely on the changed fact.
- An owner can decide whether current evidence needs qualification, replacement, or no action.
### Task steps: Assess and route staleness
1. Name the changed object, its earlier and current versions, and the reason for change.
2. Identify downstream artifacts that rely on the changed fact: linked standards, selected Cases, Snapshots, Runs, comparisons, or customer-owned human review context.
3. Classify each artifact as historically valid, current and unaffected, current but requiring qualification, or requiring replacement.
4. Route the correction to the owning workflow: edit a Case, govern a new Policy or Rubric version, refresh coverage, create a Snapshot, or run a saved Harness again.
5. Preserve the old version and its evidence. Add a note that states which boundary the evidence still supports.
6. Confirm that current navigation and handoff material point to the new canonical version.
## Object and state changes
A changed object does not make every connected artifact unusable. If a Policy wording change does not affect a particular Case or Rubric, document that determination. If a candidate configuration changes, create a new Run rather than a new Benchmark Version. If selected membership changes, create a new Snapshot rather than editing an old one.
Comparison Directions have a narrower advisory state. **Potentially stale** applies to untouched AI-suggested directions considered inconsistent with newer project context. It does not automatically apply to user-created or user-edited directions, and dismissing the label does not delete the direction or govern any downstream artifact.
## Success criteria
- The trigger and affected version boundaries are named.
- Historical evidence remains interpretable under its original boundary.
- Current artifacts are explicitly unaffected, qualified, or routed to the owning workflow.
- New Snapshots or Runs are created only when their respective evidence boundary changed.
## Common failure modes
- Treating every connected artifact as invalid after one upstream change.
- Editing a historical Snapshot or Run to resemble current state.
- Creating a new Benchmark Version when only the Harness changed.
- Confusing the advisory Comparison Direction label with governed-object staleness.
{% example-demo title="Source document superseded" %}
A new service policy supersedes the source used by twelve refund Cases. The operator preserves the old Snapshot and its Runs, updates affected current Cases, checks the linked Policy and Rubrics, and creates a new Snapshot. Two Cases describe historical behavior and remain unchanged with an explicit time boundary; ten move to the current version.
{% /example-demo %}
## Source confidence
Code-backed: Policy and Rubric types, Benchmark Version and Snapshot surfaces, and Run detail preserve the version boundaries needed for this assessment. Comparison Directions explicitly expose their narrower advisory stale label. Cross-object dependency assessment and the decision to rerun or revise remain owner-reviewed work.
## Related reference pages
{% related-card-grid title="Related workflows" %}
- [Versioning and Staleness](/docs/governance/versioning-and-staleness)
- [Coverage Refresh](/docs/coverage-engineering/coverage-refresh)
- [Benchmark Versioning](/docs/governance/benchmark-versioning)
- [Comparison Directions](/docs/assets/comparison-directions)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Diagnose stale evidence" %}
- [Stale Dimensions](/docs/troubleshooting/stale-dimensions)
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
- [Weak applicability logic](/docs/troubleshooting/weak-applicability-logic)
{% /related-card-grid %}
---
id: assets.comparison-directions
title: Comparison Directions
summary: Create reusable guidance for meaningful candidate-output differences in comparative expert work.
kind: reference
product_area: assets
status: stable
updated: 2026-09-07
canonical: /docs/assets/comparison-directions
---
# Comparison Directions
## Definition
A Comparison Direction is a reusable project Asset that describes how candidate outputs should differ during comparative expert work. It can focus attention on a meaningful contrast such as evidence grounding, uncertainty handling, or response strategy without declaring which candidate is correct.
Comparison Directions are project-scoped. A Contribution can select or allow a pool of directions for its comparative component, while the benchmark Contribution still owns the objective, cases, candidates, and expert task.
## Fields, states, or lifecycle rules
- A direction has a name or label and a description of the intended contrast.
- Users can create, edit, pin, archive, and remove directions from **Assets → Comparison Directions**.
- Pinned directions are surfaced when a Contribution request selects comparative output guidance.
- The suggestion experience creates draft candidates in a suggestion run. Nothing enters the reusable library until a user accepts it.
- A direction can carry a staleness advisory when its source context has changed. Dismissing that advisory records a review decision; it does not approve a Policy, Rubric, Case, or Benchmark.
- A direction guides comparative presentation or generation. It does not create a Case, change coverage structure, or replace expert judgment.
## Correct scope
Use Dimensions, Project Topics, and Case Construction Patterns for the behavior space a benchmark should represent. Use Comparison Directions for how candidate outputs should be contrasted within a comparative Contribution. Use Correctness Governance for the approved standard that determines how an output is judged.
{% example-demo title="Example: source-grounding contrast" %}
A project creates one Comparison Direction asking for a response that cites the current source conservatively and another asking for a focused clarification when the source hierarchy is unresolved. A comparative Contribution can use those directions to elicit an expert preference. The direction does not approve either response or create the governing rubric.
{% /example-demo %}
## Source confidence
Code-backed: the active Assets routes expose the Comparison Directions library, detail controls, suggestion runs, accept or dismiss decisions, pinning, and staleness review. The API keeps legacy compatibility names internally, but this page uses the current product label.
## Related task pages
{% related-card-grid title="Related task pages" %}
- [Assets](/docs/assets)
- [Request an Expert Contribution](/docs/expert-contributions/request-contribution)
- [Manage benchmark coverage](/docs/coverage-management)
{% /related-card-grid %}
---
id: benchmark-evaluations.compare
title: Compare Harness Versions
summary: Compare two or more saved Harness Versions in a symmetric evidence matrix across Cases, evaluators, and Coverage Facets.
kind: task
product_area: benchmark_evaluations
status: stable
updated: 2026-09-13
canonical: /docs/benchmark-evaluations/compare
---
# Compare Harness Versions
Compare shows aggregate observed evidence for saved Harness Versions. Choose the Versions, compatible evaluation configuration, and measurement to compare. Use **Inspect individual Run comparisons** for the detailed result matrix.
## Prerequisites
Saved Harness Versions with evaluation results are shown for the selected Benchmark Version. Choose one or more Versions to inspect; Run counts may differ. An output-only imported reference Run cannot become a Harness column because it has no executable saved Version.
## Select a row mode
The individual Run matrix offers row modes including Cases, all results, Policies, Rubrics, Dimension ontology values, Project Topics, Topic Groups, and Case Construction Patterns. Use Cases to inspect concrete disagreement, Policies or Rubrics to locate correctness movement, and Coverage Facets to see whether gains concentrate in one behavior slice.
## Steps
1. Confirm the immutable Benchmark Version and choose at least two visible Harness Versions.
2. Select an aggregate measurement, or open individual Run comparisons and choose a row mode.
3. Check evidence completeness for each Harness column. A blank or incomplete cell is not a failure.
4. Locate rows with material disagreement and connect them back to Case and evaluator evidence.
5. Preserve regressions and required-criterion failures next to gains.
6. Use the exact candidate and row evidence when starting an Improvement Session or requesting an Expert Contribution.
Compare reads existing evidence and does not mutate Runs. Selecting aggregate Harness Versions recomputes their comparison on compatible evidence; individual matrix visibility is local presentation. It does not activate Harnesses or choose a winner.
{% example-demo title="Example: facet-local improvement" %}
Three Harness Versions look similar overall. The Project Topic row mode shows that Version 14 improves source-authority Topics but regresses escalation Topics. Switching to Cases identifies two regressions, and the team starts Improve with those exact failures instead of claiming a uniform improvement.
{% /example-demo %}
## Common mistakes
- Comparing different Benchmark Versions as though only the candidate moved.
- Treating missing evidence as a failed cell.
- Reading a facet aggregate without checking the distinct Cases behind it.
- Describing an imported output-only Run as a Harness Version.
- Selecting the newest Version solely because it is newest.
## Object and state changes
Compare reads existing evidence. Selecting aggregate Versions reads their observed evidence; selecting individual matrix row modes changes presentation; it does not activate a Harness, mutate a Run, or retain a candidate.
## Success criteria
- At least two exact Harness Versions share the same Benchmark Version.
- Incomplete cells remain distinct from failed evidence.
- Material movement resolves to Cases, evaluators, or Coverage Facets.
## Common failure modes
- Comparing moved evidence boundaries as candidate-only change.
- Treating local column visibility as product configuration.
- Using an output-only Run as a Harness column.
## Related reference pages
{% related-card-grid title="Related reference pages" %}
- [Benchmark Evaluations](/docs/benchmark-evaluations)
- [Arena and Rankings](/docs/benchmark-evaluations/arena-and-rankings)
- [Harnesses](/docs/assets/harnesses)
{% /related-card-grid %}
## Related troubleshooting pages
{% related-card-grid title="Related troubleshooting pages" %}
- [Benchmark results changed unexpectedly](/docs/troubleshooting/benchmark-results-changed-unexpectedly)
- [Unbalanced coverage](/docs/troubleshooting/unbalanced-coverage)
{% /related-card-grid %}
## Source confidence
Code-backed: the active Compare route and Evaluation Matrix define Harness columns, local visibility, symmetric comparison, and the current evidence row modes.