Teammately Docs
Docs menu

concept

Troubleshooting

Diagnose common Teammately setup, upload, review, classification, and benchmark issues.

Troubleshooting

Use Troubleshooting when the correctness loop produces a confusing state: cases will not import, outputs are missing, reviewers cannot proceed, policies apply too broadly, coverage is unbalanced, or a benchmark result changes for reasons the team cannot yet explain.

What this area is

Start from the visible symptom and identify the owning object before changing anything. A missing response, an unmatched external output row, a stale Dimension, and an overbroad Policy can all distort evidence, but each requires a different correction.

When the cause is uncertain, collect the exact Project, Benchmark Version, saved Harness Version or imported output identity, affected Case IDs, and visible error state before editing Cases, standards, coverage, or versions.

Decision checkpoint

SymptomDiagnose firstLikely owner artifact
Cases will not importUpload format, queue state, required columnsCase import or upload queue
A managed Run has no responseRun state, Harness Version, and Case resultEvaluation Run
Imported responses are missingCase IDs and output-only Run mappingExternal output mapping
Experts cannot proceedContribution assignment, access, or source contextExpert Contributions or permissions
Policies apply too broadlyApplicability logic and rubric wordingCorrectness Governance
Coverage looks unbalancedDimensions, ontology values, and benchmark membershipCoverage Engineering
Score changed unexpectedlyCase, output, standard, benchmark, and run metadata versionsBenchmark Evaluations and Governance

Who uses it

AI engineers usually diagnose import, mapping, benchmark, and output problems. Review owners diagnose reviewer access, unclear cases, and disagreement. Product leads and accountable owners use troubleshooting notes to understand whether a failure is a model behavior issue, an artifact issue, a coverage issue, or an access issue.

Artifacts created or changed

Troubleshooting can lead to changes in Cases, imports, output mappings, Policies, applicability, Rubrics, Dimensions, current Dataset membership, Benchmark Versions, Run configuration, or Project membership. Record which object changed and why; do not manufacture an unsupported generic recovery-note object.

Preserve historical Snapshots, Runs, governed versions, Contribution attribution, and activity already recorded by the owning surfaces. Fix current state through the normal object workflow.

How to use the recovery library

Start with the observed symptom. Do not immediately change standards, rerun benchmarks, or edit cases until the failure source is clear. Missing outputs, weak applicability logic, stale dimensions, and unbalanced coverage can all make benchmark evidence look wrong, but they require different fixes.

After the immediate fix, preserve the lesson. A recovery path should make the next correctness loop stronger: clearer case context, tighter applicability, better output mapping, better benchmark coverage, or more explicit run metadata.

If the reader cannot name the symptom yet, use the task index only after identifying the blocked artifact. Planned operations belong in the operating manual; ambiguous or broken states belong here first.

Recovery proof

Symptom classThe issue is actually fixed when...Keep diagnosing if...
Case import or uploadThe affected queue or case record is complete enough for preparation work.Rows moved forward but required context or mapping is still missing.
Managed Run responseThe Run reaches a terminal state and each evaluable Case has the generated response and Rubric outcomes expected for that Run.The response is absent, the attempt failed, or the Case remains unevaluable.
Imported output mappingEvery intended external row maps to the immutable Case ID in one output-only Run.Rows are unmatched, duplicated, or joined by position.
Review blockageThe reviewer has access, source context, assignment state, and questions needed to judge.The reviewer can enter the screen but cannot make an accountable decision.
Weak applicabilityThe standard now applies to a named case boundary instead of broad intent.The same failure could pass or fail depending on reviewer interpretation.
Unexpected score changeThe team can name whether candidate behavior, case membership, standards, outputs, or metadata changed.The cause is still described as "the benchmark changed."

Common starting tasks

Worked example

Unexpected benchmark change

A score changes after a Dataset refresh even though the saved Harness Version is unchanged. The team confirms that the second Run used a new Benchmark Version with additional boundary Cases. It treats the result as evidence under an expanded benchmark, not as a candidate regression, and compares Case-level Rubric outcomes within each named boundary.

Source confidence

Code-backed: the cited routes cover the principal access, membership, Case upload, output mapping, coverage, and Policy surfaces routed from this index. Each linked troubleshooting page narrows its own claims to the current owning implementation.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context