Teammately Docs
Docs menu

error

Weak Applicability Logic

Fix standards that are applied to the wrong cases or skipped where they matter.

Weak Applicability Logic

Use this when policies or rubrics fire on irrelevant cases or miss cases where they should apply.

Symptom

Benchmark results show failures that reviewers consider irrelevant, or important cases skip the standards that should govern them. The issue appears as false positives, false negatives, or confusing policy-level result counts.

Likely causes

  • The applicability condition uses a broad keyword or metadata field that does not prove the behavior is in scope.
  • Cases lack the metadata or source context the applicability rule depends on.
  • A policy boundary changed but applicability was not revised.
  • Conversation context or retrieved-source state is not represented in the case.

Diagnostic checks

  • Inspect included and excluded cases side by side.
  • Identify the exact source signal the applicability rule depends on.
  • Check whether case metadata, dimensions, or context fields are missing or stale.
  • Review recent policy or rubric revisions for boundary changes.

Fix

  • Rewrite applicability around source-backed signals, not broad topic labels.
  • Add missing metadata or context before relying on the rule.
  • Create examples that should be included and excluded, then test the boundary.
  • Approve and version the corrected governed objects, then create the Benchmark boundary and Runs needed to evaluate the revised applicability.

Prevention

  • Define applicability before benchmark runs, not after reading failures.
  • Keep case metadata and dimensions current.
  • Review applicability whenever a policy or rubric version changes.
  • Keep included and excluded example Cases linked to the governed standard where the current surface supports them.

Source confidence

Code-backed: Policy and Rubric detail routes expose governed scope and linked Cases, while Benchmark Evaluation results expose which evaluator outcomes appeared for Cases. The boundary-testing method remains a human interpretation of those inspectable artifacts.

Found something unclear?

Report outdated, unsupported, or confusing docs so we can fix the source page.

Report a docs issue

Continue learning

Related docs

AI context