Teach the distinction behind the benchmark failure.
Use expert judgment and coverage design to direct training-data augmentation. Establish what the model needs to learn, construct examples that teach it, and assess whether the improvement transfers.
For post-training teams and AI engineers developing domain-specific model behavior
Development focusExpert-grounded training-data augmentation
What changes the decision
More examples help only if they teach the missing distinction.
Repeated failures can have different causes: the task lacks necessary context, the harness uses the model poorly, the intended criterion is ambiguous, or the weights do not reliably express the required behavior. Dataset augmentation needs a learning target before it needs volume.
Engineering practice · These are development and assessment methods, not reported customer results.
01
The benchmark failure has an unclear standard
The context
Experts agree that a response is weak but disagree about the alternative. Generating more examples at this point can multiply an unresolved preference.
Judgment to capture
Which criterion is missing, and what condition explains the disagreement?
02
The training set repeats the same easy distinction
The context
Examples vary in wording while keeping the underlying reasoning and conditions nearly identical. A large set can still omit the situations in which the model struggles.
Judgment to capture
Which ontology combinations and decision boundaries need deliberate representation?
03
The measured gain comes from familiar cases
The context
The candidate improves on material closely related to the examples used during development. The result may not establish the intended generalization.
Judgment to capture
Which independently designed cases would test the same capability under different conditions?
The people behind the standard
Define what an example should teach before constructing it.
Teammately’s agents prepare output curation, comparisons, and adaptive conversations around the failure. Experts clarify the intended distinction; agents develop the findings into proposed rubrics and coverage that can direct training-data construction.
Domain experts
The reasoning, preferences, and exceptions that distinguish acceptable behavior.
Post-training and evaluation teams
The learning objective, training format, baseline, and independent assessment design.
Coverage and improvement
Connect elicitation, representation, and post-training evidence.
Use the benchmark failure to formulate a learning objective. Expert findings establish the distinction, coverage defines where it should hold, and Weave constructs teaching material that can be evaluated through a separate assessment design.
Ontology dimensions
Target behavior, task context, evidence conditions, preference boundaries, and difficulty of the required distinction.
Combinations to exercise
A small context change reverses the preferred answer; the model handles isolated conditions but misses their interaction; a common template masks a gap.
Interpreting the case mix
Give training examples and assessment cases separate roles. Deliberate augmentation targets a learning need; the evaluation mix defines the scope of the gain.
Resolve the judgment before multiplying the examples
Agents prepare curation, comparisons, and adaptive follow-up around the failure. Experts clarify the requirement and the conditions that change it. Develop the findings into proposed rubrics and coverage.
What this makes possible
A learning target the team can explain, including where the desired answer or preference should differ.
Coverage selects ontology tuples, difficult combinations, and intended proportions. Weave constructs cases and comparisons around those distinctions, with expert criteria guiding what the examples should teach.
What this makes possible
A dataset whose additions address identifiable conditions rather than increase volume through near-duplicate wording.
Use the agreed training environment and retain the data and configuration behind the candidate. Assess it on separate cases under comparable model-serving and evaluation conditions. Trialground supplies managed trials and rubric evaluation where included.
What this makes possible
Evidence of the intended learning, preserved behavior, and remaining gaps, with the benchmark’s representation defining the scope of the conclusion.
Weave supplies training-data augmentation; training execution and compute use the agreed training environment. Coevolve provides the separate route for harness prototypes when the evidence points to an implementation change.
Assessing progress
Connect expert contribution to a change that holds beyond the examples.
The distinction experts established and how it shaped the training data.
A baseline and updated candidate assessed under comparable conditions.
Assessment cases kept separate from training and development material.
Gains and regressions across the intended coverage.
Expert participation, including review and correction.
Engineering effort, data construction, training configuration, and compute.
A benchmark gain should identify the baseline, updated score, models, harnesses, and evaluation conditions. Where practical, compare targeted augmentation with a baseline that helps isolate its contribution. Report the coverage on which learning transferred and the conditions that remain weak.
Expert time becomes meaningful when readers can see what the contribution established and how agents turned it into usable material. Report participation separately from engineering, generation, and training effort. The evidence should connect the captured judgment to the resulting improvement.
Teach how missing or conflicting evidence should change the answer or next question.
Bring a behavior that additional training data has not resolved.
We can examine whether the missing piece is the judgment, the representation, or the learning target, then define a useful augmentation and assessment scope.