Teammately models

Kestel and Bower.

Specialized models for evaluation and synthesis.

We fine-tune open-weight models for the judgments and cases that guide AI improvement. Kestel applies binary rubrics; Bower constructs cases around planned coverage. We develop them to support fair assessment and useful development material across model families.

Choose Kestel for rubric evaluation or Bower for case synthesis, alongside other models supported by Teammately.

Binary rubric evaluation

Kestel

Trained to apply the criterion.

Explore Trialground Evaluation

Assess the requirement behind the score.

A fluent response can still miss an expert requirement. Kestel is fine-tuned to judge candidate behavior against binary rubrics: whether an applicable criterion is met in the stated context. Experts establish the standard; the model’s task is to apply it.

This supports comparisons across harness and weights changes. The relevant question is whether the candidate satisfies the same requirement, even when its language, structure, or approach differs.

When to consider Kestel

Consider Kestel when your evaluation depends on explicit binary criteria and comparisons across candidate models. Check its agreement with your experts on the requirements and failure modes that matter to your work. It is one of the judge models you can choose in Trialground Evaluation.

Coverage-directed case synthesis

Bower

Trained to construct the intended case.

Explore Lemon Weave

Make the conditions that matter occur together.

A plausible example can still omit the distinction it was meant to test. Bower is fine-tuned for case synthesis guided by ontology tuples and expert criteria. Coverage defines which combinations and proportions matter; the synthesis model constructs cases to serve that design.

Within Lemon Weave, this material supports benchmark construction, expert elicitation, and targeted training. Cases can expose a boundary for expert review or supply examples of a distinction the weights need to learn.

When to consider Bower

Consider Bower when you need constructed cases to reflect specific conditions, difficult combinations, or underrepresented situations. Examine whether the cases fulfill the intended tuples and remain coherent and useful for their purpose. You can select Bower alongside other synthesis models supported by Lemon Weave.

Model choice

Choose each model for the work it must do.

Your target model, rubric judge, and synthesis model have different responsibilities. Teammately lets you choose from the models supported for each task. Kestel and Bower provide specialized options for assessment and construction separately from the candidate you are developing.

Match the specialization to your needs.

Use the criteria, inputs, model families, and coverage conditions from your own work to assess fit. We develop Kestel and Bower for these focused responsibilities; your results and requirements should guide the selection.

Keep the comparison interpretable.

When comparing candidates, preserve the benchmark and assessment settings. When comparing judges or synthesis models, examine the differences they introduce before changing the basis on which you measure improvement.

Assessment approach

Fair assessment has to be demonstrated.

Fairness across model families is a development objective. Selecting a separate judge or generator does not by itself establish independence: models can share preferences and blind spots. The assessment needs to examine how those influences affect judgments, case construction, and subsequent improvement.

For rubric judgments

Compare agreement with expert decisions across unfamiliar rubrics and outputs from different model families. Examine whether judgments change with wording or presentation when the criterion-relevant behavior stays the same. Inspect errors by criterion and candidate family, alongside aggregate agreement.

For constructed cases

Check fulfillment of the intended ontology tuples, consistency, and variation within each coverage need. When using the material for development, assess the resulting candidate on cases kept separate from training and optimization. This helps distinguish useful learning from familiarity with a generator’s patterns.

Put the models to work

Start with your criteria and coverage.

Bring the behavior you need to assess or the cases you need to construct. We can discuss where Kestel or Bower fits, the other supported model options, and the evidence your team needs to make a choice.