Teammately’s AI agent

Meet Lemon.

Your AI agent for alignment.

Work with Lemon to develop harnesses, prepare expert contributions, and construct benchmarks. Its specialized capabilities carry that work through elicitation, synthesis, and benchmark-driven improvement, with your experts establishing judgment and your engineers deciding what to adopt.

Lemon, Teammately’s cockatiel AI agent

Working with your engineers.
Learning from your experts.

Lemon · Assistance across your workspace

Give the next piece of work to Lemon.

Develop coverage facets, explore which expert contribution would resolve an uncertainty, or investigate a conflict in the standards your team has established. Lemon prepares suggestions around the relevant project context, giving you concrete work to inspect and refine.

Included in your platform subscription Pricing details ↗

Make a starting point concrete

Project context and reference materials inform coverage-schema and topic-group suggestions. Your team can refine the dimensions that matter before directing case construction.

Find the next useful contribution

Coverage gaps and unresolved judgments can inform a focused expert contribution. Lemon proposes work around the distinction that needs clarification.

Keep work inspectable

The AI Work surface brings agent activity, findings, and requests for attention into view. Open the relevant context to inspect what happened and decide how to continue.

Lemon Elicit · Adaptive expert elicitation

Develop expert contributions into usable standards.

Lemon Elicit prepares and adapts contributions to uncover the reasoning, exceptions, and preferences behind a judgment. It learns during reviews, asks directly through chat or voice when useful, and develops findings into proposed policies, rubrics, and coverage. Your experts contribute their judgment while Lemon undertakes the preparation and interpretation needed to put it to use.

Included in your platform subscription Pricing details ↗

Learn through the right comparison

Output curation, pairwise and listwise comparisons, and trajectory reviews make a distinction concrete. Lemon Elicit plans the contribution around what needs to be learned and adapts as the expert responds.

Ask when the judgment needs context

Forms, chat, and live voice interviews let experts explain which evidence, condition, or tradeoff changes their decision. Direct questioning is part of a continuing elicitation process.

Return an interpretation for review

Findings become proposed policies, benchmark rubrics, and coverage observations. Experts inspect whether those proposals preserve their intent; unresolved boundaries guide further work.

Lemon ElicitInteractive product preview
Loading Elicitation…
The expert perspectiveYour experts shape what correct means.

Explore participation, review, and the contribution to your correctness loop.

Lemon Weave · Cases, worlds, and training material

Construct what your coverage calls for.

Lemon Weave synthesizes cases and worlds from ontology tuples, following coverage designs informed by product managers’ objectives and experts’ experience. It constructs the conditions and supporting material that make a task meaningful. The same foundation guides targeted dataset augmentation for weights training.

Metered usage Pricing details ↗

Cases with a reason to exist

Coverage Stories and construction patterns give each case a representation need. Selected ontology tuples define the conditions that must occur together, including difficult combinations and consequential exceptions.

Worlds an agent can act within

A world can supply connected information, tools, state, and consequences within the configured environment. Supporting documents, workbooks, code, and other artifacts carry the context the task requires.

Teach a missing distinction

Expert criteria guide supervised demonstrations, preference pairs, and other training material. Target the distinctions exposed by failures, then assess the trained candidate on cases kept separate from that teaching material.

Lemon WeaveInteractive product preview
Loading Weave…
Construction, training material, and scope

Cases and comparisons

Make a missing distinction observable.

Lemon Weave turns coverage needs into candidate cases and controlled alternatives. Construction patterns shape how the selected conditions appear in the task. Comparisons can vary the evidence, constraint, or response choice an expert needs to examine, while retaining enough common context to explain why the judgment changes.

Supporting materials

Give the task the evidence it actually requires.

A task may depend on a document, image, structured record, or source file. Lemon Weave constructs supporting materials within the configured artifact scope and reference context. Their role is to carry the evidence needed for the case, including relevant uncertainty or disagreement between sources.

Executable worlds

Test decisions that unfold through action.

World synthesis supplies the materials, initial state, and scoped environment behavior an agent task needs. Tools and state changes give a trajectory consequences the benchmark can examine. This extends coverage to decisions whose quality depends on what happens between the initial request and the final response.

Training datasets

Teach the distinction behind a repeated failure.

Lemon Weave augments training datasets around missing distinctions and underrepresented conditions. Expert rubrics guide what the examples should teach; ontology tuples guide their variation. Your team can target a learning need revealed by the benchmark and assess the trained candidate on separate cases.

Lemon Weave constructs the cases, supporting artifacts, worlds, and training examples. Trialground supplies managed harness execution and rubric evaluation. Weights-training execution and compute are scoped separately for your training environment.

Lemon Code · Interactive harness development

Develop a harness through a working conversation.

Work with Lemon Code on the code, context, and instructions around your model. Bring the relevant harness, policies, and reference material into the conversation. Inspect proposed changes in the editor, refine the approach, and decide what to apply to your draft.

Metered usage Pricing details ↗

Work from the implementation

A coding conversation stays connected to its harness draft and selected context. Lemon Code can help explain existing behavior and develop changes across the files that implement it.

Inspect the proposed change

Review file changes and their rationale before accepting them. Continue the conversation to refine the implementation; the proposal remains connected to the draft it was prepared against.

Take the candidate into evaluation

A plausible edit needs behavioral evidence. Use Trialground to evaluate the candidate against your rubrics and coverage, or give Lemon Goal a benchmark objective for further optimization.

Lemon CodeInteractive product preview
Loading Code…

Lemon Goal · Benchmark-driven harness optimization

Give Lemon the benchmark objective.

Lemon Goal investigates failures, develops competing harness prototypes in a sandbox, and tests them through Trialground. It pursues the behavior you want to improve while tracking protected requirements, permitted changes, and the search budget. Engineers receive tested candidates and the evidence behind them.

Metered usage Pricing details ↗

Define useful progress

The goal connects a starting candidate and benchmark to the outcomes that should improve. Protected behavior and permitted changes keep the search directed toward an acceptable result.

Investigate competing explanations

Lemon Goal examines case evidence, tests changes to prompts, retrieval, tools, or orchestration, and coordinates subagents where useful. Experiments inform which approach to retain or investigate next.

Keep the evidence with the candidate

Compare gains, regressions, and rejected approaches. Independent confirmation and preserved requirements help establish whether an apparent gain holds. Your team decides which ideas to integrate and deploy.

Lemon GoalInteractive product preview
Loading Improve…
Search, confirmation, and engineering handoff

The benchmark goal

Make the desired improvement and its constraints explicit.

A Goal Contract connects the starting candidate and benchmark evidence to the outcomes the agent should pursue. Protected behavior, permitted changes, and the search budget define useful progress. This gives Lemon Goal a continuing objective as it moves between investigation, implementation, and evaluation.

Investigation

Turn observed failures into testable explanations.

Lemon Goal examines case evidence and the current harness to identify possible causes. It can investigate prompts, retrieval, tool use, and orchestration within the permitted scope. A proposed intervention connects an observed failure to a change the next trial can support or rule out.

Code and Goal

Follow a focused intervention or explore competing approaches.

Lemon Code supports interactive harness development. Lemon Goal pursues a benchmark objective through competing candidate approaches, with experiments informing what to retain or investigate next. Lemon Goal uses Trialground to evaluate prototypes and compare the gains, regressions, and tradeoffs they produce.

Confirmation

Establish what supports the candidate you retain.

A plausible edit or promising preliminary score is a reason to investigate further. Lemon Goal’s search uses comparable evaluation evidence, protected requirements, and independent confirmation when establishing an improvement. The retained prototype comes with the evidence and limitations behind that decision.

Lemon Goal produces tested harness prototypes and experimental evidence. Your team owns production integration and deployment. Lemon Weave’s training-data augmentation provides the separate path for teaching new distinctions to weights.

Scope and pricing

Included assistance. Capacity for the work you choose.

General Lemon assistance and Lemon Elicit are included in your platform subscription. Lemon Weave, Code, and Goal use metered capacity. Trialground execution and rubric evaluation have their own usage measures.

Expert work stays included

Preparation, interviews, review learning, and proposed standards are part of Elicit. If a contribution calls for newly synthesized Weave cases, that paid work is identified separately before generation.

Evaluation remains attributable

Trialground’s rubric evaluators apply configured criteria to recorded behavior. Inspect the rubric, evaluation configuration, and supporting evidence behind a result. Lemon can use those findings to direct further work.

Choose a useful first objective

Begin with a harness, a coverage need, or an expert question. Agree on the required context, participation, integrations, and capacity for that work. A finished benchmark is not required to start.

Start from the work you have

Bring Lemon a useful first objective.

Tell us what you are building, what you have today, and where progress gets difficult.

Talk to our team ↗