Skip to main content
The Evaluations API tests agents against defined examples and metrics. It keeps the evaluation definition separate from each execution and its results.

Evaluation definitions

An evaluation definition identifies a target agent, a dataset, metrics, and repetitions. The dataset supplies the cases to evaluate. Metrics specify what to measure for each case or across the run. Definitions are stored in project source. The target agent is managed through the Agents API.

Runs and reports

Starting an evaluation creates a new run. The run records progress and results without replacing the definition or the history of earlier runs. An evaluation run produces a report. Check the run status before retrieving it; creating the run does not make the report immediately available. The Execution API provides shared run operations, including event inspection and cancellation. Evaluation-specific operations provide the definition, evaluation run details, and report.

Assessing a change

For a support agent, a dataset can represent expected questions and known failure cases. Metrics define the evidence collected from those cases; repetitions can help expose variation across executions. When assessing a changed instruction or tool configuration, keep the cases and relevant settings consistent enough for the results to answer the question you are testing. A completed evaluation means execution finished; its report and your acceptance criteria determine whether behavior is acceptable. Your release process decides how evaluation results inform approval. The evaluation API does not make every successful run a deployment approval.

Get started

Evaluate an agent walks through creating a definition, starting a run, and inspecting its report. The Code evaluation guide covers source-level authoring.

API references