Skip to main content
Create a small evaluation for a project agent, execute it, and inspect the measured result before expanding the dataset.

Before you start

You need curl, jq, an API credential, and a project agent you can run. The example uses triage-agent in support-assistant. The example checks whether an answer contains billing. That demonstrates one metric; choose cases and metrics that reflect your application’s requirements.

1. Create the evaluation

Call POST /projects/{project_reference}/evals to create a source definition:
create-evaluation.sh
The returned id identifies the evaluation. URL-encode it for path parameters; the jq command above performs that encoding.

2. Make the definition available to the runtime

Deploy a release containing the new evaluation and its target agent. Set ENVIRONMENT_ID to the available environment where that release is deployed. The evaluation’s source selector and runtime selector must identify the intended version and execution environment. Other target combinations are documented in the start-run reference.

3. Start an evaluation run

Call POST /projects/{project_reference}/evals/{eval_id}/runs:
run-evaluation.sh
The evaluation response uses runId, unlike the generic run-creation response’s run.run_id.

4. Inspect status and retrieve the report

Read GET /projects/{project_reference}/evals/{eval_id}/runs/{run_id}:
read-evaluation-run.sh
If the run is still pending or running, wait and read it again. For a failed execution, inspect the error before treating the result as an agent-quality failure. Once a report is available, retrieve it with GET /projects/{project_reference}/evals/{eval_id}/runs/{run_id}/report:
read-evaluation-report.sh
Check the case results and metric outcomes. A completed execution does not imply every quality check passed. Keep the same cases and relevant settings when assessing a changed agent definition.

API references