Before you start
You needcurl, jq, an API credential, and a project agent you can run. The example uses triage-agent in support-assistant.
The example checks whether an answer contains billing. That demonstrates one metric; choose cases and metrics that reflect your application’s requirements.
1. Create the evaluation
CallPOST /projects/{project_reference}/evals to create a source definition:
create-evaluation.sh
id identifies the evaluation. URL-encode it for path parameters; the jq command above performs that encoding.
2. Make the definition available to the runtime
Deploy a release containing the new evaluation and its target agent. SetENVIRONMENT_ID to the available environment where that release is deployed.
The evaluation’s source selector and runtime selector must identify the intended version and execution environment. Other target combinations are documented in the start-run reference.
3. Start an evaluation run
CallPOST /projects/{project_reference}/evals/{eval_id}/runs:
run-evaluation.sh
runId, unlike the generic run-creation response’s run.run_id.
4. Inspect status and retrieve the report
ReadGET /projects/{project_reference}/evals/{eval_id}/runs/{run_id}:
read-evaluation-run.sh
GET /projects/{project_reference}/evals/{eval_id}/runs/{run_id}/report:
read-evaluation-report.sh
API references
- Evaluation definitions and runs lists source selectors, execution targets, and report operations.
- Evaluation GraphQL operations and MCP tools provide the supported equivalents.