> ## Documentation Index
> Fetch the complete documentation index at: https://veryfront.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluate an agent

> Create an evaluation case, run it against an agent, and inspect its report.

Create a small evaluation for a project agent, execute it, and inspect the measured result before expanding the dataset.

## Before you start

You need `curl`, `jq`, an [API credential](/docs/cloud/authentication), and a project agent you can run. The example uses `triage-agent` in `support-assistant`.

The example checks whether an answer contains `billing`. That demonstrates one metric; choose cases and metrics that reflect your application's requirements.

## 1. Create the evaluation

Call `POST /projects/{project_reference}/evals` to create a source definition:

```bash title="create-evaluation.sh" theme={null}
curl --fail-with-body https://api.veryfront.com/projects/support-assistant/evals \
  -H "Authorization: Bearer $VERYFRONT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "kind": "eval-source-create",
    "fields": {
      "name": "Triage quality",
      "target": "agent:triage-agent",
      "dataset": {
        "kind": "inline",
        "editable": true,
        "dynamic": false,
        "examples": [{
          "id": "duplicate-charge",
          "input": "I was charged twice. Classify this request as billing.",
          "reference": "billing"
        }]
      },
      "metrics": [{
        "name": "answer.contains",
        "family": "answer",
        "severity": "gate",
        "config": {"text":"billing"},
        "editable": true,
        "dynamic": false
      }],
      "repetitions": 1
    }
  }' > evaluation.json

export EVAL_REFERENCE=$(jq -er '.id | @uri' evaluation.json)
```

The returned `id` identifies the evaluation. URL-encode it for path parameters; the `jq` command above performs that encoding.

## 2. Make the definition available to the runtime

[Deploy a release](/docs/cloud/deploy) containing the new evaluation and its target agent. Set `ENVIRONMENT_ID` to the available environment where that release is deployed.

The evaluation's source selector and runtime selector must identify the intended version and execution environment. Other target combinations are documented in the start-run reference.

## 3. Start an evaluation run

Call `POST /projects/{project_reference}/evals/{eval_id}/runs`:

```bash title="run-evaluation.sh" theme={null}
export ENVIRONMENT_ID="<ENVIRONMENT_ID>"

curl --fail-with-body \
  "https://api.veryfront.com/projects/support-assistant/evals/$EVAL_REFERENCE/runs?source_target_kind=environment&runtime_target_kind=environment&target_environment_id=$ENVIRONMENT_ID" \
  -H "Authorization: Bearer $VERYFRONT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{}' > evaluation-run.json

export EVAL_RUN_ID=$(jq -er '.runId' evaluation-run.json)
```

The evaluation response uses `runId`, unlike the generic run-creation response's `run.run_id`.

## 4. Inspect status and retrieve the report

Read `GET /projects/{project_reference}/evals/{eval_id}/runs/{run_id}`:

```bash title="read-evaluation-run.sh" theme={null}
curl --fail-with-body \
  "https://api.veryfront.com/projects/support-assistant/evals/$EVAL_REFERENCE/runs/$EVAL_RUN_ID" \
  -H "Authorization: Bearer $VERYFRONT_API_KEY" > evaluation-result.json

jq '{status, summary, error}' evaluation-result.json
```

If the run is still pending or running, wait and read it again. For a failed execution, inspect the error before treating the result as an agent-quality failure.

Once a report is available, retrieve it with `GET /projects/{project_reference}/evals/{eval_id}/runs/{run_id}/report`:

```bash title="read-evaluation-report.sh" theme={null}
curl --fail-with-body \
  "https://api.veryfront.com/projects/support-assistant/evals/$EVAL_REFERENCE/runs/$EVAL_RUN_ID/report" \
  -H "Authorization: Bearer $VERYFRONT_API_KEY" > report.json

jq . report.json
```

Check the case results and metric outcomes. A completed execution does not imply every quality check passed. Keep the same cases and relevant settings when assessing a changed agent definition.

## API references

* [Evaluation definitions and runs](/docs/cloud/rest/apis/evaluations-api) lists source selectors, execution targets, and report operations.
* [Evaluation GraphQL operations](/docs/cloud/graphql/apis/evaluations) and [MCP tools](/docs/cloud/mcp/apis/evaluations) provide the supported equivalents.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.