> ## Documentation Index
> Fetch the complete documentation index at: https://veryfront.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate a model response

> Discover a model and send a chat-completion request through the AI Gateway.

Select an available model and generate a text response using the OpenAI-compatible chat-completion endpoint.

## Before you start

You need `curl`, `jq`, and an [API credential](/docs/cloud/authentication) authorized for inference. The examples attribute the request to `support-assistant`; replace that reference with your project.

## 1. Discover model IDs

Call `GET /ai/v1/models`:

```bash title="list-models.sh" theme={null}
curl --fail-with-body https://api.veryfront.com/ai/v1/models \
  -H "Authorization: Bearer $VERYFRONT_API_KEY" \
  -H "x-veryfront-project-reference: support-assistant" > models.json

jq -r '.data[].id' models.json
```

Choose a model that supports chat completions. The [model discovery reference](/docs/cloud/rest/apis/ai-gateway-api#model-discovery) describes capabilities and deployment metadata. Use a returned ID rather than guessing a model or version name.

## 2. Generate a response

Set `MODEL_ID`, then call `POST /ai/v1/chat/completions`:

```bash title="generate-response.sh" theme={null}
export MODEL_ID="<MODEL_ID>"

jq -n --arg model "$MODEL_ID" '{
  model: $model,
  messages: [{role: "user", content: "Explain what an API key is in one sentence."}],
  stream: false
}' | curl --fail-with-body https://api.veryfront.com/ai/v1/chat/completions \
  -H "Authorization: Bearer $VERYFRONT_API_KEY" \
  -H "x-veryfront-project-reference: support-assistant" \
  -H "Content-Type: application/json" \
  --data-binary @- > response.json
```

The model ID selects the deployment. Project policies, credential limits, and spending checks still apply to the request.

## 3. Read the answer

Inspect the response before extracting the assistant message:

```bash title="read-response.sh" theme={null}
jq '{model, choices, usage}' response.json
jq -r '.choices[0].message.content' response.json
```

For this text-only example, expect an assistant answer in the first choice. Inspect the error body if the request fails, or `finish_reason` if generation ends unexpectedly.

This is a direct inference request. To execute a configured agent and track its run history, follow [Run an agent](/docs/cloud/run-agent).

## API references

* [Create a chat completion](/docs/cloud/rest/api-reference/inference/create-a-chat-completion) and [model discovery](/docs/cloud/rest/apis/ai-gateway-api#model-discovery).
* [List models through MCP](/docs/cloud/mcp/tools/list-models). The available operations differ across interfaces.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.