Skip to main content
Select an available model and generate a text response using the OpenAI-compatible chat-completion endpoint.

Before you start

You need curl, jq, and an API credential authorized for inference. The examples attribute the request to support-assistant; replace that reference with your project.

1. Discover model IDs

Call GET /ai/v1/models:
list-models.sh
Choose a model that supports chat completions. The model discovery reference describes capabilities and deployment metadata. Use a returned ID rather than guessing a model or version name.

2. Generate a response

Set MODEL_ID, then call POST /ai/v1/chat/completions:
generate-response.sh
The model ID selects the deployment. Project policies, credential limits, and spending checks still apply to the request.

3. Read the answer

Inspect the response before extracting the assistant message:
read-response.sh
For this text-only example, expect an assistant answer in the first choice. Inspect the error body if the request fails, or finish_reason if generation ends unexpectedly. This is a direct inference request. To execute a configured agent and track its run history, follow Run an agent.

API references