Skip to main content
The AI Gateway API provides model discovery, inference, embeddings, and image generation. Clients choose a request format and model while project policy constrains eligible deployments.

Models and deployments

A model describes capabilities and accepted inputs. A deployment identifies where that model is served, including its residency, precision, retention, and pricing metadata. The public catalog exposes customer-facing metadata. Authenticated discovery can reflect a project’s inference policy. An available model does not imply that every deployment or protocol is available for a particular request.

Request formats

Inference endpoints support provider-compatible formats, including OpenAI and Anthropic formats. The provider gateway forwards requests using a provider-specific path. The endpoint reference defines supported paths, parameters, and response handling. Image generation produces project uploads. The Projects API manages those stored assets. Embedding generation produces vectors; the Knowledge API stores and searches indexed vectors. A support application can request a model response directly or use model inference as part of an agent run. A gateway request and an agent run are different resources: the Execution API provides the orchestration history for the latter. A model deployment identifies where inference is served. It is distinct from an application deployment, which assigns a project release to an environment.

Policies and spending

Project inference policies restrict model and deployment choices within the environment’s policy. Stored policy and active enforcement are distinct: inspect the documented enforcement fields and limitations before relying on a constraint. Budgets and usage belong to the Billing and Usage API. Authentication and API key ownership belong to the Identity and Access API.

Get started

Generate a model response shows model discovery and a chat-completion request.

API references