> ## Documentation Index
> Fetch the complete documentation index at: https://veryfront.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# AI Gateway API

> How model discovery, inference formats, and project policies affect AI requests.

The AI Gateway API provides model discovery, inference, embeddings, and image generation. Clients choose a request format and model while project policy constrains eligible deployments.

## Models and deployments

A model describes capabilities and accepted inputs. A deployment identifies where that model is served, including its residency, precision, retention, and pricing metadata.

The public catalog exposes customer-facing metadata. Authenticated discovery can reflect a project's inference policy. An available model does not imply that every deployment or protocol is available for a particular request.

## Request formats

Inference endpoints support provider-compatible formats, including OpenAI and Anthropic formats. The provider gateway forwards requests using a provider-specific path. The endpoint reference defines supported paths, parameters, and response handling.

Image generation produces project uploads. The [Projects API](/docs/cloud/apis/projects) manages those stored assets. Embedding generation produces vectors; the [Knowledge API](/docs/cloud/apis/knowledge) stores and searches indexed vectors.

A support application can request a model response directly or use model inference as part of an agent run. A gateway request and an agent run are different resources: the [Execution API](/docs/cloud/apis/execution) provides the orchestration history for the latter.

A model deployment identifies where inference is served. It is distinct from an application deployment, which assigns a project release to an environment.

## Policies and spending

Project inference policies restrict model and deployment choices within the environment's policy. Stored policy and active enforcement are distinct: inspect the documented enforcement fields and limitations before relying on a constraint.

Budgets and usage belong to the [Billing and Usage API](/docs/cloud/apis/billing-and-usage). Authentication and API key ownership belong to the [Identity and Access API](/docs/cloud/apis/identity-and-access).

## Get started

[Generate a model response](/docs/cloud/generate-response) shows model discovery and a chat-completion request.

## API references

* [REST endpoints](/docs/cloud/rest/apis/ai-gateway-api)
* [GraphQL queries and mutations](/docs/cloud/graphql/apis/ai-gateway)
* [MCP tools](/docs/cloud/mcp/apis/ai-gateway)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.