At a glance
- Availability: Experimental (how to enable).
- Authentication: API key.
- Connection: The key comes from
REPLICATE_API_TOKEN. - Provider documentation: Authentication reference.
Credentials
Set these per environment. See Connect an integration.Setup
- Create a Replicate account: Go to https://replicate.com and sign up (GitHub sign-in supported). New accounts get a small amount of free usage before billing is required.
- Create an API token: Open https://replicate.com/account/api-tokens, give the token a descriptive name (e.g. ‘Veryfront Integration’), and create it.
- Store the token: Copy the token and add it to your .env file as REPLICATE_API_TOKEN=r8_…
- Verify access: Run the List Models tool to confirm the token works. A 401 means the token is wrong or revoked.
Provider notes
- Predictions are billed per second of compute - costs vary widely by model and hardware
- Create Prediction sends Prefer: wait=60 to return synchronously when possible; long-running models still return status ‘starting’ or ‘processing’ - poll with Get Prediction
- Use the version ID from Get Model’s latest_version.id field when creating predictions
Tools
Verify the connection
Call a read tool such asreplicate__list_models with arguments for your account. Confirm that the result comes from the intended account or workspace before enabling write tools.