Cloudflare AI Gateway
Everruns runs agents through Cloudflare’s AI REST API, which fronts many upstream model providers and Cloudflare’s own Workers AI models behind one account-scoped, OpenAI-compatible endpoint. Every call routes through an AI Gateway for caching, rate limiting, and logging.
What you get
Section titled “What you get”- Many vendors, one token: OpenAI, Anthropic, Google, xAI and the rest of the catalog, plus Workers AI models, from a single Everruns provider.
- No upstream keys: Cloudflare authenticates and bills the whole call against your account, so there are no per-vendor API keys to manage.
- Gateway controls: response caching, rate limiting, retries, and per-request logging applied by the gateway the call routes through.
- Full chat capabilities: streaming, tool/function calling, and structured output, through the same uniform driver as every other provider.
- Workers AI model discovery: the
@cf/catalog is synced with its advertised context windows and tool-calling support, so those ids appear in the model pickers without being typed by hand.
Configure in Everruns
Section titled “Configure in Everruns”- Create an API token with the Account → Workers AI → Read permission in the Cloudflare dashboard.
- Go to Settings → Providers and click Add provider.
- Choose Cloudflare AI Gateway.
- Enter:
- API Token — the token from step 1.
- Account ID — the account it belongs to.
- Gateway name — optional. Calls route through the account’s default gateway when it is unset.
- Save.
Everruns derives the endpoint from the account id:
https://api.cloudflare.com/client/v4/accounts/<account-id>/ai/v1Set a base URL only to route through a proxy in front of Cloudflare; it replaces the derived one.
Models
Section titled “Models”Model ids are namespaced and passed through unchanged:
- Third-party:
openai/gpt-6-luna,anthropic/claude-opus-5— see the upstream providers AI Gateway supports. - Workers AI:
@cf/meta/llama-3.3-70b-instruct-fp8-fast— see the Workers AI model catalog.
Which third-party models a gateway can reach depends on the upstreams Cloudflare has enabled for your account, so verify an id with a real request rather than assuming the catalog.
Third-party models are billed to your Cloudflare account and return 402
(“Insufficient balance”) until it is funded or you configure BYOK. Workers AI
models draw on the account’s own allocation instead.
Model sync covers the Workers AI half of the catalog: Everruns reads Cloudflare’s model search endpoint and keeps its text-generation models, which arrive with the context window and tool-calling support Cloudflare advertises. Third-party models are never listed — which ones an account can reach depends on what it can bill — so add those by id. Everruns matches a namespaced id against its model profile registry, so a model it recognizes still gets its capability and cost metadata.
Why Chat Completions
Section titled “Why Chat Completions”Cloudflare’s REST API offers four formats. Everruns uses
/ai/v1/chat/completions because it is the only one that serves the whole
catalog and streams:
| Endpoint | Why not |
|---|---|
/ai/v1/responses | Rejects Workers AI models — they are not translated into the Responses shape |
/ai/run | Cannot stream: with stream: true it returns an empty JSON body rather than events |
/ai/v1/messages | Excludes Workers AI models |
The older gateway.ai.cloudflare.com/.../compat/chat/completions endpoint is
deprecated by Cloudflare for single-model calls and kept only for dynamic
routes (dynamic/{route}), which this provider does not target.