agents/models/ai-sdk is an AI SDK ↗︎ provider for Workers AI and AI Gateway. createAI returns a full ProviderV4, so its models work anywhere the AI SDK takes a model: generateText, streamText, tools, structured output, and Think.
The provider needs ai@^7 and @ai-sdk/provider@^4. Both are optional peer dependencies of agents:
npm i agents ai @ai-sdk/provideryarn add agents ai @ai-sdk/providerpnpm add agents ai @ai-sdk/providerbun add agents ai @ai-sdk/providerAdd a vendor's AI SDK provider only if you call that vendor's models, for example @ai-sdk/anthropic or @ai-sdk/openai.
Add the AI binding to your Wrangler configuration:
{
"ai": {
"binding": "AI",
},
}[ai]
binding = "AI"Pass a @cf/ id. Model ids autocomplete from @cloudflare/workers-types:
import { createAI } from "agents/models/ai-sdk";
import { generateText } from "ai";
export default {
async fetch(request, env) {
const ai = createAI({ binding: env.AI });
const { text } = await generateText({
model: ai("@cf/zai-org/glm-4.7-flash"),
prompt: "Say hello in three words",
});
return Response.json({ text });
},
};import { createAI } from "agents/models/ai-sdk";
import { generateText } from "ai";
export default {
async fetch(request: Request, env: Env) {
const ai = createAI({ binding: env.AI });
const { text } = await generateText({
model: ai("@cf/zai-org/glm-4.7-flash"),
prompt: "Say hello in three words",
});
return Response.json({ text });
},
};Any string that is not a @cf/ id throws a TypeError that names the provider package to install.
Build the model with the vendor's own AI SDK provider, then pass it to ai(). The request goes through AI Gateway, which holds the credential, so the vendor's apiKey is a placeholder:
import { createAnthropic } from "@ai-sdk/anthropic";
import { createOpenAI } from "@ai-sdk/openai";
import { createAI } from "agents/models/ai-sdk";
import { generateText } from "ai";
export default {
async fetch(request, env) {
const ai = createAI({ binding: env.AI });
const anthropic = createAnthropic({ apiKey: "cloudflare" });
const openai = createOpenAI({ apiKey: "cloudflare" });
const claude = await generateText({
model: ai(anthropic("claude-opus-4-8")),
prompt: "Say hello in three words",
});
const gpt = await generateText({
model: ai(openai.responses("gpt-5-mini")),
prompt: "Say hello in three words",
});
return Response.json({ claude: claude.text, gpt: gpt.text });
},
};import { createAnthropic } from "@ai-sdk/anthropic";
import { createOpenAI } from "@ai-sdk/openai";
import { createAI } from "agents/models/ai-sdk";
import { generateText } from "ai";
export default {
async fetch(request: Request, env: Env) {
const ai = createAI({ binding: env.AI });
const anthropic = createAnthropic({ apiKey: "cloudflare" });
const openai = createOpenAI({ apiKey: "cloudflare" });
const claude = await generateText({
model: ai(anthropic("claude-opus-4-8")),
prompt: "Say hello in three words",
});
const gpt = await generateText({
model: ai(openai.responses("gpt-5-mini")),
prompt: "Say hello in three words",
});
return Response.json({ claude: claude.text, gpt: gpt.text });
},
};Use the vendor's own model id. The provider does not rewrite ids, so claude-opus-4.8 reaches Anthropic as written and Anthropic rejects it.
Providers on Unified Billing answer with no key of yours. For any other provider, store a key on the gateway. Until you do, the vendor's own error comes back unchanged.
ai() clones the model for each call and never changes the object you passed. If you add middleware, wrap the model with ai() first:
wrapLanguageModel({ model: ai(anthropic("claude-opus-4-8")), middleware });A model that is already wrapped in middleware hides the settings the provider needs, so ai() rejects it.
Both kinds of model work with every AI SDK call:
import { createAI } from "agents/models/ai-sdk";
import { generateText, isStepCount, Output, streamText, tool } from "ai";
import { z } from "zod";
const getWeather = tool({
description: "Get the current weather for a city.",
inputSchema: z.object({ city: z.string() }),
execute: ({ city }) => ({ city, conditions: "clear", temperatureC: 21 }),
});
export default {
async fetch(request, env) {
const ai = createAI({ binding: env.AI });
const model = ai("@cf/zai-org/glm-4.7-flash");
const { pathname } = new URL(request.url);
if (pathname === "/stream") {
const result = streamText({
model,
prompt: "Write a haiku about Durable Objects",
});
return result.toTextStreamResponse();
}
if (pathname === "/tools") {
const { text } = await generateText({
model,
prompt: "What is the weather in Lisbon?",
tools: { getWeather },
stopWhen: isStepCount(3),
});
return Response.json({ text });
}
const { output } = await generateText({
model,
prompt: "One surprising fact about octopuses.",
output: Output.object({ schema: z.object({ fact: z.string() }) }),
});
return Response.json(output);
},
};import { createAI } from "agents/models/ai-sdk";
import { generateText, isStepCount, Output, streamText, tool } from "ai";
import { z } from "zod";
const getWeather = tool({
description: "Get the current weather for a city.",
inputSchema: z.object({ city: z.string() }),
execute: ({ city }) => ({ city, conditions: "clear", temperatureC: 21 }),
});
export default {
async fetch(request: Request, env: Env) {
const ai = createAI({ binding: env.AI });
const model = ai("@cf/zai-org/glm-4.7-flash");
const { pathname } = new URL(request.url);
if (pathname === "/stream") {
const result = streamText({
model,
prompt: "Write a haiku about Durable Objects",
});
return result.toTextStreamResponse();
}
if (pathname === "/tools") {
const { text } = await generateText({
model,
prompt: "What is the weather in Lisbon?",
tools: { getWeather },
stopWhen: isStepCount(3),
});
return Response.json({ text });
}
const { output } = await generateText({
model,
prompt: "One surprising fact about octopuses.",
output: Output.object({ schema: z.object({ fact: z.string() }) }),
});
return Response.json(output);
},
};Return a model from getModel():
import { Think } from "@cloudflare/think";
import { createAI } from "agents/models/ai-sdk";
export class Assistant extends Think {
getModel() {
return createAI({ binding: this.env.AI })("@cf/moonshotai/kimi-k2.7-code");
}
}import { Think } from "@cloudflare/think";
import { createAI } from "agents/models/ai-sdk";
export class Assistant extends Think<Env> {
getModel() {
return createAI({ binding: this.env.AI })("@cf/moonshotai/kimi-k2.7-code");
}
}AI Gateway can cache responses. Set cacheTtl in seconds, and optionally a cacheKey. Pass skipCache: true on a call to bypass the cache:
const ai = createAI({ binding: env.AI, id: "prod", cacheTtl: 60 });
// A longer TTL for one model.
const faq = ai("@cf/zai-org/glm-4.7-flash", {
cacheTtl: 3600,
cacheKey: "faq-v1",
});
// Skip the cache for one call.
await generateText({
model: faq,
prompt: "What changed today?",
providerOptions: { cloudflare: { skipCache: true } },
});Attach metadata to the gateway log entry. Set it once on the provider and add to it per call. Metadata merges key by key:
const ai = createAI({ binding: env.AI, metadata: { app: "support-bot" } });
await generateText({
model: ai(anthropic("claude-opus-4-8")),
prompt: "Summarize this ticket",
providerOptions: {
cloudflare: { metadata: { userId: "u_123" }, eventId: "ticket-42" },
},
});eventId groups several requests into one gateway event, such as all the model calls in one agent turn.
fallback lists models to try in order if the primary model fails before it produces any output. Mix Workers AI ids and model objects:
const model = ai(anthropic("claude-opus-4-8"), {
fallback: [ai(openai.responses("gpt-5-mini")), "@cf/zai-org/glm-4.7-flash"],
});A request that the vendor rejects falls back. So does a stream that fails before its first output part. Once a model has produced output, its answer is final, errors included. providerMetadata.cloudflare.model names the model that answered.
Each fallback model uses the chain's gateway options. The chain's id, cacheTtl, and the other gateway options win over a fallback model's own, and metadata merges. Headers and Workers AI options a fallback model was built with stay its own.
Fallback runs in your Worker today.
Set gateway options on the provider, on a model, or per call under providerOptions.cloudflare. A per-call option wins over a per-model option, which wins over the provider's:
// Provider
const ai = createAI({ binding: env.AI, id: "prod", metadata: { app: "demo" } });
// Model
const model = ai(anthropic("claude-opus-4-8"), { cacheTtl: 3600 });
// Call
await generateText({
model,
prompt: "hi",
providerOptions: {
cloudflare: { skipCache: true, metadata: { userId: "u1" } },
},
});The provider removes providerOptions.cloudflare before the call reaches a vendor's model, so the vendor never sees it.
Gateway options can be given flat or nested: { cacheTtl: 60 } and { gateway: { cacheTtl: 60 } } are the same.
| Option | Type | Description |
|---|---|---|
id |
string |
Gateway id. Defaults to "default", which is created the first time you use it. |
gateway |
string | GatewayOptions |
The gateway id, or the gateway options nested. |
skipCache |
boolean |
Bypass the gateway cache. |
cacheTtl |
number |
Cache lifetime in seconds. |
cacheKey |
string |
Cache key to use instead of the one derived from the request. |
collectLog |
boolean |
Whether the gateway stores the request and response in its log. |
metadata |
Record<string, string | number | boolean | null> |
Values attached to the gateway log entry. |
eventId |
string |
Groups several requests into one gateway event. |
requestTimeoutMs |
number |
Timeout for the upstream request. |
retries |
{ maxAttempts, retryDelayMs, backoff } |
Gateway retry policy for the upstream call. maxAttempts is 1 to 5. |
fallback |
(WorkersAIModelId | LanguageModelV4)[] |
Models to try in order. Model and call only. |
headers |
Record<string, string> |
Extra request headers, merged last. Model and call only. |
sessionAffinity |
string |
Workers AI only. Sends requests with the same key to one replica for prefix caching. |
reasoningEffort |
"low" | "medium" | "high" | null |
Workers AI only. null turns reasoning off on models that allow it. |
chatTemplateKwargs |
Record<string, unknown> |
Workers AI only. Sent as chat_template_kwargs, for example { enable_thinking: false }. |
For a vendor's model, ask for reasoning with the AI SDK's reasoning call option or the vendor's own providerOptions.
Every result, and every stream's finish part, carries providerMetadata.cloudflare. For a vendor's model it sits next to the vendor's own metadata:
const { providerMetadata } = await generateText({
model: ai(anthropic("claude-opus-4-8")),
prompt: "hi",
});
providerMetadata?.cloudflare;
// {
// model: "claude-opus-4-8",
// provider: "anthropic",
// gateway: "default",
// logId: "01M1KZWN069WWNPC18V05NKHSS",
// cacheStatus: "MISS",
// ...
// }| Field | Description |
|---|---|
model |
The model that answered, after any fallback. |
gateway |
The gateway id the request went through. |
provider |
The gateway provider, for a vendor's model. |
logId |
The gateway log entry, from cf-aig-log-id. |
eventId |
The gateway event, from cf-aig-event-id. |
requestId |
The gateway request id, from cf-aig-request-id. |
traceId |
The trace this request belongs to, from cf-aig-trace-id. |
runId |
From cf-aig-run-id, when the gateway sends one. |
cacheStatus |
HIT or MISS, from cf-aig-cache-status. |
step |
The step of a gateway route that answered, from cf-aig-step. |
Only model and gateway are always present. Check for the other fields before you use them.
The provider throws CloudflareAIError, which extends the AI SDK's APICallError. The AI SDK's retry loop reads its isRetryable flag:
import { CloudflareAIError } from "agents/models/ai-sdk";
try {
await generateText({ model: ai("@cf/zai-org/glm-4.7-flash"), prompt: "hi" });
} catch (error) {
if (error instanceof CloudflareAIError) {
console.log(error.status, error.code, error.logId, error.model);
// error.attempts lists every fallback model that failed.
}
}code is one of "auth", "rate-limit", "not-found", "bad-request", "provider-error", "gateway-error", or "unknown".
For a vendor's model, CloudflareAIError covers the gateway's own failures, such as a missing gateway or an unpaid account. The vendor's own errors, such as an unknown model id or a rate limit, come from the vendor's provider in its own error type.
Embeddings, images, speech, transcription, and reranking use Workers AI models from the same provider. Each method also has its ProviderV4 name, such as ai.embeddingModel():
import {
embedMany,
generateImage,
generateSpeech,
rerank,
transcribe,
} from "ai";
const { embeddings } = await embedMany({
model: ai.embedding("@cf/baai/bge-base-en-v1.5"),
values: ["hello", "world"],
});
const { image } = await generateImage({
model: ai.image("@cf/black-forest-labs/flux-1-schnell"),
prompt: "A cyberpunk lizard",
});
const { audio } = await generateSpeech({
model: ai.speech("@cf/deepgram/aura-1"),
text: "Hello from Cloudflare.",
});
const { text } = await transcribe({
model: ai.transcription("@cf/openai/whisper-large-v3-turbo"),
audio: audio.uint8Array,
});
const { rerankedDocuments } = await rerank({
model: ai.reranking("@cf/baai/bge-reranker-base"),
documents: ["a cyberpunk lizard", "a rainy Tuesday"],
query: "Which one is cooler?",
});Model-specific settings go under providerOptions.cloudflare:
| Modality | Call options used | providerOptions.cloudflare |
|---|---|---|
| Image | prompt, n, size, seed, files, mask |
steps, guidance, negativePrompt, strength |
| Speech | text, voice (as speaker), outputFormat (as encoding), language (MeloTTS only) |
None |
| Transcription | audio, mediaType |
language, task, initialPrompt, prefix, vadFilter |
| Reranking | query, documents, topN (as top_k) |
None |
An option a model does not take is listed in the result's warnings instead of being sent. For example, flux-1-schnell takes only a prompt and steps. Image generation makes one image per request, so generateImage with a larger n sends parallel requests.
To route a vendor's own embedding, image, speech, or transcription model through AI Gateway, pass it to ai.routed(model). ai.routed() takes the gateway options, headers, and sessionAffinity.
createAI returns a ProviderV4, so it works with createProviderRegistry and as the AI SDK's global default provider:
import { createProviderRegistry, generateText } from "ai";
const registry = createProviderRegistry({ cloudflare: ai });
await generateText({
model: registry.languageModel("cloudflare:@cf/zai-org/glm-4.7-flash"),
prompt: "Hi",
});
// Bare strings now resolve through this provider.
globalThis.AI_SDK_DEFAULT_PROVIDER = createAI({ binding: env.AI });
await generateText({ model: "@cf/zai-org/glm-4.7-flash", prompt: "Hi" });A string still has to be a @cf/ id. Reach a vendor's model with a model object.
- AI Gateway's universal request carries JSON only. A multipart body, such as a file upload sent as
FormData, throws aTypeError. @cf/deepgram/nova-3and@cf/deepgram/fluxare not supported. They throw aCloudflareAIErrorbefore any request is sent. Use@cf/openai/whisper-large-v3-turbofor transcription.reasoningEffort,chatTemplateKwargs, andsessionAffinityapply to Workers AI models only.