Vercel AI Gateway tutorial

By

Use Vercel AI Gateway with the AI SDK and OpenAI client, then add model routing, fallbacks, caching, budgets, BYOK, and observability.

~~~

Vercel AI Gateway gives you one API for models from OpenAI, Anthropic, Google, xAI, and many other providers.

Your application calls Vercel. Vercel picks an upstream provider, sends the request, and returns the response.

So you can switch models by changing a string, without installing another provider SDK. Usage logs, spend tracking, routing, and fallbacks all live in one place.

In this tutorial we’ll make a request, stream a response, point an existing OpenAI client at the gateway, and then set up the features you want before real users hit it.

What is an AI gateway?

Without a gateway, your application talks to each provider directly:

your app -> OpenAI
your app -> Anthropic
your app -> Google

Every provider has its own key, billing account, API details, and logs.

With Vercel AI Gateway, the shape changes:

your app -> Vercel AI Gateway -> model provider

Your application uses one AI Gateway key. Model names follow the creator/model format:

openai/gpt-5.4-mini
anthropic/claude-sonnet-4.6
google/gemini-3.1-flash-lite

Change that string and you change the model.

You don’t have to host on Vercel to use it. A VPS, a Cloudflare Worker, another serverless platform, or a script on your computer can all call the gateway.

Keep the API key on the server. Never call the gateway from browser JavaScript with a secret key.

Why use a gateway?

The code you write is about the same length with or without a gateway. What you gain is one place where every model call passes through.

That one place gives you a single key for many providers, a unified API, provider routing with automatic retries, model fallbacks, and logs for usage, tokens, latency, and spend. You can put a spending budget on each key. You can bring your own provider keys (BYOK) if you already have them. And there are OpenAI and Anthropic-compatible endpoints for clients you already wrote.

The cost is another service sitting between your application and the model.

For a small script that always calls one provider, I would call that provider directly. For a production application using several models, the gateway pays for itself quickly.

How pricing works

AI Gateway charges the upstream provider’s list price with no token markup. Check the current price in the model catalog before choosing a model.

Every Vercel team gets a monthly free AI Gateway credit. At the time of writing, it is $5. The free credit starts with the first request.

Once you purchase credits, the account moves to pay-as-you-go and the monthly free credit stops. Credits are prepaid and requests consume the balance.

BYOK also has no gateway markup. You pay the provider through your own account.

Prices change often, so don’t copy model prices into your code. Read them from the catalog when you need them.

Create an API key

Open your Vercel dashboard and go to AI Gateway → API Keys.

Click Create key, give it a clear name, and copy it immediately. Vercel does not show the value again.

For local development, create a .env file:

AI_GATEWAY_API_KEY=your_ai_gateway_api_key

Add the file to .gitignore:

.env

Create separate keys for separate applications. If one key leaks, you can revoke it without breaking every project.

Set a budget before the first request

An AI key without a budget can spend the team’s entire credit balance.

When you create the key, enable its budget. Choose a dollar limit and a daily, weekly, monthly, or non-resetting period.

You can also create a budgeted key with the Vercel CLI:

vercel ai-gateway api-keys create \
  --name tutorial \
  --budget 5 \
  --refresh-period monthly

The minimum budget is $1.

The budget is a soft cap. AI Gateway checks it before each request, so a request that starts below the limit can finish above it.

You still need rate limits in your application. The budget stops one key from draining the account. It does nothing to stop one user from calling your API in a loop.

Make the first request with the AI SDK

The AI SDK is the shortest path from TypeScript to AI Gateway.

Create a project:

mkdir vercel-ai-gateway-demo
cd vercel-ai-gateway-demo
npm init -y

Install the AI SDK, dotenv, and tsx:

npm install ai dotenv tsx

Create index.ts:

import 'dotenv/config'
import { generateText } from 'ai'

const { text, usage } = await generateText({
  model: 'openai/gpt-5.4-mini',
  prompt: 'Explain what an API gateway does in 3 sentences.',
})

console.log(text)
console.log(usage)

Run it:

npx tsx index.ts

There is no provider import in that file. The AI SDK sees the creator/model string, defaults to AI Gateway as the provider, and reads AI_GATEWAY_API_KEY from the environment.

The usage object tells you how many input and output tokens the request used.

Stream a response

For chat and longer generations you want to show text as it arrives, instead of waiting for the whole answer.

Replace generateText() with streamText():

import 'dotenv/config'
import { streamText } from 'ai'

const result = streamText({
  model: 'openai/gpt-5.4-mini',
  prompt: 'Write a short story about a robot learning Git.',
})

for await (const textPart of result.textStream) {
  process.stdout.write(textPart)
}

console.log()
console.log('Usage:', await result.usage)

The text shows up in the terminal while the model is still writing.

Keep in mind that streaming only changes when you see the text. You pay for the same number of tokens.

Use the OpenAI client

AI Gateway also exposes an OpenAI-compatible Chat Completions API.

If an application already uses the OpenAI client, you don’t have to rewrite it. Change the API key, the base URL, and the model name:

npm install openai

Then:

import 'dotenv/config'
import OpenAI from 'openai'

const client = new OpenAI({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh/v1',
})

const response = await client.chat.completions.create({
  model: 'anthropic/claude-sonnet-4.6',
  messages: [
    {
      role: 'user',
      content: 'Explain immutable deployments in 3 sentences.',
    },
  ],
})

console.log(response.choices[0].message.content)

Look at the model name. We’re calling an Anthropic model through the OpenAI client. The client speaks the OpenAI format, and AI Gateway translates the request for whichever provider is behind that model.

You don’t even need an SDK. https://ai-gateway.vercel.sh/v1/chat/completions works with plain fetch() or cURL.

Find available models

The public model endpoint does not require authentication:

curl https://ai-gateway.vercel.sh/v1/models

It returns model IDs, capabilities, context windows, and pricing.

In code:

const response = await fetch('https://ai-gateway.vercel.sh/v1/models')
const { data: models } = await response.json()

const textModels = models.filter((model) => model.type === 'language')

console.log(textModels.map((model) => model.id))

Don’t pass that list to a browser and let users pick anything they want. One expensive model in a loop can burn through a budget in minutes. Keep an allowlist of the models your application supports and reject everything else.

Control provider routing

A model and a provider are two different things. Anthropic creates Claude, but the same Claude model might be hosted by Anthropic, Amazon Bedrock, or Google Vertex AI.

By default, AI Gateway picks between the available providers based on recent uptime and latency.

You can control the order with providerOptions.gateway:

const { text } = await generateText({
  model: 'anthropic/claude-sonnet-4.6',
  prompt: 'Explain provider routing in one paragraph.',
  providerOptions: {
    gateway: {
      order: ['anthropic', 'bedrock'],
      only: ['anthropic', 'bedrock'],
    },
  },
})

order sets the preference. only prevents the gateway from using providers outside that list.

I would leave routing alone unless you have a specific reason, like a compliance rule, a data location requirement, or an existing agreement with one provider. The default already picks a healthy provider for you.

Add model fallbacks

Provider routing tries another host for the same model.

A model fallback tries a different model altogether when the primary one fails or is unavailable.

const result = streamText({
  model: 'openai/gpt-5.4-mini',
  prompt: 'Write a short description for a coffee shop.',
  providerOptions: {
    gateway: {
      models: [
        'google/gemini-3.1-flash-lite',
        'anthropic/claude-sonnet-4.6',
      ],
    },
  },
})

AI Gateway tries the primary model first, then each fallback in order.

Pick fallbacks that can do the same job. A text-only model can’t replace a vision model when your prompt contains an image.

And expect the answer to look different. A fallback keeps the request alive, but a different model writes differently and might format JSON differently too. If you parse structured output, validate it before your application trusts it.

Use automatic prompt caching

Long system prompts and agent conversations repeat a lot of text from one request to the next.

Some providers cache those repeated prefixes on their own. Others want you to mark them explicitly. AI Gateway can deal with that difference for you:

const { text } = await generateText({
  model: 'anthropic/claude-sonnet-4.6',
  system: 'You are a patient JavaScript tutor.',
  prompt: 'Explain closures with a small example.',
  providerOptions: {
    gateway: {
      caching: 'auto',
    },
  },
})

With caching: 'auto' the gateway adds the markers for providers that need them and leaves the others alone.

This saves money and time only when your requests share a stable prefix. If every prompt is different, there is nothing to cache.

Bring your own provider keys

By default, Vercel pays the provider and deducts the cost from your AI Gateway credits.

With Bring Your Own Key, AI Gateway uses credentials from your OpenAI, Anthropic, or other provider account.

Add provider credentials under AI Gateway → Bring Your Own Key in the Vercel dashboard.

The credentials are shared across the Vercel team. Your application keeps using its AI Gateway key, so the provider keys never have to live in the application.

Vercel tries your BYOK credentials first. If they fail, AI Gateway may fall back to Vercel’s own credentials to keep the request running. Because of that fallback, the team still needs AI Gateway credits even with BYOK turned on.

BYOK makes sense when you already have provider credits, negotiated pricing, or access to features only your provider account has. Otherwise, paying Vercel for everything is less to manage.

Use OIDC on Vercel deployments

An API key works everywhere, but it’s a long-lived secret you have to store and rotate.

If the application is deployed on Vercel, you can skip the key. Vercel generates an OIDC token for the deployment, and the AI SDK uses it to authenticate with the gateway.

Link the local directory to a Vercel project:

vercel link

Pull the development environment:

vercel env pull

That pulls VERCEL_OIDC_TOKEN into your local environment file. The AI SDK picks it up with no code changes.

Local tokens expire after 12 hours. Run vercel env pull again to get a fresh one.

Code running outside Vercel can’t get one of these tokens, so there you keep using an API key.

Read the request logs

After a few requests, open AI Gateway in the Vercel dashboard.

The overview shows requests by model, time to first token, input and output token counts, and spend, grouped by project and API key. Click a single generation and you see its model, the provider that served it, latency, token usage, cost, and finish reason.

The AI SDK also returns a gateway generation ID in providerMetadata. Store it next to your own request ID. When a user reports a bad answer, that ID takes you straight to the generation in the dashboard.

My advice is to look there before touching the prompt. Often the problem is latency, a provider error, a token limit, or a fallback that kicked in, and no prompt change would fix any of those.

Before you expose this to users

A few things I would check before an AI feature goes live.

Give each application its own gateway key, and put a budget on every key. Keep the keys out of browser code and out of Git.

Rate-limit your users before the request reaches the model, and only accept the models your application needs.

If you add fallbacks, test their output first. If you rely on structured responses, validate them. Log the gateway generation ID so you can debug later.

Then look at spend and failed requests every so often, and read the provider’s data policy before you send it anything sensitive.

The gateway sits between you and the model providers. It doesn’t authenticate your users, authorize them, rate-limit them, or check their input, so your application still has to do all of that.

When should you use Vercel AI Gateway?

Use it when you want to switch models, compare providers, see spend in one place, or add fallbacks without maintaining several integrations.

Skip it when a small script calls one provider and you don’t care about shared logs or routing.

If you’re unsure, start with one key, one cheap model, and one budget. Add routing, BYOK, caching, and fallbacks later, when the application gives you a reason.

Tagged: AI · All topics

Want me to talk about your product? You can sponsor this site.

~~~

Related posts about ai: