Cloudflare Workers AI: run LLMs without API keys

By

Run AI models on Cloudflare's GPUs directly from a Worker with the AI binding. Models, pricing in neurons, streaming, and a real production setup.

~~~

Most AI features start the same way: sign up for OpenAI, create an API key, store it as a secret, write a fetch call.

Workers AI skips all of that. The model runs on Cloudflare’s own GPUs, and you call it from a Worker like you’d call KV or D1: through a binding.

No API key. No external provider. One line of config.

I use this in production on this site — the app idea generator runs on Workers AI. In this post I’ll show you how it works and what I learned running it.

What is Workers AI?

Workers AI is Cloudflare’s inference service. They host a catalog of open models — Llama, Mistral, Qwen, embedding models, Whisper for audio, image generation models — and you run them on their infrastructure.

You don’t manage GPUs. You don’t pick a region. You call env.AI.run() with a model name and get a response.

The catalog changes over time. You can list what’s available with:

npx wrangler ai models list

Why use it?

Three reasons made me pick it over OpenAI for my tools:

The tradeoff is honest: these are open models, not frontier models. A 3B Llama won’t write like Claude. But for many features — generating ideas, summarizing, classifying, embeddings — a small model is plenty.

Setting up the binding

Add the AI binding to wrangler.jsonc:

{
  "ai": {
    "binding": "AI"
  }
}

That’s the whole setup. The binding shows up in your Worker as env.AI.

Running a model

Call env.AI.run() with the model name and the input:

export default {
  async fetch(request, env) {
    const result = await env.AI.run('@cf/meta/llama-3.2-3b-instruct', {
      messages: [
        { role: 'system', content: 'You are a helpful assistant. Be brief.' },
        { role: 'user', content: 'What is a Cloudflare Worker?' },
      ],
      max_tokens: 256,
    })

    return Response.json(result)
  },
}

The response contains the generated text plus token counts:

{
  "response": "A Cloudflare Worker is a serverless function...",
  "usage": {
    "prompt_tokens": 27,
    "completion_tokens": 89
  }
}

The messages format is the same one OpenAI uses, so if you’ve written a chat completion call before, this will feel familiar.

Notice max_tokens. My advice is to always set it, and keep it low. It caps your cost per request, and it forces you to design prompts that get to the point.

Streaming

For anything user-facing, you want tokens to appear as they generate. Pass stream: true and you get back a ReadableStream:

const stream = await env.AI.run('@cf/meta/llama-3.2-3b-instruct', {
  messages: [{ role: 'user', content: 'Explain HTTP caching' }],
  stream: true,
})

return new Response(stream, {
  headers: { 'content-type': 'text/event-stream' },
})

The stream uses server-sent event framing. For a chat request sent with POST, call the Worker with fetch() and read its ReadableStream response body. EventSource only makes GET requests, so it does not fit this pattern.

How pricing works: neurons

Workers AI doesn’t bill in tokens. It bills in neurons, a unit that normalizes cost across very different models (text, images, audio).

Two things matter in practice:

Each model’s page on developers.cloudflare.com shows how its usage converts to neurons. Small text models are cheap: on the 3B Llama I use, a day of normal traffic on my tool stays within the free allocation.

The usage field in each response tells you the token counts, so you can track spend yourself. I accumulate those counters in KV and check them on a small dashboard.

Be careful: models get deprecated

This is the part that bit me.

Workers AI retires models over time. Llama 3.1 8B, which a lot of tutorials still reference, was shut down in May 2026. If your Worker hardcodes a dead model, the call just starts failing.

Two defenses:

My tool falls back to a curated static list when the AI call errors. Users still get something, and I get time to swap the model name.

A real production setup

Calling a model is the easy part. Exposing it on a public endpoint is where you need to think.

My app idea generator endpoint does this, in order:

  1. verify a Turnstile token, so bots can’t hit the endpoint
  2. bump a per-visitor daily counter in KV — reject over the cap
  3. bump a global daily counter in KV — reject over the cap
  4. only then call env.AI.run()

The counters run before the AI call. That means the worst-case daily spend is a number I chose, not a number an attacker chose. With a global cap of 300 generations a day on a 3B model, the worst case is pennies.

If you’re putting Workers AI behind a public form, I’d copy this structure. The model being cheap doesn’t matter if someone scripts a million requests against it.

Beyond text

The same binding runs other model types. Embeddings, for example:

const result = await env.AI.run('@cf/baai/bge-base-en-v1.5', {
  text: 'Cloudflare Workers run JavaScript at the edge',
})

// result.data[0] is the embedding vector

That pairs naturally with Vectorize, Cloudflare’s vector database, if you’re building search or RAG on the platform.

When to skip it

If your feature needs frontier-model quality — long reasoning, nuanced writing, complex code generation — a small open model will disappoint you, and you should call OpenAI or Anthropic instead. You can still route those calls through Cloudflare with AI Gateway and keep the logging benefits.

But for small, well-scoped AI features living inside a Worker you already have, Workers AI is the shortest path I know from idea to production. One binding, one function call, no keys.

~~~

Related posts about cloudflare: