# How to use Jev with the Vercel AI SDK

> How to use Jev with the Vercel AI SDK: install the TypeSafe provider, call experimental_evaluate, read confidence, use AI Gateway, and compare Jev with an LLM.

Author: [Flavio Copes](https://flaviocopes.com/about/) | Published: 2026-10-06 | Topics: [AI](https://flaviocopes.com/tags/ai/) | Canonical: https://flaviocopes.com/jev-vercel-ai-sdk/

To use Jev with the Vercel AI SDK, install `ai` and `@ai-sdk/typesafe-ai`, set `TYPESAFE_AI_API_KEY`, and pass `typeSafeAi.evaluationModel('jev-latest')` to the `experimental_evaluate` function. If you go through Vercel's AI Gateway instead, you skip the provider package and pass the model string `typesafe-ai/jev` with an `AI_GATEWAY_API_KEY`.

Jev is TypeSafe AI's decision model. You send it some text and a list of typed questions, and it answers each one with a yes/no probability, one option from a list, or a position on a scale, always as probabilities rather than generated text. I covered the model itself in my [deep dive into Jev](https://flaviocopes.com/jev/). This post is about the AI SDK side.

## The quick answer

| What you want | How |
| --- | --- |
| Install | `npm install ai @ai-sdk/typesafe-ai` (Node.js 22 or newer) |
| API key, direct | `TYPESAFE_AI_API_KEY` |
| Model, direct | `typeSafeAi.evaluationModel('jev-latest')` |
| Model, through AI Gateway | `'typesafe-ai/jev'` with `AI_GATEWAY_API_KEY` |
| Yes/no question | `type: 'boolean'`, read `.probability` |
| Pick one option | `type: 'choice'`, read `.choice` and `.probabilities` |
| Position on a scale | `type: 'score'`, read `.score` and `.probabilities` |
| TypeSafe's confidence | `result.providerMetadata?.typesafe?.confidence` |
| Zero data retention | `providerOptions: { gateway: { zeroDataRetention: true } }` |

## Why use the AI SDK instead of TypeSafe's own SDK?

TypeSafe ships its own JavaScript client, `@typesafe-ai/sdk`, and it's the closest thing to the raw API (I cover it in [how to use Jev in Node.js](https://flaviocopes.com/jev-nodejs/)). The AI SDK path makes sense when your app already uses the AI SDK for its LLM features. You keep one package, one set of error classes, one retry policy and one telemetry setup for everything. If you're new to it, my [Vercel AI SDK tutorial](https://flaviocopes.com/vercel-ai-sdk/) covers the LLM side.

It also makes comparisons cheap. `experimental_evaluate` works with OpenAI, Anthropic and Google models too, with the same questions and the same answer shapes, so you can test Jev against an LLM on your own data by swapping one line.

And there's billing. Through the [Vercel AI Gateway](https://flaviocopes.com/vercel-ai-gateway/), Jev calls land in the same logs, budgets and invoice as your other models, at the same $0.042 per million input tokens TypeSafe charges directly. To estimate what your workload costs, see [How much does Jev cost?](https://flaviocopes.com/jev-pricing/). As of late September 2026 TypeSafe has paused new signups for its direct API because of demand, while the Gateway route works with a Vercel account.

I'd stay on TypeSafe's SDK if you're on Node.js 20 (the AI SDK packages need Node.js 22 or newer), or if you want `confidence` right on each answer with the same names you see in TypeSafe's docs.

## How do I install the Jev provider?

Check `node --version` first. Both `ai` and `@ai-sdk/typesafe-ai` require Node.js 22 or newer.

```bash
npm install --save-exact ai @ai-sdk/typesafe-ai
```

The `experimental_` prefix is a real warning. The AI SDK docs say the evaluation API and its model specification may change in patch releases, not only in major versions. That's why I'd install with `--save-exact` and read the changelog before upgrading. The function arrived partway through AI SDK 7, so if the import fails on an existing project, update `ai`.

Then create a key in the [TypeSafe console](https://console.typesafe.ai/keys) and export it:

```bash
export TYPESAFE_AI_API_KEY=your_key_here
```

In a Next.js app it goes in `.env.local` instead, as we'll see later.

Be careful with the name. TypeSafe's own SDK reads `TYPESAFE_API_KEY`, while the AI SDK provider reads `TYPESAFE_AI_API_KEY`. If you already have the first one set, pass it explicitly with `createTypeSafeAi()`:

```ts
import { createTypeSafeAi } from '@ai-sdk/typesafe-ai'

const typeSafe = createTypeSafeAi({
  apiKey: process.env.TYPESAFE_API_KEY,
})

const jev = typeSafe.evaluationModel('jev-1.13.0')
```

`createTypeSafeAi()` also accepts `baseURL`, `headers` and a custom `fetch`. Notice I used the versioned ID `jev-1.13.0` here instead of the `jev-latest` alias. Pin the version once you've tuned your thresholds against it.

## How do I run my first evaluation?

`experimental_evaluate` takes a `model`, a `state` (the text or JSON the questions are about) and a map of `questions`. Each key in `questions` becomes a key in `result.answers`.

Here's a support ticket with one question of each type:

```ts
import { experimental_evaluate as evaluate } from 'ai'
import { typeSafeAi } from '@ai-sdk/typesafe-ai'

const result = await evaluate({
  model: typeSafeAi.evaluationModel('jev-latest'),
  state: {
    ticket: 'I was charged twice for my October invoice. Please refund the duplicate payment.',
  },
  questions: {
    department: {
      type: 'choice',
      instructions: 'Which team should handle `ticket`?',
      criteria: {
        billing: 'Charges, invoices, refunds, subscriptions',
        technical: 'Bugs, outages, integration problems',
        other: null,
      },
    },
    urgency: {
      type: 'score',
      instructions: 'How urgent is `ticket`?',
      criteria: [
        'No deadline and no consequence mentioned',
        'The customer wants it solved soon, but nothing is blocked',
        'The customer is losing money or customers right now',
      ],
    },
    requests_refund: {
      type: 'boolean',
      instructions: 'Does the customer ask for money back in `ticket`?',
    },
  },
})

console.log(result.answers.department.choice)
console.log(result.answers.urgency.score)
console.log(result.answers.requests_refund.probability)
console.log(result.response.modelId)
```

I import it as `evaluate` to keep the calls short. The state can be a string, a JSON object or an array. An object lets each question point at a field with backticks, like `` `ticket` ``. A `null` description is fine for a catch-all `other` option.

`result.answers` comes back in this shape. The numbers are illustrative, not from a real call:

```json
{
  "department": {
    "type": "choice",
    "choice": "billing",
    "probabilities": { "billing": 0.97, "technical": 0.02, "other": 0.01 }
  },
  "urgency": {
    "type": "score",
    "score": 0.9,
    "probabilities": { "0": 0.2, "1": 0.7, "2": 0.1 }
  },
  "requests_refund": { "type": "boolean", "probability": 0.98 }
}
```

`result.response.modelId` holds the version that answered, such as `jev-1.13.0`, so log it next to the answers.

You also get real types. `result.answers.department.choice` is typed as `'billing' | 'technical' | 'other'`, inferred from your `criteria` keys, and a typo like `department.choice === 'tech'` fails type checking.

All the questions go to TypeSafe in one request against the same state. Jev evaluates them in parallel, so an extra question costs its own tokens and barely changes the response time.

## How do Jev's question types map to the AI SDK?

The AI SDK uses neutral names, because the same function works with other providers. Choice and Score keep their names, and Noul becomes `boolean`:

| TypeSafe API | AI SDK | What you read |
| --- | --- | --- |
| `noul` | `boolean` | `probability` (instead of `noul`) |
| `choice` | `choice` | `choice`, `probabilities` |
| `score` | `score` | `score`, `probabilities` |

If you copy a question from TypeSafe's docs and keep `type: 'noul'`, TypeScript rejects it. The provider translates `boolean` to `noul` on the way out and maps `noul` back to `probability` on the way in.

`probability` is the probability that the answer is yes, so 0.98 is a strong yes and 0.02 a strong no. When the line between yes and no is subtle, add optional `true` and `false` criteria:

```ts
requests_refund: {
  type: 'boolean',
  instructions: 'Does the customer ask for money back in `ticket`?',
  criteria: {
    true: 'The customer asks for a refund, a chargeback, or account credit',
    false: 'The customer only reports a problem or asks a question',
  },
},
```

The answer objects drop the `legend` and `confidence` fields of the raw API. You don't need the legend, because it's your own `criteria` array and the `probabilities` keys (`'0'`, `'1'`, `'2'`) are indexes into it. Confidence moved, and we'll get to it in a second.

The limits are TypeSafe's: up to 255 options per Choice and 2 to 10 levels per Score. Go over and the provider throws an `InvalidArgumentError` before sending anything. TypeSafe also rounds probabilities to two decimals, so a distribution can add up to 0.99. The SDK allows for that, so don't renormalize the values yourself.

## Where is TypeSafe's confidence?

It lives in `result.providerMetadata.typesafe.confidence`, keyed by question ID. You get it for Choice and Score answers only, because a boolean's probability already tells you how sure the model is.

The AI SDK types provider metadata as generic JSON, so you cast it. Continuing the ticket example:

```ts
const confidence = result.providerMetadata?.typesafe?.confidence as
  | Record<string, number>
  | undefined

const departmentConfidence = confidence?.department ?? 0
const { department, requests_refund } = result.answers

if (departmentConfidence < 0.6) {
  console.log('Send the ticket to a person')
} else if (department.choice === 'billing' && requests_refund.probability > 0.8) {
  console.log('Open a refund review in the billing queue')
} else {
  console.log(`Assign the ticket to ${department.choice}`)
}
```

TypeSafe computes [confidence](https://docs.typesafe.ai/confidence) from the shape of the distribution. It's close to 1 when one option takes most of the probability and close to 0 when the probability is spread out. It's not the selected option's probability.

My post on [how to use Jev confidence scores](https://flaviocopes.com/jev-confidence/) goes deeper into picking thresholds. Notice the `?? 0`. If the metadata is ever missing, the ticket goes to a person instead of being routed on a guess. The 0.6 and 0.8 thresholds are starting points, so move them once you've run labeled tickets through.

## How do I use Jev through the Vercel AI Gateway?

Pass the model as a string. In the AI SDK, string model IDs resolve through the AI Gateway by default, and `ai` already depends on the Gateway provider, so you don't need `@ai-sdk/typesafe-ai` at all.

Create a key in the AI Gateway dashboard and export it:

```bash
export AI_GATEWAY_API_KEY=your_gateway_key
```

On a Vercel deployment you can skip the key, because the Gateway also accepts the project's OIDC token. Locally, `vercel env pull` writes one to your env file, and it expires after 12 hours.

Here's a sponsor email checked through the Gateway:

```ts
import { experimental_evaluate as evaluate } from 'ai'

const result = await evaluate({
  model: 'typesafe-ai/jev',
  state: 'Hi Flavio, we run a managed Postgres service and would like to sponsor your newsletter in November.',
  questions: {
    sponsor_inquiry: {
      type: 'boolean',
      instructions: 'Is this message asking to sponsor the site or the newsletter?',
    },
  },
  providerOptions: {
    gateway: { zeroDataRetention: true },
  },
})

console.log(result.answers.sponsor_inquiry.probability)
```

Vercel's [Jev model page](https://vercel.com/ai-gateway/models/jev) lists it as `typesafe-ai/jev` at $0.042 per million input tokens, with no charge for output. Calls appear in the Gateway logs and count toward your budgets. Vercel's own [Jev and AI SDK guide](https://vercel.com/kb/guide/typesafe-jev-and-ai-sdk) reads confidence from the same `providerMetadata.typesafe.confidence` path, and the Gateway adds its routing and cost data under `providerMetadata.gateway`. On this path `result.response.modelId` is `typesafe-ai/jev`, not the version number.

`zeroDataRetention: true` makes the Gateway route the request only to providers with a zero data retention agreement, and TypeSafe AI is on Vercel's list of ZDR providers. If no ZDR provider can serve a model, the request fails with a 400 instead of going through. Per the [Gateway ZDR docs](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr), checked in September 2026, ZDR needs a Pro or Enterprise plan. Per-request ZDR costs nothing extra, while turning it on team-wide from the dashboard costs $0.10 per 1,000 requests.

Already have code on TypeSafe's SDK? You don't have to rewrite it to bill through Vercel. The Gateway exposes a TypeSafe-compatible API at `https://ai-gateway.vercel.sh/typesafe`, so you change the client's `baseURL` and use your Gateway key.

## How do I call Jev from a Next.js route handler?

Keep every Jev call on the server. A key in client-side JavaScript is readable by anyone who opens the browser dev tools. In the Next.js App Router a route handler only runs on the server, so it's the natural place for the call.

Put the key in `.env.local`, without a `NEXT_PUBLIC_` prefix, so Next.js never bundles it for the browser:

```bash
TYPESAFE_AI_API_KEY=your_key_here
```

Then create `app/api/triage/route.ts`:

```ts
import { experimental_evaluate as evaluate } from 'ai'
import { typeSafeAi } from '@ai-sdk/typesafe-ai'

const jev = typeSafeAi.evaluationModel('jev-latest')

async function triage(message: string) {
  const result = await evaluate({
    model: jev,
    state: { message },
    questions: {
      department: {
        type: 'choice',
        instructions: 'Which team should handle `message`?',
        criteria: {
          billing: 'Charges, invoices, refunds, subscriptions',
          technical: 'Bugs, outages, integration problems',
          sales: 'Pricing questions, upgrades, new accounts',
          other: null,
        },
      },
      requests_refund: {
        type: 'boolean',
        instructions: 'Does the author of `message` ask for money back?',
      },
    },
  })

  const confidence = result.providerMetadata?.typesafe?.confidence as
    | Record<string, number>
    | undefined
  const { department, requests_refund } = result.answers

  if ((confidence?.department ?? 0) < 0.6) {
    return { route: 'human', reason: 'unclear department' }
  }

  return {
    route: department.choice,
    refundRequested: requests_refund.probability > 0.8,
    model: result.response.modelId,
  }
}

export async function POST(request: Request) {
  const { message } = (await request.json()) as { message?: unknown }

  if (typeof message !== 'string' || message.length === 0 || message.length > 5000) {
    return Response.json({ error: 'message must be 1 to 5,000 characters' }, { status: 400 })
  }

  try {
    return Response.json(await triage(message))
  } catch (error) {
    console.error(error)
    return Response.json({ route: 'human', reason: 'evaluation failed' })
  }
}
```

The handler checks the input before spending any tokens, asks both questions in one call, and returns a routing decision instead of the raw answers. If the evaluation throws, the message goes to a person and the form still gets a response.

The AI SDK retries rate limits (`429`) and overload errors (`529`) twice by default, and you can change that with the `maxRetries` option. Failed requests throw an `APICallError`, and an answer that doesn't match your questions throws an `InvalidResponseDataError`.

Start the app with `npm run dev` and try it:

```bash
curl -X POST http://localhost:3000/api/triage \
  -H "Content-Type: application/json" \
  -d '{"message": "I was charged twice for my October invoice. Please refund the duplicate payment."}'
```

To test the thresholds without calling Jev, make the model a parameter of `triage()` and pass `Experimental_EvaluationMockModelV4` from `ai/test`. It returns whatever answers and metadata you give it.

## How do I compare Jev with an LLM on the same questions?

`experimental_evaluate` doesn't care which model answers. The OpenAI, Anthropic and Google providers each have an `evaluationModel()` factory that wraps a language model in an adapter. The adapter sends your state and questions in one prompt and asks for structured JSON back.

Install the OpenAI provider and set its key:

```bash
npm install @ai-sdk/openai
export OPENAI_API_KEY=your_openai_key
```

Then run the same questions through both models over labeled tickets, counting correct answers, time and tokens:

```ts
import {
  experimental_evaluate as evaluate,
  type Experimental_EvaluationModel,
  type Experimental_EvaluationQuestion,
} from 'ai'
import { typeSafeAi } from '@ai-sdk/typesafe-ai'
import { openai } from '@ai-sdk/openai'

const questions = {
  department: {
    type: 'choice',
    instructions: 'Which team should handle `ticket`?',
    criteria: {
      billing: 'Charges, invoices, refunds, subscriptions',
      technical: 'Bugs, outages, integration problems',
      other: null,
    },
  },
} satisfies Record<string, Experimental_EvaluationQuestion>

const labeled = [
  { ticket: 'I was charged twice for my October invoice.', expected: 'billing' },
  { ticket: 'The CSV export returns a 500 error since this morning.', expected: 'technical' },
  { ticket: 'Can I change the email address on my account?', expected: 'other' },
]

const models: Record<string, Experimental_EvaluationModel> = {
  jev: typeSafeAi.evaluationModel('jev-latest'),
  'gpt-5.6-luna': openai.evaluationModel('gpt-5.6-luna'),
}

for (const [name, model] of Object.entries(models)) {
  let correct = 0
  let inputTokens = 0
  let outputTokens = 0
  const start = performance.now()

  for (const { ticket, expected } of labeled) {
    const result = await evaluate({ model, state: { ticket }, questions })
    if (result.answers.department.choice === expected) correct++
    inputTokens += result.usage.inputTokens ?? 0
    outputTokens += result.usage.outputTokens ?? 0
  }

  const ms = Math.round((performance.now() - start) / labeled.length)
  console.log(`${name}: ${correct}/${labeled.length} correct, ${ms} ms per call, ${inputTokens} in, ${outputTokens} out`)
}
```

`satisfies` keeps the option names as literal types, so both models return a typed `department.choice`. Three tickets only show the idea. Use a few dozen real ones with answers you've checked. `gpt-5.6-luna` is the example ID from the AI SDK docs, so swap in whichever model you'd otherwise pay for.

Before reading the results, know what the adapters don't do. This comes from the [AI SDK evaluation docs](https://ai-sdk.dev/docs/ai-sdk-core/evaluation) and the adapter source:

- Choice and Score answers from an LLM carry no `probabilities`, only the pick. Code like `answer.probabilities?.[answer.choice]` has to handle `undefined`.
- A boolean `probability` from an LLM is the model's own estimate, written into its JSON. The docs say it isn't guaranteed to be calibrated.
- There's no `confidence` in `providerMetadata`.
- The LLM sees all the questions in one prompt, while Jev evaluates each one independently.
- The adapters request `reasoning: 'none'` by default. You can turn reasoning on through `providerOptions`, and pay for it in time and output tokens.

So compare accuracy, latency and cost per task, and leave confidence out, because only Jev has it. When you price the runs, remember the LLM bills its output tokens and Jev doesn't.

## What should I do next?

Pick one decision your app makes today and write the questions for it. Run them through `experimental_evaluate` beside your current code and log what comes back, including the confidence and the model ID. Let Jev route anything only after its answers agree with what your code decided. When an upgrade breaks something, the [AI SDK evaluation docs](https://ai-sdk.dev/docs/ai-sdk-core/evaluation) are the first place to look.
