How to use Jev confidence scores

By

How to use Jev confidence scores to decide when code acts, asks for confirmation, or hands off to a person, with thresholds tuned on labeled data.

~~~

Every Choice and Score answer from Jev carries a confidence value between 0 and 1. It says how concentrated the answer’s probabilities are: 1.0 when all the probability sits on one option, lower as it spreads across several. You use it by splitting it into bands in your code: let the code act alone when it’s high, ask someone to confirm in the middle, hand off to a person when it’s low. The cost of a wrong answer decides where each band starts, and labeled examples tell you the exact numbers. Noul answers have no confidence, because the yes probability is already the signal.

Jev is TypeSafe’s decision model: you send it text and typed questions, and it returns yes/no probabilities, one option from a list, or a position on a scale, with probabilities instead of generated text. The deep dive into Jev covers the whole model. Here we focus on one field and the code around it.

The quick answer

Answer typeWhat to readWhat your code does
Choiceconfidence, then probabilitiesActs, asks to confirm, or hands off, based on thresholds
Scoreconfidence, probabilities and scoreSame, and checks which levels share the weight
Noulthe noul value itselfYes above a high line, no below a low line, a person in between
Low confidence on a label with a parentthe parent labelReports the broader category instead of guessing
Any, once thresholds are tunedthe model field of the responsePins that exact version

What does confidence mean in Jev?

A probability belongs to one option. Confidence describes the whole distribution: is there one clear winner, or is the weight spread around?

Here’s a Choice answer from the TypeSafe docs. The support ticket says “Shoes arrived two weeks late and in the wrong size. Also I see two charges of $120 on my card. What are you going to do about this?”, and the question asks which team should handle it:

{
  "department": {
    "type": "choice",
    "choice": "returns",
    "confidence": 0.42,
    "probabilities": {
      "shipping": 0.04,
      "billing": 0.35,
      "returns": 0.61
    }
  }
}

returns wins with 0.61, but billing keeps 0.35 because of the double charge, and confidence drops to 0.42 to tell you the ticket belongs to two teams.

TypeSafe computes confidence from the distribution it already returns, without publishing an exact formula for every case. The interactive demo on the Confidence page approximates it for three options as (3 × largest probability − 1) / 2, which gives 1.0 when one option has everything and 0 when all three sit at a third. The ticket above fits that: (3 × 0.61 − 1) / 2 is about 0.42.

TypeSafe calls confidence a default. If your domain needs a different measure, like the gap between the top two options, compute it from probabilities.

Jev is trained to return calibrated probabilities. Across many answers, outcomes given 0.8 should happen about 80% of the time. That describes groups of answers, so a single answer at confidence 1.0 can still be wrong, and your thresholds need checking against your own data.

Why doesn’t a Noul answer have a confidence score?

A Noul answers a yes/no question with one number, noul, the probability that the answer is yes. With only two outcomes, that number describes the whole distribution, so a separate confidence would add nothing.

The certainty is the distance from the middle. 0.95 is a confident yes, 0.04 a confident no, and 0.5 means the model can’t tell. If you ask whether a customer wants a refund and get 0.5, you learned that the message is ambiguous, not that the answer is “half yes”.

In TypeScript, reading .confidence on a Noul answer doesn’t even compile, because the SDK’s Noul response type only has type and noul.

So for a Noul you draw two lines: one for yes, one for no, and the space between goes to a person. The self-consistency cookbook for Nouls uses 0.30 and 0.70 as an illustration:

export function readNoul(value, { yes = 0.7, no = 0.3 } = {}) {
  if (value > yes) return 'yes'
  if (value < no) return 'no'
  return 'unsure'
}

console.log(readNoul(0.93)) // yes
console.log(readNoul(0.48)) // unsure
console.log(readNoul(0.95, { yes: 0.9 })) // yes

Move the lines with the cost of each mistake. Raise the yes line when acting on a false yes is expensive, like issuing a refund. Lower it when missing a true yes is expensive, like failing to flag a safety issue.

The middle band also absorbs run-to-run wobble. In that cookbook, TypeSafe ran the same insurance claim through Jev 15 times, and one answer moved between 0.43 and 0.53. With a single line at 0.5, the same claim would get opposite automatic decisions. With the band it goes to a person every time, although a value right on 0.30 or 0.70 could still flip between bands.

Why read probabilities and not only confidence?

confidence tells you how sure the model is, and probabilities tells you what it’s torn between.

On a Choice, a runner-up with a real share can get a copy of the work. For the shoes ticket, returns takes it and billing gets a copy, because 0.35 is over 0.25. The Choice docs use the same pattern:

const { department } = answers

for (const [team, probability] of Object.entries(department.probabilities)) {
  if (team !== department.choice && probability > 0.25) {
    notifyTeam(team, message)
  }
}

On a Score it matters more. score is the probability-weighted average of the level numbers, so different distributions produce the same number. A score of 1.0 can mean all the weight on level 1, or half on level 0 and half on level 2.

The first is a clear “broken, with a workaround”. In the second the model sees a cosmetic issue and a blocking one at once, which often means the question mixes two things or the message doesn’t say enough. Confidence is low there, and probabilities shows where the weight went:

const { answers } = await client.systemOne({
  model: MODEL,
  state: { message },
  questions: {
    severity: score('How severe is the problem in `message`?', [
      'Cosmetic; no impact on functionality',
      'Broken or degraded feature, but a workaround exists',
      'Blocking issue; no workaround exists',
    ]),
  },
})

const { probabilities } = answers.severity
const torn = probabilities['0'] > 0.3 && probabilities['2'] > 0.3
console.log(torn)

When torn is true, don’t treat 1.0 as a middle severity. Ask for more detail, or split the Score into two questions that each measure one thing.

How do I turn confidence into act, confirm, or hand off?

Let’s build this on a support inbox. Every incoming message goes to billing, technical, returns, or other. Billing messages that ask for money back also get an automatic reply with the refund form.

The questions live in one file, so they’re easy to review. MODEL pins a version (more on that below):

import { choice, noul } from '@typesafe-ai/sdk'

export const MODEL = 'jev-1.13.0'

export const QUESTIONS = {
  department: choice('Which team should handle `message`?', {
    billing: 'Charges, invoices, refunds, subscriptions',
    technical: 'Bugs, outages, login or API problems',
    returns: 'Sending an item back or exchanging it',
    other: 'Anything that fits none of the teams above',
  }),
  wants_refund: noul('Does `message` ask for money back?'),
}

Save that as questions.js, after installing the SDK with npm install @typesafe-ai/sdk (it needs Node.js 20 or newer and reads TYPESAFE_API_KEY from the environment). If you haven’t used the SDK yet, start with how to use Jev in Node.js, and for writing the labels themselves see how to classify text with Jev.

The two actions have different costs. A ticket assigned to the wrong team costs a reassignment, so 0.8 is enough to assign it automatically. Between 0.5 and 0.8 the code suggests a team and an agent confirms with one click, and below 0.5 the ticket goes to the triage queue. A refund form sent to someone who didn’t ask for money is something the customer sees, so it needs 0.9 on both the department and the Noul.

Here’s triage.js:

import { TypeSafeClient } from '@typesafe-ai/sdk'
import { MODEL, QUESTIONS } from './questions.js'

const client = new TypeSafeClient()

const ASSIGN = { act: 0.8, confirm: 0.5 }
const REFUND_REPLY = { department: 0.9, noul: 0.9 }

export async function triage(message) {
  const { answers, model } = await client.systemOne({
    model: MODEL,
    state: { message },
    questions: QUESTIONS,
  })

  const { department, wants_refund } = answers

  if (department.confidence < ASSIGN.confirm) {
    return { action: 'triage', model }
  }

  if (department.confidence < ASSIGN.act) {
    return { action: 'suggest', team: department.choice, model }
  }

  if (
    department.choice === 'billing' &&
    department.confidence >= REFUND_REPLY.department &&
    wants_refund.noul >= REFUND_REPLY.noul
  ) {
    return { action: 'send_refund_form', team: 'billing', model }
  }

  return { action: 'assign', team: department.choice, model }
}

The floor check comes first, so a low-confidence answer never reaches an action, and the thresholds sit at the top of the file where a reviewer can find them. For actions that can’t be undone, confidence should be one gate among several, and I wrote about the others in how to let an AI agent perform irreversible actions safely.

These numbers are starting points, close to the ones in TypeSafe’s confidence-gated routing pattern, which uses a 0.6 floor and 0.85 before approving a bank transfer.

The middle band steadies Choices across repeat runs too. In the self-consistency cookbook for Choices, a borderline moderation post went through 8 questions 15 times, and the picked labels repeated 90.8% of the time. With a top probability under 0.60 treated as “uncertain” (that cookbook thresholds the top probability, not confidence), decisions repeated 99.2% of the time, and 74.2% of answers were still automatic. That’s repeatability, not accuracy.

How do I pick the thresholds from labeled data?

A threshold is a trade between how often the model is right above it and how much work it leaves for people.

Start from what a mistake costs. Say a misrouted ticket costs an agent 5 minutes to read, notice, and reassign, and a manual triage costs 1 minute. Auto-assigning pays off while fewer than 1 in 5 automatic assignments are wrong, so the target is at least 80% correct above the threshold. A wrong refund reply costs more, so its target is higher.

Then collect labeled examples, past messages with the team that actually handled them, a few hundred if you can. Put them in labeled.json:

[
  {
    "message": "I was charged twice for order A-104, please refund one of the charges.",
    "label": "billing"
  },
  {
    "message": "The export button crashes the settings page in Safari.",
    "label": "technical"
  }
]

collect.js runs every message through the same questions and saves the answers next to the labels. It’s the only step that needs an API key:

import { readFile, writeFile } from 'node:fs/promises'
import { TypeSafeClient } from '@typesafe-ai/sdk'
import { MODEL, QUESTIONS } from './questions.js'

const client = new TypeSafeClient()
const labeled = JSON.parse(await readFile('labeled.json', 'utf8'))
const rows = []

for (const { message, label } of labeled) {
  const { answers, model } = await client.systemOne({
    model: MODEL,
    state: { message },
    questions: QUESTIONS,
  })

  const { choice, confidence, probabilities } = answers.department
  rows.push({ message, label, choice, confidence, probabilities, model })
}

await writeFile('answers.json', JSON.stringify(rows, null, 2))
console.log(`Saved ${rows.length} answers`)

Each row in answers.json looks like this:

[
  {
    "message": "I returned the chair two weeks ago and still have no refund.",
    "label": "returns",
    "choice": "billing",
    "confidence": 0.71,
    "probabilities": {
      "billing": 0.78,
      "technical": 0.0,
      "returns": 0.22,
      "other": 0.0
    },
    "model": "jev-1.13.0"
  }
]

thresholds.js is the part you’ll rerun often, and it makes no API calls. It groups the saved answers into confidence buckets and prints how many were right in each. Then, for each candidate threshold, it prints the share of messages handled automatically and how many of those were right:

import { readFile } from 'node:fs/promises'

const file = process.argv[2] ?? 'answers.json'
const text = await readFile(file, 'utf8')
const rows = file.endsWith('.jsonl')
  ? text.trim().split('\n').map((line) => JSON.parse(line))
  : JSON.parse(text)

const percent = (part, total) =>
  total === 0 ? '-' : `${Math.round((part / total) * 100)}%`

console.log('confidence   answers   correct')
for (let i = 0; i < 5; i++) {
  const low = i / 5
  const high = (i + 1) / 5
  const bucket = rows.filter(
    (row) => row.confidence >= low && (row.confidence < high || i === 4),
  )
  const correct = bucket.filter((row) => row.choice === row.label).length
  const range = `${low.toFixed(1)}-${high.toFixed(1)}`
  console.log(
    `${range.padEnd(13)}${String(bucket.length).padEnd(10)}${percent(correct, bucket.length)}`,
  )
}

console.log('\nthreshold   automatic   correct when automatic')
for (const threshold of [0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9]) {
  const automatic = rows.filter((row) => row.confidence >= threshold)
  const correct = automatic.filter((row) => row.choice === row.label).length
  console.log(
    `${threshold.toFixed(1).padEnd(12)}${percent(automatic.length, rows.length).padEnd(12)}${percent(correct, automatic.length)}`,
  )
}

Run it with node thresholds.js. Here’s what it prints on a file of 20 made-up answers:

confidence   answers   correct
0.0-0.2      0         -
0.2-0.4      4         25%
0.4-0.6      3         67%
0.6-0.8      3         33%
0.8-1.0      10        90%

threshold   automatic   correct when automatic
0.3         90%         67%
0.4         80%         75%
0.5         75%         73%
0.6         65%         77%
0.7         60%         83%
0.8         50%         90%
0.9         35%         100%

The top bucket is right 90% of the time and the bottom one 25%, which is the shape you want. The two middle buckets are out of order because three rows per bucket is noise. With a few hundred rows they settle down.

With the 80% target, 0.7 is the lowest threshold that passes: 60% of messages get assigned automatically, and 83% of those correctly. 0.8 gets to 90% correct but only handles half. On 20 rows I would pick 0.8 and look again once there’s more data.

Repeat this for every action with its own target. For the refund reply, label whether each customer really asked for money back and compare against the saved Noul value.

What should I do when confidence is low?

When your labels form a hierarchy, you can answer one level up instead of handing off.

TypeSafe’s classification using confidence cookbook does this with SEC annual reports: one Choice over 75 industry groups, each belonging to a broader division. A 0.9 cutoff split the 60 filings in half. The confident half was right 27 times out of 30, the other half 12 times out of 30, and reporting that second half as divisions took it to 70%. Those numbers come from jev-1.12 in August 2026.

The same idea works for a support inbox with specific queues under each team:

import { choice, TypeSafeClient } from '@typesafe-ai/sdk'
import { MODEL } from './questions.js'

const client = new TypeSafeClient()

const queue = choice('Which queue should handle `message`?', {
  refunds: 'Wants money back for a charge or an order',
  invoices: 'Needs an invoice or a receipt, or a change to one',
  plans: 'Upgrading, downgrading, or cancelling a subscription',
  bugs: 'Something in the product is broken',
  login: 'Cannot sign in, reset a password, or verify an account',
  api: 'Problems calling the API or using an integration',
  exchanges: 'Sending an item back or swapping it for another',
  other: 'Anything that fits none of the queues above',
})

const TEAM = {
  refunds: 'billing',
  invoices: 'billing',
  plans: 'billing',
  bugs: 'technical',
  login: 'technical',
  api: 'technical',
  exchanges: 'returns',
  other: 'other',
}

export async function pickQueue(message) {
  const { answers } = await client.systemOne({
    model: MODEL,
    state: { message },
    questions: { queue },
  })

  const answer = answers.queue

  if (answer.confidence >= 0.9) {
    return { level: 'queue', label: answer.choice }
  }
  return { level: 'team', label: TEAM[answer.choice] }
}

It’s still one request, because the team follows from the queue. A message torn between refunds and invoices gets a low queue confidence, but both lead to billing. If the broad label is too coarse to act on, that branch is where a person steps in.

Why should I pin the model version once thresholds are tuned?

jev-latest is an alias. As of September 2026 it points to jev-1.13.0, and it moves when a new release ships. Your 0.8 was measured against one model, and a new version can spread its probabilities differently. The Models page recommends pinning the version you tuned against and moving on your own schedule.

Pin it per request with model, as MODEL does above, or once on the client:

const client = new TypeSafeClient({ defaultModel: 'jev-1.13.0' })

The SDK also reads a TYPESAFE_DEFAULT_MODEL environment variable. Every response includes the versioned model that answered, and the scripts above save it in each row.

When a new version comes out, rerun collect.js with the new ID, run thresholds.js again, and compare the tables before you switch. Do the same after you reword a question or a criteria description, because wording moves confidence too.

How do I log answers in shadow mode before automating?

Before a threshold controls a single ticket, run Jev next to your current process and let people keep deciding. The code only logs what it would have done.

The easiest place to hook it is where a person already makes the decision: when an agent assigns a ticket to a team. At that point you have the message and the right answer together:

import { appendFile } from 'node:fs/promises'
import { TypeSafeClient } from '@typesafe-ai/sdk'
import { MODEL, QUESTIONS } from './questions.js'

const client = new TypeSafeClient()

export async function logShadow(ticketId, message, assignedTeam) {
  const { answers, model } = await client.systemOne({
    model: MODEL,
    state: { message },
    questions: QUESTIONS,
  })

  const { choice, confidence, probabilities } = answers.department
  const row = {
    ticketId,
    label: assignedTeam,
    choice,
    confidence,
    probabilities,
    wantsRefund: answers.wants_refund.noul,
    model,
    at: new Date().toISOString(),
  }

  await appendFile('shadow.jsonl', JSON.stringify(row) + '\n')
}

Call it from your helpdesk’s “ticket assigned” webhook after the assignment is saved, so a slow or failed call never gets in the agent’s way. If tickets often bounce between teams, call it when the ticket closes instead, so the label is the team that really handled it.

Every row is a labeled example you didn’t have to label by hand. After a couple of weeks of traffic, point the same script at the log:

node thresholds.js shadow.jsonl

When the table looks good, turn on the top band only: assign automatically above your threshold and leave everything else as it is today. Keep logging, and lower the threshold or widen the middle band only when the numbers keep holding.

Tagged: AI · All topics

Want me to talk about your product? You can sponsor this site.

~~~

Related posts about ai: