How to classify text with Jev
By Flavio Copes
How to classify text with Jev: pick Noul, Choice or Score, write labels that work, label comments in bulk, and test accuracy before you trust it.
To classify text with Jev, you send the text as state, list your labels as the options of a Choice question, and read back the winning label plus a confidence number. Your code decides what to do with it. There’s no prompt asking for JSON and no generated output to parse.
Jev is TypeSafe’s decision model, and instead of generating text it returns typed decisions, like a yes/no probability, one option from a list you wrote, or a position on a scale you described, each with probabilities. If you’re new to it, start with my deep dive into Jev.
The same recipe works for support tickets, reviews and emails. Our running example is comments on a programming blog, which we want to sort into help requests, corrections, feedback and spam without reading every one.
The quick answer
| What you want to know | Question type | Example |
|---|---|---|
| Is this true? | Noul | Is this comment spam? |
| Which one label, from a set with no order? | Choice, up to 255 options | help, correction, feedback, spam, other |
| Where on a scale you can describe? | Score, 2 to 10 levels | How much the commenter needs a reply |
| Which labels apply, maybe several or none? | One Noul per label | Mentions an error, mentions deployment |
How do I set up the SDK?
As of September 2026 the current model is jev-1.13.0. You create an API key in the TypeSafe console. In late September 2026 TypeSafe paused new signups because of demand, while existing accounts keep working, so check typesafe.ai for the current state. The steps are in how to get access to Jev and an API key, and the SDK itself is covered in how to use Jev in Node.js. If you work in Python, see how to use Jev with Python.
Install the official JavaScript SDK, which needs Node.js 20 or newer:
npm install @typesafe-ai/sdk
The client reads the key from the TYPESAFE_API_KEY environment variable:
export TYPESAFE_API_KEY=your_key_here
In fish, it’s set -x TYPESAFE_API_KEY your_key_here.
How do I classify one piece of text?
Let’s classify a single comment. The choice() helper takes the question and an object of labels, each with a short description:
import { choice, TypeSafeClient } from '@typesafe-ai/sdk'
const client = new TypeSafeClient()
const comment = 'I followed every step but npm run dev fails with "Cannot find module astro". What did I miss?'
const { answers, model } = await client.systemOne({
state: comment,
questions: {
kind: choice('What is the commenter doing in this comment?', {
help: 'Asks for help getting the steps in the post to work',
correction: 'Says the post is wrong or outdated',
feedback: 'Thanks, praise, or an opinion, with no request',
spam: 'Promotes an unrelated product, service, or link',
other: 'Fits none of the above',
}),
},
})
console.log(answers.kind.choice)
console.log(answers.kind.confidence)
console.log(model)
The response has this shape. The numbers are illustrative:
{
"model": "jev-1.13.0",
"answers": {
"kind": {
"type": "choice",
"choice": "help",
"confidence": 0.94,
"probabilities": {
"help": 0.95,
"correction": 0.04,
"feedback": 0.0,
"spam": 0.0,
"other": 0.01
}
}
},
"usage": { "input_tokens": 190, "output_tokens": 30 }
}
choice is the most probable label, and probabilities covers every label and sums to 1. confidence is close to 1 when one label wins clearly and drops when probability spreads across several. In TypeScript, answers.kind.choice is typed as the union of your labels, so a typo fails the type check.
Jev never sees the key kind, only the instructions and every label name with its description, so write the whole question in the instructions.
Which question type should I use?
Pick by the shape of the answer. A Noul returns noul, the probability that the answer is yes, so phrase it so a high value means yes. A Choice fits one label from a set with no order, and since each extra option costs only a few tokens, give it the full list.
When the labels do have an order, like how much the commenter needs a reply, use a Score. You write 2 to 10 levels from low to high, and the score that comes back is a probability-weighted position, like 1.4, that can land between two levels.
A Choice always picks exactly one label, so for multi-label classification ask one Noul per label. As the docs put it, a Choice is relative and settles which option wins, while each Noul is absolute and can be low for all of them. That’s what you want for tags, because a comment can have none.
Here we build one Noul per topic tag in code:
import { noul } from '@typesafe-ai/sdk'
const TAGS = {
install: 'installing or upgrading a tool or package',
error: 'an error message or a crash',
deploy: 'deploying a project to a hosting service',
pricing: 'what a product or service costs',
}
const tagQuestions = Object.fromEntries(
Object.entries(TAGS).map(([tag, topic]) => [tag, noul(`Does \`comment\` talk about ${topic}?`)])
)
const { answers } = await client.systemOne({
state: { comment },
questions: tagQuestions,
})
const tags = Object.keys(answers).filter((tag) => answers[tag].noul > 0.5)
The Nouls run in parallel against the same state, so adding tags barely changes the response time. Treat 0.5 as a starting point until you’ve measured.
How do I write labels that work?
Describe what belongs to each label and how it differs from its closest neighbor. In our list, “the command in step 2 fails” could be help or correction. The difference is where the problem lives: a help request says something doesn’t work on the reader’s machine, a correction says the post itself is wrong.
Add an other option whenever your labels might not cover every input. Jev must pick one of your options, so without an exit a weird comment gets pushed into whichever real label fits least badly.
Keep the instructions and the criteria saying the same thing. If the question asks what the commenter wants and the descriptions talk about the post’s topic, answers get worse, and TypeSafe lists that mismatch as a known weak spot.
When the state is an object, point at its fields with backticked paths like comment or post.title, so there’s no doubt about which part a question is judging.
Here’s our Choice with those rules applied:
const kind = choice('What does the author of `comment` want from the post `post_title`?', {
help: 'Asks for help because the steps in the post do not work on their machine',
correction: 'Says the post itself is wrong or outdated, like a changed command or a broken link',
feedback: 'Thanks, praise, or an opinion about the post, with no request and no correction',
spam: 'Promotes a product, service, or link unrelated to the post',
other: 'Fits none of the above',
})
const { answers } = await client.systemOne({
state: {
post_title: 'How to install Astro',
comment,
},
questions: { kind },
})
The post title helps tell spam from a relevant link. The full post stays out, since no question needs it.
When should I add examples to a label?
Start with one line per label. When the model keeps confusing two labels on comments you consider clear, turn each description into an object. The docs’ Advanced: structure page uses fields like what, not_for and examples:
const kind = choice(
{
question: 'What does the author of `comment` want from the post `post_title`?',
focus: 'Classify the main request, not every topic mentioned.',
},
{
help: {
what: 'The steps in the post do not work on their machine',
not_for: 'Saying the post itself is wrong',
examples: ['I get a 404 after step 3, what am I missing?'],
},
correction: {
what: 'The post is wrong or outdated',
not_for: 'Problems caused by their own setup',
examples: ['The --template flag was renamed, so the command in step 2 fails'],
},
feedback: {
what: 'Thanks, praise, or an opinion, with no request',
not_for: 'Anything that asks for a fix or an answer',
examples: ['Great explanation, closures finally make sense'],
},
spam: {
what: 'Promotes a product, service, or link unrelated to the post',
not_for: 'A relevant link that answers a question',
examples: ['Cheap hosting deals, check the link in my profile'],
},
other: {
what: 'Fits none of the above',
not_for: 'Comments that match another option',
examples: ['First!'],
},
}
)
The field names are yours and none are reserved. The model reads names and values together, so keep them short and use the same fields on every option.
Examples help only when they look like your real inputs. The Score docs show it on a bug report about an export button that crashes in Safari but works in Chrome. With plain string levels it scored 1.43 at 0.35 confidence. Adding the example “export fails in one browser but works in another” to the middle level moved it to 1.03 at 0.96, while an unrelated example left it at 1.43 and 0.35.
Higher confidence doesn’t prove an answer is right, though. Check the change on inputs whose correct label you know before keeping it.
How do I classify hundreds of items?
Make one request per item, with every question about that item in the same request, so they cost a single round trip. TypeSafe calls this speculative fan-out: ask everything you might need, and let the code ignore what doesn’t apply.
Here’s the full question set, reusing kind from above, and a function that classifies one comment:
import { noul, score } from '@typesafe-ai/sdk'
const COMMENT_QUESTIONS = {
kind,
has_code: noul('Does `comment` include code or an error message?'),
needs_reply: score('How much does the author of `comment` need a reply?', [
'No reply needed, the comment is complete on its own',
'A reply would be nice, but they are not blocked',
'They are stuck and waiting for an answer to continue',
]),
}
async function classifyComment({ id, postTitle, text }) {
const { answers, model } = await client.systemOne({
state: { post_title: postTitle, comment: text },
questions: COMMENT_QUESTIONS,
})
return {
id,
model,
kind: answers.kind.choice,
confidence: answers.kind.confidence,
hasCode: answers.has_code.noul > 0.5,
needsReply: answers.needs_reply.score,
}
}
To run many of them, cap how many requests are in flight. As of September 2026 the limits are 1,200 requests per minute and 250,000 tokens per second, adjusting dynamically during early access. Above them you get a 429, which the SDK retries with backoff (twice by default, honoring retry-after). This worker pool raises the retries and keeps five requests running at a time:
const client = new TypeSafeClient({ retry: { maxRetries: 5 } })
async function classifyAll(comments, concurrency = 5) {
const results = []
let next = 0
async function worker() {
while (next < comments.length) {
const comment = comments[next++]
results.push(await classifyComment(comment))
}
}
await Promise.all(Array.from({ length: concurrency }, worker))
return results
}
Throughput is about concurrency divided by latency, so five workers at 300 milliseconds per call make roughly 1,000 requests a minute, just under the limit. If the SDK still throws a RateLimitError after its retries, lower the concurrency. Results arrive in completion order, which is why each one carries its id.
Input costs $0.042 per million tokens and output is free, so 10,000 comments at about 300 input tokens each come to 3 million tokens, around 13 cents. The Jev pricing post shows how to estimate other workloads.
For a short list there’s a shortcut: put the list in the state and ask one Noul per item in one request. The docs use the same trick for counting:
const comments = [
'Thanks, this fixed my Docker build!',
'Cheap hosting deals, check the link in my profile',
'Is there a version of this for Deno?',
'Buy followers fast, DM me',
]
const spamQuestions = Object.fromEntries(
comments.map((_, i) => [`comment_${i}`, noul(`Is \`comments[${i}]\` spam?`)])
)
const { answers } = await client.systemOne({
state: { comments },
questions: spamQuestions,
})
const spam = comments.filter((_, i) => answers[`comment_${i}`].noul > 0.8)
Keep this for a handful of short items. Every question sees the whole list, the request must fit in 64k tokens (32k for the state plus the longest question), and accuracy drops as unrelated content fills the state. For a real backlog, stick to one request per item.
What if I have hundreds of labels?
One Choice takes up to 255 options, and TypeSafe’s classification using confidence cookbook says it works reliably up to roughly 240. That cookbook sorts SEC annual reports into 75 industry groups with a single question.
Past that, or when your labels form a tree like product categories, walk the tree: one Choice for the top level, then one among the winner’s children, until you reach a leaf. This is one of the few cases where a second request makes sense, because you can’t write the second question before you know the first answer.
The taxonomy below would fit in one Choice, but it shows the walk. Each top-level option is described by its children, so the model sees what lives under a branch before committing to it:
const TAXONOMY = {
javascript: ['node', 'react', 'astro', 'typescript'],
css: ['flexbox', 'grid', 'tailwind'],
devops: ['docker', 'cloudflare', 'linux'],
other: [],
}
async function classifyTechnology(comment) {
const first = await client.systemOne({
state: { comment },
questions: {
area: choice('Which area of web development is `comment` about?', TAXONOMY),
},
})
const area = first.answers.area
if (area.choice === 'other' || area.confidence < 0.5) {
return { level: 'none', label: null }
}
const second = await client.systemOne({
state: { comment },
questions: {
technology: choice(
'Which technology is `comment` about?',
Object.fromEntries(TAXONOMY[area.choice].map((name) => [name, null]))
),
},
})
const technology = second.answers.technology
if (technology.confidence < 0.8) {
return { level: 'area', label: area.choice }
}
return { level: 'technology', label: technology.choice }
}
A null description means the label name speaks for itself. The last if comes from the same cookbook: when the model is unsure about the narrow label, report the broader one. In its test on 60 filings, answers at 0.9 confidence or higher were right 90% of the time. The rest were right 40% of the time as industry groups, and 70% of the time as the broader division.
This walk is greedy, so one early mistake can’t be undone. The hierarchical classification cookbook adds beam search: at each level it keeps the best three paths, asks the next Choice for each of them in parallel, and ranks paths by the geometric mean of their Choice probabilities, so shallow and deep paths compare fairly. On its four test documents, greedy search found the expected leaf twice and beam search four times. That’s a small sample, and both cookbooks ran on the older jev-1.12.
How do I know the labels are accurate?
Build a small test set of real comments you label by hand, with every label represented and a few hard cases on the boundaries. I would start with 50 to 100.
Run it with the model pinned to a version, so a new release behind jev-latest doesn’t shift your numbers, and check accuracy at a few confidence thresholds:
const TEST_SET = [
{ postTitle: 'How to install Astro', comment: 'The create command now asks for a template name, the post skips that step', expected: 'correction' },
{ postTitle: 'How to install Astro', comment: 'Getting EACCES when I run npm install -g, any idea?', expected: 'help' },
{ postTitle: 'CSS Grid tutorial', comment: 'Clear and short, thanks!', expected: 'feedback' },
{ postTitle: 'CSS Grid tutorial', comment: 'We build websites for dentists, visit our page', expected: 'spam' },
]
const results = []
for (const item of TEST_SET) {
const { answers } = await client.systemOne({
model: 'jev-1.13.0',
state: { post_title: item.postTitle, comment: item.comment },
questions: { kind },
})
results.push({ ...item, got: answers.kind.choice, confidence: answers.kind.confidence })
}
for (const threshold of [0.5, 0.7, 0.9]) {
const kept = results.filter((r) => r.confidence >= threshold)
const right = kept.filter((r) => r.got === r.expected)
console.log(`>= ${threshold}: labeled ${kept.length}/${results.length}, correct ${right.length}/${kept.length}`)
}
console.log(results.filter((r) => r.got !== r.expected))
The list of mistakes at the end is the most useful output. Each wrong answer points at a description to sharpen, or at a test label you got wrong yourself. Change one thing at a time and rerun the set to see if it helped.
Then pick the threshold where accuracy is good enough for what happens next, starting conservative as the docs suggest. Hiding spam automatically needs a higher bar than sorting a review queue, and everything below it goes to a person.
What are the common mistakes?
Most of them are on TypeSafe’s jaggedness page for jev-1.13.
The first is asking Jev for math. It doesn’t count reliably, it reads dates as text, and it isn’t a calculator. “Does this comment have more than two links?” is a regular expression, and “was it posted this week?” is date math in code. To count items that match a meaning, ask one Noul per item and add them up in code, like the spam list above.
Score levels written as degrees are next. Each level is judged on its own, without its number or its neighbors, so “somewhat” and “very” give the model nothing to match. In the docs, levels that were only “0”, “1” and “2” scored a misaligned-button report 0.55 at 0.33 confidence, while descriptive levels scored it 0.0 at 1.0. Describe situations, like “They are stuck and waiting for an answer to continue”.
Then there are questions that measure two things. “Is this comment polite and on topic?” needs two Nouls combined in code, and a Score level that mixes qualities has the same problem.
Don’t stuff the state either. Accuracy falls as it fills with content the question doesn’t need, so send the comment and the post title, not the whole post and its thread.
Watch out for a Noul where a high value means “no”, like “Is this comment free of spam?”, which confuses the model and whoever reads the code later. And a threshold tuned on a Noul doesn’t carry over to a yes/no Choice, because the two don’t return comparable numbers.
Finally, English is Jev’s primary language. Other languages work less well, so test on your own non-English comments first.
Where do I go from here?
Open the Playground and try your labels on a few real comments before writing any code. Then hand-label your test set, run the script above, and only let the classifier act on its own above a threshold you’ve measured.
Want me to talk about your product? You can sponsor this site.
Related posts about ai: