NewsAnalysis
What Jev is actually for: seven jobs for a model that only decides
Two weeks after launch, TypeSafe's own docs and cookbooks show where a typed, calibrated decision model beats an LLM call, a classifier or plain rules, and where it plainly does not.

Two weeks after TypeSafe AI released Jev, the interesting question is no longer what it is but where it belongs in a working system. The answer, read from TypeSafe’s own documentation and cookbooks, is narrow and useful: Jev earns its place wherever code needs a judgement about messy text, the possible answers can be listed in advance, and the system needs to know how sure the judgement is before acting on it.
That covers a lot of production plumbing. Agent routing, ticket triage, moderation, guardrails, risk flags, checking another model’s output and bulk classification all reduce to “pick one of these, score this, or tell me whether this is true”, jobs now done with an LLM call that returns JSON, a trained classifier or a pile of rules. This piece takes each in turn, with worked examples on Jev’s documented API, TypeSafe’s published numbers and an account of where a decision model is the wrong choice. The short version: the price is a headline, but the confidence score is the product.
What Jev returns, briefly
Jev launched in early access on 15 September 2026 alongside a reported US$40 million seed round led by DCVC; the launch coverage has the funding and architecture detail. For this article only the interface matters.
A request sends a state (a string, JSON object or array of text) and a map of named questions to POST /v1/systemone, according to the API reference. There are three question types. A Choice picks one option from up to 255 you define and returns the choice, a probability for every option and a confidence. A Score rates the state against up to ten ordered, described levels and returns a probability-weighted score, the distribution and a confidence. A Noul answers a yes-or-no question and returns the probability that the answer is yes, with no separate confidence because the probability already says everything.
Confidence is derived from the shape of the distribution: per the confidence page, the number of options times the peak probability, minus one, over the number of options minus one. A winner at 0.90 over three options gives 0.85. TypeSafe says the model is trained with a method it calls Reinforcement Learning for Calibrated Decisions so that “higher confidence means higher accuracy”, and its docs add the important caveat that calibration “is measured across groups of predictions; it does not guarantee that an individual answer is correct”.
As of September 2026, per the models page, jev-1.13.0 costs $0.042 per million input tokens with output free, takes 64,000 tokens per request (32,000 for the state plus the longest question), accepts text only, and is limited to 250,000 tokens per second and 1,200 requests per minute, limits TypeSafe says “are adjusting dynamically”. The launch post claims 70 to 500 ms end to end; the docs say most queries take about 100 ms. Questions in one request run in parallel, so adding one “barely changes the response time”.
Rules, classifier, LLM or Jev: the order to ask in
TypeSafe’s own build guide starts with “use code when you can”, and that is the right first question. If a decision can be computed exactly (an invoice is 30 days overdue, an amount exceeds a limit), code is free, instant and deterministic.
The second question is whether the output is prose. A customer reply, a summary, a rationale for an auditor: those need a generative model. Jev cannot produce them.
The third is whether a trained classifier fits: a fixed label set, plenty of labelled examples and very high volume favour a small fine-tuned model on owned hardware, with near-zero marginal cost and no external rate limit.
Jev fits the remaining space: inputs too fuzzy for rules, labels that change with policy, no training set yet, decisions that need instructions and context, and a need to know when the answer is shaky. Jev is not fine-tuned on customer data; domain knowledge goes into the state and question criteria. That helps when labels change weekly and hurts when a team has years of labelled data a classifier could learn from.

Figure 1: Ask the questions in this order. Jev is the answer to the last one, not the first.
Agent control flow and tool routing
The most common use TypeSafe documents is putting Jev in front of expensive or risky steps so that code, not an LLM loop, owns the control flow. The intent routing pattern classifies a customer message into four intents with a Choice and rates its complexity with a three-level Score, in one call. Order-status questions go to deterministic lookup code with no LLM involved; product and returns questions go to two different specialist LLMs; complaints go to a human if the complexity score exceeds 1 or its confidence is under 0.5. Any intent below 0.5 confidence goes to a person.
The confidence-gated routing example makes the stakes explicit with a voice banking assistant. A Choice decides between check_balance, approve_transfer and other. Below 0.6 confidence on anything, the call goes to a support agent. A balance check proceeds at 0.6, because the worst outcome is reading out the wrong screen. A transfer approval proceeds automatically only above 0.85; between 0.6 and 0.85 the assistant asks the user to confirm. Same model, same answer type, three different bars set by the cost of being wrong.
The function-calling cookbook maps trading requests onto ten Python functions with one Choice per closed-set argument; free-text, numeric and date arguments get no question, which is the approach’s honest limit. The skill-suggestion cookbook is more striking. An agent built on claude-haiku-4-5 choosing among 182 skills from Nous Research’s Hermes catalogue loaded the wrong skill on 16.8 percent of 488 requests and loaded a skill when none fitted on 9.8 percent. With two Jev requests per turn (one Choice to rank all 182, one pass to re-check the top three with full descriptions), those fell to 7.3 and 4.0 percent. Handing the agent the right answer only got them to 2.5 and 1.2 percent, so Jev closed most of the gap to that floor. These are TypeSafe’s numbers on TypeSafe’s test set.
Why Jev here. Rules cannot read “I want to move money to my sister” as a transfer. A classifier could, but tool sets change every release and retraining is what pushed teams to LLM routing. An LLM router adds latency to every turn and returns a label with no usable measure of doubt. The Register quoted TypeSafe’s Eugene Shvarts framing the design question as: “If many useful semantic judgments were affordable within our application’s response-time budget, what would we design differently?” Routing is where that question bites first.
Support ticket triage
Triage is the cleanest fit: every ticket needs several small, independent judgements. The Choice page walks a five-question triage request: department, return reason, shipping issue, requested resolution and tone, all in one call. Its routing code sends a ticket to manual triage if department confidence is under 0.3, copies a second team if that team holds more than 0.25 of the probability, and asks the customer what they want if the resolution confidence is under 0.5. On its example ticket, the returns team gets it, billing is copied at a 0.35 share, and the customer is asked to clarify at a resolution confidence of 0.20.
A worked example in that shape, using the documented request format (the ticket text and labels are illustrative):
{
"model": "jev-1.13.0",
"state": {
"ticket": "Third time asking. I was charged twice for order A-104 and nobody has replied.",
"customer_tier": "business"
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle `ticket`?",
"criteria": { "billing": "Charges, refunds, invoices", "shipping": "Delivery and tracking", "account": "Login and settings", "other": null }
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer in `ticket`?",
"criteria": ["Calm", "Mildly annoyed", "Clearly frustrated", "Angry or threatening to leave"]
},
"wants_human": {
"type": "noul",
"instructions": "Is the customer asking to speak to a person?"
}
}
}
Code does the rest. The Noul page suggests YES = 0.8 and NO = 0.2, with anything between sent to a reviewer; in its example “Can I please just talk to a real person?” scores 0.99 on wanting a human, while “Are you a bot?” lands at 0.40, exactly what a middle band is for. The request pins a versioned model ID because, per the models page, the jev-latest alias moves with each release.
Cost, worked through. Assume about 1,000 input tokens per call for the ticket, context and four questions. At $0.042 per million, that is about $0.000042 per ticket, or roughly $42 per million tickets, as of September 2026. On gpt-5.4-mini ($0.75 input, $4.50 output per million, as priced in TypeSafe’s SDE cascade cookbook) with 150 output tokens, the same call is about $1,400 per million. On GPT-6 Luna ($0.10 and $0.50, per GenAI Brief’s inference pricing tracker) it is about $175. Against the cheapest small model Jev is four times cheaper, not four hundred. The large multiples in TypeSafe’s launch post (“193.6x faster, 444.6x cheaper” on its workflow evaluations) are measured against frontier models, and the company itself describes them as “the higher end of real world gains”.
So the stronger arguments for triage are the per-question probabilities and a response time that fits inside a web request. The real cost of an AI product is rarely the classification line; nobody should migrate for price alone.
Content moderation and LLM guardrails
Moderation is where calibrated confidence stops being a nicety. The self-consistency cookbook for choices ran an eight-question rubric over one deliberately borderline post fifteen times per condition. Jev repeated its plurality labels 90.8 percent of the time, against 87.5 to 100 percent for the LLM settings, and flipped on two questions. It did not simply win. What changed the picture was a rule: below a top probability of 0.60, label the answer uncertain and send it to a person. Agreement then rose to 99.2 percent, with 74.2 percent of answers handled automatically. The same cookbook measured a mean round trip of 114 ms for the Jev call against 826 ms to 13.0 seconds for the LLM conditions, which included claude-haiku-4-5, gpt-5.4-mini, gpt-5.5 and claude-opus-4-8.
That is the pattern in miniature: the model is not claimed to be right on hard cases, only to know which cases are hard.
The LLM guardrails cookbook applies the same idea to both sides of a chatbot. Each message gets one call with four hazard Nouls (jailbreak, harm or crime, diagnosis or dosage, self-harm) and a four-level severity Score. Under the “strict” policy a hazard at 0.35 goes to review, at 0.70 it triggers its action (block, review or a crisis-support path), and severity of 2.0 turns a review into a block; “permissive” raises the action bar to 0.85. In the cookbook’s run on jev-1.12, a real jailbreak prompt scored 0.98 and was blocked, a novelist asking how a detective would describe a poisoning passed, and a self-harm message scored 0.96 and went to support rather than a refusal. A borderline jailbreak at 0.74 is blocked under strict and reviewed under permissive, from the same probabilities.
Why Jev rather than an LLM judge. The cookbook’s own argument is that a second LLM in front of the first costs a full call’s latency and money every turn, and “an attacker can talk that one past too”. The first half is plainly true. The second needs a counterweight from TypeSafe’s own jaggedness page, which says the current model “does not treat [state] as hostile by default” and that adversarial content “can move the answer”. A guardrail built on Jev is cheaper and faster, not immune. It should sit alongside, not replace, provider-side safety and rules for known patterns.
Fraud, risk and compliance flags
TypeSafe’s use-case map lists financial crime (transaction narratives, KYC documents, entity matching), insurance claims (fraud indicators, escalation to an adjuster) and compliance (missing clauses, prohibited claims). The Cloudflare model page for Jev uses refund processing, support routing and account-risk scoring as its three worked examples. No named customer has published results in these areas, so what follows is an illustration built on the documented pattern, not a reported deployment.
A payments team triaging alerts might send the alert narrative, the customer’s recent transaction descriptions and the relevant policy excerpt as state, and ask: a Noul for “Does the narrative describe a purchase the customer says they did not make?”, a Noul for “Does the merchant description match the customer’s stated business?”, a Score for evidence quality from “no supporting detail” to “specific and consistent”, and a Choice for the alert type. Code combines them: strong evidence of an unrecognised purchase goes to the fast-track queue, any mid-band Noul goes to an investigator, and nothing closes automatically without both a low risk probability and a confident evidence Score.
Two documented details shape how that should be built. First, amounts and dates stay in code. The jaggedness page is blunt that Jev “is not a calculator”, cannot reliably compare dates, and counts badly; a rule computes “three transactions over the limit in 24 hours” and passes the result in as a field. Second, compliance checks are exactly where parallel questions pay off. TypeSafe ran a 13-question regulatory briefing (eight Nouls, two Choices, three Scores) over the 54,000-character Wikipedia article on the GDPR and found that one batched call was 12.2 times cheaper and 10 times faster than 13 single-question calls, with no change in the answers. The document dominates the token count, so asking twenty questions of one contract costs little more than asking one.
Why not a classifier. Fraud teams usually have one, trained on labelled chargebacks. Jev belongs upstream of it, turning free-text narratives into probabilistic features; an autoresearch cookbook feeds Jev probabilities into a CatBoost regressor. Replacing a validated risk model with a general decision model would hurt auditability.
Judging and verifying LLM output
The widest use may be checking other models, which the use-case map calls “universal verification”: prompts, extractions, reasoning traces and tool calls, at a fraction of the cost of the call being checked.
In the SDE cascade, gpt-5.4-mini extracts structured fields from web pages, Jev asks one Noul per field (“is this value absent from the source?”, “was it lifted from unrelated text?”), and the record escalates to gpt-5.5 at high reasoning effort if any field flag exceeds 0.7. It catches a schema-valid fabricated field that JSON Schema could not. Across 100 prompts, TypeSafe reports the cascade beat every single model on cost for a given quality, reaching most of the reasoning model’s quality (about 0.81 at about $0.10 per extraction alone) for a fraction of the spend. TypeSafe labels these “internal results” and notes the chart has not been recalculated at current Jev pricing.
The citation-check cookbook splits the job the way the docs recommend: an ordinary string match catches quotes missing from the source, then one Choice per surviving citation reads the context and returns verified, unsupported or contradicted. On eight citations against RFC 7519, the four accurate ones came back verified at confidence 0.93 or higher, and all four planted failures were caught, two of them routed to a human. Eight citations is a demonstration, not an evaluation.
The RAG passage cookbook puts four Nouls between retrieval and generation (relevant, usable, contradicts the query’s assumption, tries to instruct the model) so that passages arrive as evidence, as flagged conflict, or not at all. The re-ranking cookbook sorts BM25 shortlists by a per-pair Noul on 40 legal queries and lifts top-1 accuracy from 5 to 18 percent and top-10 from 38 to 62 percent.
Why Jev rather than an LLM-as-judge. An LLM judge produces a score and a paragraph, is slow enough that teams sample rather than check everything, and tends to be overconfident. A decision model is cheap enough to check every output and returns a probability that can gate an action. The trade-off is that Jev gives no explanation of why it flagged something, which matters for evaluation work where the explanation is the point, and it inherits every limit on literal reading and indirection. Vendor numbers are marketing until they are reproduced, and none of these cookbook numbers has been reproduced independently yet.
Classification and data validation at scale
TypeSafe pitches “AI map-reduce over big data”: classify, score and extract features across datasets too large to put through an LLM. The best-evidenced example is the classification-using-confidence cookbook. It classifies 60 SEC annual reports into 75 industry groups with one Choice each. A confidence cutoff of 0.9 split the filings roughly in half: the confident half was right 90 percent of the time, the other half 40 percent. Reporting the uncertain half one level up the SIC hierarchy (its division rather than its group) lifted those to 70 percent, at no extra calls. The three least confident filings were readable as hard cases: two development-stage companies describing businesses they had not yet started, and one that had sold one of its two segments weeks before filing.
That is the most useful single result in the documentation, because it shows confidence tracking real difficulty on a real public dataset, and it shows a design pattern (fall back to a coarser label, not to a person) that most classification pipelines could copy. The sample is 60 documents from jev-1.12, so it is a demonstration of the mechanism rather than a measured accuracy figure.
For data validation the documented shapes are entity alignment (which of 450 candidate product pairs match) and value extraction, where regex finds candidates and Jev picks the right span. Jev never produces the value; it chooses among values code has found.
The practical constraint is throughput, not price. At 1,200 requests per minute, the published limit as of September 2026, one question set per row would take about four weeks to cover 50 million rows. The docs’ answer is to pack many items into one request as separate keyed questions, which the counting example on the jaggedness page does with one Noul per list item, and to ask sales for higher limits. Anyone planning a backfill should do that arithmetic before choosing Jev over a batch LLM job or a self-hosted classifier.
Model routing
The use-case map lists building a custom router that decides which LLM receives each prompt: classify intent and domain, estimate difficulty and risk, and escalate to a more expensive model when needed. TypeSafe has not published a routing benchmark, so this remains a documented intended use rather than a measured one.
As an illustration: a Choice over domains, a Score for difficulty from “lookup” to “multi-step reasoning” and a Noul for “does this need current information?”, with code mapping the combination to a model. A low-confidence difficulty estimate is itself a reason to pick the stronger model. Teams already running traffic through an AI gateway would put this decision in front of the gateway’s own routing rules rather than replace them.
Use cases, question types and thresholds at a glance
| Use case | Question types | Documented threshold or rule | Below the bar |
|---|---|---|---|
| Agent intent and tool routing | Choice (intent), Score (complexity) | Act at 0.5 to 0.6 confidence; irreversible actions above 0.85 | Human, or ask the user to confirm |
| Support triage | Choice (team, resolution), Score, Nouls | Team confidence under 0.3 to manual triage; copy a team above 0.25 share; Noul review band 0.2 to 0.8 | Manual triage or clarify with customer |
| Content moderation | Choice per rubric question | Top probability at least 0.60 | Label uncertain, human review |
| LLM input and output guardrails | Nouls per hazard, Score (severity) | Review at 0.35; act at 0.70 (strict) or 0.85 (permissive); severity 2.0 blocks | Human review or crisis path |
| Extraction verification | Noul per field | Escalate if any field flag exceeds 0.7 | Reasoning model re-extracts |
| Citation checking | Choice (verdict) | Accurate citations returned at 0.93 or higher | Human checks the citation |
| Taxonomy classification | Choice (up to 255 options) | Report fine label at 0.9 or higher | Report the parent category |
| Fraud and compliance flags | Nouls, Score, Choice | None published; tune on labelled alerts | Investigator or counsel |
| Model routing | Choice (domain), Score (difficulty) | None published | Send to the stronger model |
Every threshold above comes from a TypeSafe example tuned on that example’s data. None is a default to copy. The docs say so directly: “Start with conservative thresholds, test with your own data, and adjust as you observe results.”

Figure 2: The same answer can be acted on, confirmed or escalated. Confidence and the cost of a mistake decide which.
Using confidence without fooling yourself
Calibration is a population property. If Jev is well calibrated, answers at 0.9 confidence are right about 90 percent of the time across many cases. It says nothing certain about the next case. The SEC cookbook’s confident half was 90 percent accurate, which still means one confident answer in ten was wrong. Systems need a path for confident mistakes, usually sampling auto-handled cases for audit.
Thresholds scale with stakes, not with the model. The confidence page puts it as “a confidence threshold is not one number”. A read-only action can proceed at a lower bar than a destructive one inside the same product.
Do not carry thresholds across question types. The jaggedness page shows the same refund question asked as a Noul (0.22) and as a yes/no Choice (0.01 yes, 0.99 no) on the same ticket, and a question and its negation as two Nouls summing to 1.19. A threshold tuned on one formulation does not transfer to another.
Pin the version. jev-latest moves with releases. A threshold tuned on jev-1.13.0 should be re-tuned before moving to the next version, and the response’s model field records which version answered.
Use three bands. Act above a high bar, confirm or review in the middle, hand off below a floor. Where a hierarchy exists, a coarser answer beats a person.
Where Jev is the wrong tool
TypeSafe’s jaggedness page, last reviewed on 17 September 2026, is unusually candid, and it defines most of this list.
- Anything that generates. Replies, summaries, code and explanations need an LLM. Jev can be forced to spell out text by chaining Choices, and the docs say “this will not work well and will be very slow”.
- Arithmetic, counting and dates. Keep them in code. Score outputs are also “weak in numerical calibration”, so interpolating a number between two rubric levels is unreliable.
- Compound or multi-hop questions, and large noisy state. Both cost accuracy. Decompose and filter first.
- Adversarial input with no other defence. Injected instructions can move answers, as noted above.
- Non-text and non-English workloads. Input is text only, and the models page says English is where accuracy is currently best, with other languages “handled but not equally well”.
- Stable, high-volume labels with training data. A small trained classifier is cheaper per call, has no external rate limit and can run where the data lives.
- Cases needing an explanation. Jev returns probabilities, not reasons.
- Throughput beyond the current limits. Rate limits are provisional, and TypeSafe documents no self-hosted option.
The evidence gap belongs on this list too. Almost every number in this article is TypeSafe’s, run on TypeSafe’s examples. The Register quoted the YouTuber Mo Bitar putting the open question simply: “I know it’s fast. I know it’s cheap, but is it good?” Two weeks in, GenAI Brief has found no independent evaluation of Jev on standard classification or moderation benchmarks.
How to adopt it without regret
- Find the LLM calls that return a label: prompts ending in “respond with one of” or “return JSON with a boolean”. Anything returning prose stays put.
- Build a labelled set of a few hundred real cases per decision, hard ones included. Without it, thresholds are guesses.
- Decompose. Atomic Choices, Scores and Nouls; arithmetic and dates in code; only the state each question needs.
- Run in shadow. Call Jev beside the existing path, log both, and plot accuracy against confidence. If accuracy does not rise with confidence on your data, stop.
- Set three bands per action, stricter for irreversible ones, and define the fallback for each: human queue, confirmation step, coarser label or bigger model.
- Pin the model version and re-check thresholds before moving to a new one.
- Route it through the same controls as other model traffic. Bifrost, the AI gateway from Maxim AI, exposes TypeSafe’s native API under a
/typesafeprefix so the official SDKs work by changing the base URL, with per-key budgets, rate limits, logging and cost tracking applied, and retries on 429 and 529 responses. Decision calls are frequent and small, which is exactly the traffic that escapes budget tracking when it is called directly. - Sample the automated band. Audit a fixed share of high-confidence decisions every week to catch confident mistakes.
The judgement, two weeks in: Jev is not a cheaper LLM, and teams that treat it as one will be disappointed by how little it can do. It is a new kind of component for the parts of a system that currently ask an LLM to act like a function. Where those calls sit in a hot path, run at volume, or need to know when to hand off, the documented patterns are strong enough to justify a shadow trial now. Where they do not, the right move is to wait for independent evaluations and for the rate limits to settle.
Sources
- Introducing System One models and Jev TypeSafe AI
- Introduction TypeSafe AI docs
- Example use cases TypeSafe AI docs
- Confidence TypeSafe AI docs
- Confidence-gated routing TypeSafe AI docs
- Intent routing TypeSafe AI docs
- Models TypeSafe AI docs
- Jev 1.13 jaggedness TypeSafe AI docs
- API reference TypeSafe AI docs
- Noul TypeSafe AI docs
- Choice TypeSafe AI docs
- Guardrails for LLMs TypeSafe AI docs
- Self-consistency: choices TypeSafe AI docs
- Classification using confidence TypeSafe AI docs
- SDE cascade TypeSafe AI docs
- Skill suggestion TypeSafe AI docs
- Double-checking citations TypeSafe AI docs
- Parallel questions TypeSafe AI docs
- Re-ranking TypeSafe AI docs
- Jev (typesafe) model page Cloudflare
- TypeSafe SDK integration Bifrost docs
- Shut up and calculate: Jev's new AI primitives for coders The Register
Frequently asked questions
What is Jev used for?
Jev is used where software needs a fast judgement over unstructured text with a bounded answer: routing requests and tool calls, triaging support tickets, moderating content, screening LLM inputs and outputs, verifying extractions and citations, and classifying large datasets. It returns typed answers with probabilities rather than text.
Can Jev replace an LLM?
No. Jev does not generate text, write code or hold a conversation, and TypeSafe's docs say it is not a drop-in model for coding agents. It replaces the part of an LLM workflow where the LLM is asked to pick a label or return a yes or no, and it can decide when a bigger model is needed.
What confidence threshold should I use with Jev?
There is no single number. TypeSafe's examples use a floor of 0.5 to 0.6 below which a person decides, a higher bar such as 0.85 for irreversible actions like approving a transfer, and 0.9 for reporting a fine-grained label. Tune thresholds on your own labelled data and pin the model version they were tuned on.
How much does Jev cost per decision?
As of September 2026 Jev is priced at $0.042 per million input tokens with free output tokens. A triage call carrying about 1,000 tokens of ticket and questions therefore costs about $0.00004, or roughly $42 per million tickets, before any volume discount or rate-limit constraint.
Is Jev better than a fine-tuned classifier?
Not for a fixed label set with plenty of training data and high volume, where a small trained classifier is cheaper per call and can run on your own hardware. Jev wins when labels change often, when there is no labelled data yet, or when the decision needs instructions and context that a classifier cannot take.
How fast is Jev?
TypeSafe reports end-to-end responses of 70 to 500 ms, and its docs say most queries complete in about 100 ms. In one TypeSafe cookbook, an eight-question moderation rubric averaged 114 ms per call, against 826 ms to 13 seconds for the LLMs it was compared with.
What is Jev bad at?
TypeSafe's jaggedness page lists literal reading of instructions, arithmetic, counting, date comparison, multi-hop indirection, large states full of irrelevant detail, adversarial content and text generation. It also accepts text only and is best in English.


