IndustryExplainer
What are System One models, and will the category outlast Jev?
TypeSafe coined the term for models that return typed, probability-weighted decisions instead of text. Where the name comes from, how the mechanism differs, and what the economics mean.

A System One model is a model that answers questions instead of writing text. It takes unstructured input, a set of questions whose possible answers are fixed in advance, and returns a typed answer to each with a probability distribution attached, in a single pass. The term was coined by TypeSafe AI, whose model Jev went into early access on 15 September 2026. It is the first commercial product in what TypeSafe presents as a new class of model. Whether it is a class or a product depends on things that have not happened yet.
The question matters to engineers who pay LLM prices, and LLM latency, for decisions that are really classifications: route this ticket, flag this transaction, block this prompt. It matters to model providers and classifier vendors, who now face a product aimed at that traffic. Investors have already reacted, on reports that TypeSafe discussed a valuation roughly fifty times its seed price within ten days of launch.
This explainer covers the idea rather than the launch. The funding and product details are in the launch story, and worked applications are in the use-case guide. What follows is the definition, the mechanism, calibration, the economics and the case against.
The definition, in TypeSafe’s words
TypeSafe’s founder, Diogo Almeida, defined the term in the launch post. A System One model is “a new class of frontier models built to make fast, structured decisions that software can use directly.” The same post compresses Jev into one line: “Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.”
The documentation fills in the contract. A caller sends a state, which is text, a JSON object or an array of text, and a set of questions. Each question is one of three primitives. A Choice picks one option from a list you define and returns a probability for every option. A Score rates the state against ordered levels you describe, such as calm, frustrated and very frustrated. A Noul, TypeSafe’s name for a yes/no question, returns the probability that the statement is true. Choice and Score answers also carry a confidence value. All the questions are evaluated “in parallel and in isolation against the same state in one go”, in the docs’ wording. A single endpoint, POST /v1/systemone, serves them.
What the definition excludes matters as much. The concept page is blunt: “System One models do not write replies, produce code, or generate explanations of their reasoning.” The launch post turns the limitation into a pitch. Jev “gives up string generation” and, because every possible answer is enumerated in advance, “can’t hallucinate” in the sense of returning something outside the schema.
Two claims define the class, and they are worth keeping apart. The first is structural and falsifiable with one counter-example: outputs are always one of the declared values. The second is statistical and much harder to check: the probabilities are calibrated, meaning “higher confidence means higher accuracy.”
Where the name comes from, and where it stops fitting
The TypeSafe FAQ says the name draws on Daniel Kahneman’s Thinking, Fast and Slow and “the distinction between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning.” The distinction predates the 2011 book. In his 2002 Nobel lecture Kahneman credited the labels to the psychologists Keith Stanovich and Richard West. He summarised them this way: “The operations of System 1 are fast, automatic, effortless, associative, and difficult to control or modify. The operations of System 2 are slower, serial, effortful, and deliberately controlled; they are also relatively flexible and potentially rule-governed.”
How AI borrowed the vocabulary
Machine learning adopted the terms well before TypeSafe. Yoshua Bengio’s invited talk at NeurIPS 2019, “From System 1 Deep Learning to System 2 Deep Learning”, argued that deep learning had mastered fast, perceptual System 1 tasks and lacked the slow, compositional System 2 kind. When reasoning models arrived, the label stuck to them. A 2025 survey of reasoning LLMs by Li and colleagues is titled “From System 1 to System 2”. It describes foundational LLMs as excelling “at fast decision-making” while reasoning models such as o1 and R1 are “closely mimicking the deliberate reasoning of System 2”.
In practice, the industry’s System 2 is a training recipe plus a token budget. As the explainer on reasoning models describes, a reasoning model is a standard LLM post-trained with reinforcement learning to produce a long chain of thought, then allowed to spend thinking tokens at inference. The labs’ roadmaps bent around that axis, as covered in how test-time compute changed the roadmap. For two years the direction of travel was towards more thinking, more tokens and more latency.
TypeSafe’s move is to name the other half and sell it as a separate product. That is a genuine strategic idea. If the frontier labs are optimising for the slow system, a company can specialise in the fast one.
Three places the analogy breaks
The psychology should not be read too literally into the product, for three reasons.
First, in Kahneman’s work System 1 is where the errors live. The heuristics-and-biases programme is largely a catalogue of fast, intuitive judgements going wrong. TypeSafe acknowledges this in its own FAQ, noting that “System 1 thinking” has “also implied error-prone”, and says it believes System One models “can be made more reliable than its alternatives.” That is a claim about a product, not a finding borrowed from psychology. The name brings the speed from Kahneman and leaves the fallibility behind.
Second, “intuitive” describes a human experience, not a mechanism. TypeSafe’s docs describe the right workload as “gut-check” questions “a highly knowledgeable person could make in a few seconds”. That is a guideline for question size, not a theory of mind.
Third, Kahneman’s two systems share one head. In software the handoff is written by an engineer: TypeSafe’s docs recommend using low confidence to “escalate to a person or a reasoning model.” That is an architecture choice, not a property of the model.

Figure 1: In AI usage, “System 2” names a way of spending more compute per answer. “System One” names the opposite bet, with the handoff between the two written in code.
How the mechanism differs from what teams already use
TypeSafe has not published Jev’s architecture, parameter count or training data. The launch post says the company built “a new model architecture, parallel sampler for maximum efficiency, and training method” called Reinforcement Learning for Calibrated Decisions (RLCD). The AI primer sets RLCD beside the two post-training methods it is meant to succeed. RLHF “trains models to produce responses people prefer”. RLVR “created reasoning models”. RLCD “trains TypeSafe to return decisions and calibrated probabilities instead of generated text.” Almeida is a co-author of the InstructGPT paper that introduced RLHF to instruction following, so the argument that preference training damages calibration comes from someone who helped build it.
What can be described is the output contract, and that is enough to compare it with the four ways teams already turn model output into a branch in code.
Autoregressive LLMs, with or without structured outputs
An autoregressive LLM generates one token at a time, each conditioned on the last. To get a decision, a developer prompts for a label and parses the reply. Structured outputs make the parse safe. OpenAI’s guide says its feature means developers need not “worry about the model omitting a required key, or hallucinating an invalid enum value.” Anthropic’s documentation says its structured outputs “guarantee schema-compliant responses through constrained decoding.” Constrained decoding masks tokens that would break the schema at each step. The output is valid JSON by construction, which is TypeSafe’s first claim, and both OpenAI and Anthropic already ship it as a documented feature.
What constrained decoding does not give is a trustworthy probability for the decision. Token log-probabilities exist on some APIs, but they are the probabilities of tokens after preference tuning, and the evidence says post-training disturbs them. The GPT-4 Technical Report states that “the pre-trained model is highly calibrated”, but “after the post-training process, the calibration is reduced.” Asking the model to state its confidence in words is the other option. Xiong and colleagues found in 2023 that LLMs “when verbalizing their confidence, tend to be overconfident.” Tian and colleagues found that verbalised confidence from RLHF models was still usually better calibrated than their token probabilities, “often reducing the expected calibration error by a relative 50%”. None of these approaches is designed to be calibrated. Calibration is recovered afterwards, imperfectly.
Non-autoregressive decoding and classic classifiers
Producing outputs in parallel rather than token by token is an old idea. Gu and colleagues’ 2017 paper on non-autoregressive translation removed the dependence of each output word on the previous ones “and produces its outputs in parallel, allowing an order of magnitude lower latency during inference.” A text classifier goes further. An encoder such as ModernBERT reads the whole input once and a classification head produces a softmax over a fixed label set. That is one pass, one distribution and no generation. ModernBERT’s authors call encoders “the workhorse of numerous production pipelines” for classification and retrieval.
The catch is that a classifier’s labels are fixed at training time. A new routing category means new labelled data and a new model; techniques such as SetFit cut the data needed, but a model is still trained per task.
Reward models and LLM-as-judge
The third approach is scoring. A reward model, as in InstructGPT, emits a scalar trained on human preference rankings. An LLM-as-judge prompts a strong model to grade. Zheng and colleagues found GPT-4 as judge reached “over 80% agreement” with human preferences. They also documented “position, verbosity, and self-enhancement biases”. A judge is flexible and often accurate, but it is an autoregressive LLM with all the cost and latency that implies, and its scores carry no calibration guarantee.
Where the System One model sits
Seen against these, the System One model looks like a promptable classifier. Its output resembles a classifier’s: a fixed answer space, one pass, a full distribution. Its input resembles an LLM’s: the question and the answer space are written in natural language at request time, so no per-task training is needed. TypeSafe says “the same weights serve every account” and that Jev “is not fine-tuned or LoRA-adapted with customer data”. Customisation happens through the state, the instructions and the criteria. The claimed extra is that the distribution is calibrated by training rather than repaired afterwards.
| Approach | Output | Latency per decision | Calibration | Cost shape | Flexibility |
|---|---|---|---|---|---|
| LLM + structured output | Schema-valid text; label chosen by generation | Seconds; more with reasoning | Not designed in; token probabilities degraded by post-training | Input + output tokens; output priced higher | Any task, any schema, changed by prompt |
| Fine-tuned classifier | Softmax over fixed labels | Milliseconds, self-hosted | Often overconfident; fixable with temperature scaling on held-out data | Labelling + training + hosting; no per-token fee | Labels fixed at training time |
| LLM-as-judge | Score or verdict as text | Seconds | Not designed in; documented biases | Full LLM tokens per judgement | Any rubric, changed by prompt |
| System One model (Jev) | Typed answer + distribution + confidence | 70 to 500 ms end to end (vendor figure) | Trained for it (vendor claim; not independently verified) | Input tokens only, US$0.042/M | Any question within Choice, Score or Noul; no text |
Table 1: Four ways to get a decision out of a model. Jev’s latency, calibration and price are TypeSafe’s published figures as of September 2026.
The table shows what the category gives up as clearly as what it gains. Generation is gone, so an explanation, a rationale or an extracted free-text value is impossible. TypeSafe’s jaggedness page lists further limits for jev-1.13. It “does not count reliably”, reads dates “as text, not as ordered quantities”, “can be quite literal”, and loses accuracy as irrelevant state grows. Choice questions top out at 255 options. For more, the launch post describes a two-stage scoring workaround.

Figure 2: Each approach wins somewhere. The System One pitch is aimed at the gap between the classifier’s fixed labels and the LLM’s per-token bill.
Calibration, properly defined, and why software cares
Calibration is the whole bet, so it needs a precise definition. A model is calibrated when, across all the predictions to which it assigns probability p, the fraction that turn out correct is p. Guo, Pleiss, Sun and Weinberger’s 2017 paper On Calibration of Modern Neural Networks is the standard reference. It measures calibration with reliability diagrams, which plot accuracy against confidence in bins, and with Expected Calibration Error (ECE), the weighted average gap between confidence and accuracy across those bins.
The paper’s central finding still frames the debate. “Modern neural networks, unlike those from a decade ago, are poorly calibrated.” A 110-layer ResNet was more accurate on CIFAR-100 than a 1998-era LeNet but systematically overconfident. The fix was temperature scaling, a single learned parameter that softens the output distribution, which proved “surprisingly effective” on most datasets. The lesson is that calibration is not magical. It is measurable, it degrades in predictable ways, and it can often be repaired with held-out data.
TypeSafe’s documentation is careful on the same points. Calibration “is measured across groups of predictions; it does not guarantee that an individual answer is correct.” Its AI primer gives the textbook version: outcomes assigned 0.2 “should occur about 20% of the time”, and those assigned 0.8 about 80 per cent of the time.
Confidence is not the same number as probability
One detail in the docs is easy to miss. Jev’s confidence field is not the probability of being right. It is “a statistic computed from the probability distribution”, a measure of how concentrated the distribution is. The docs’ interactive example uses (3 × largest probability − 1) / 2 to approximate it for three options. A Choice with probabilities of 0.7, 0.2 and 0.1 therefore reports a confidence of about 0.55, not 0.7. The calibration claim attaches to the probabilities. Confidence is a convenience for thresholding. Teams that set a 0.9 confidence gate and assume a 90 per cent success rate will be making a category error. The docs themselves recommend starting conservatively and testing on your own data.
Why branching code needs calibrated numbers
For a human reading a chatbot answer, a miscalibrated “I’m fairly sure” is an annoyance. For code that acts without a human, it is a cost line. The expected cost of automating a decision is the number of decisions times the probability of error times the cost of an error. Only a calibrated probability lets that sum be computed in advance.
Take a refund workflow that auto-approves every case the model scores above 0.95, at 10,000 such cases a day. If the model is calibrated, roughly 500 of those approvals a day are wrong, and the business can price that. If it is overconfident and the true accuracy in that bin is 85 per cent, the figure is 1,500 a day, three times the planned loss, and nothing in the output reveals the gap. The launch post makes the same argument: “If a model can do a task 95% of the time but doesn’t say when it’s in the 5%, it can’t automate that task.”
Calibration has limits the marketing does not dwell on. A model right 60 per cent of the time that answers 0.6 to everything is perfectly calibrated and tells the code nothing; sharpness matters too. Calibration also drifts across distributions. A model calibrated on English support tickets offers no guarantee on German insurance claims, and TypeSafe’s models page says other languages are handled “not equally well”.
The economics: why a decision-shaped API changes the unit cost
The commercial argument is less about intelligence than about cost shape. An LLM bills input and output tokens, with output typically priced several times higher. Reasoning adds hidden output tokens on top. Jev bills input only. As of September 2026, TypeSafe’s models page lists jev-1.13.0 at US$0.042 per million input tokens (US$42 per billion) with output tokens free. Rate limits are 250,000 tokens per second and 1,200 requests per minute, and the context window is 64k tokens per request. The launch post claims end-to-end latency of 70 to 500 ms. It also notes that its published evals were generally run from laptops on the West Coast, where the service is currently based.
A worked comparison on list prices
Consider a ticket-triage service handling 10 million tickets a month. Each request is about 1,500 input tokens of state and questions.
- Jev: 15 billion input tokens at US$42 per billion is US$630 a month. There is no output charge.
- Claude Haiku 4.5 with structured outputs: 15 billion input tokens at US$1 per million is US$15,000. Add about 80 output tokens per ticket, 800 million tokens at US$5 per million, which is US$4,000. That makes roughly US$19,000 a month before any prompt caching, which Anthropic prices at a tenth of the input rate for cache hits.
That is about a thirtyfold difference on list prices. The arithmetic says nothing about accuracy. A fine-tuned classifier on a rented GPU could undercut both on cost, but it carries labelling and retraining costs that do not appear per token. The per-call price is one line of several in the real cost of an AI feature.
What TypeSafe’s own evals show
TypeSafe’s workflow evals give the most detailed public comparison, and the conditions matter. The four workflows are security-incident triage, agent-trace review, invoice processing and customer service. Each is broken into Noul, Choice and Score questions, with code combining the answers. “Accuracy” means agreement with reference labels produced by averaging GPT-6 Astra and Claude Fable 5.1 at high thinking. It is not agreement with human ground truth. The workflows were written by TypeSafe’s own capabilities team. The launch post says “some bias could exist” and that the headline 193.6x-faster, 444.6x-cheaper figures are “on the higher end of real world gains”.
Averaged across the four workflows, as published in September 2026:
| Model (as labelled by TypeSafe) | Agreement with reference | Cost per case | Time per case |
|---|---|---|---|
| Jev | 67.8% | US$0.0004 | 0.4 s |
| Terra | 67.9% | US$0.0304 | 10.1 s |
| Sonnet 5 | 67.8% | US$0.1174 | 78.1 s |
| Luna | 66.8% | US$0.0033 | 12.9 s |
| Sol | 74.1% | US$0.0836 | 23.3 s |
| Opus 5 | 73.1% | US$0.1761 | 37.8 s |
| Haiku 4.5 | 53.6% | US$0.0195 | 12.5 s |
Table 2: TypeSafe’s workflow-eval averages. All LLMs run at the provider’s default reasoning setting, in the workflow form. Source: evals.typesafe.ai, September 2026.
Read carefully, the table supports a narrower claim than the headlines. Jev matches the mid-tier models at a fraction of their cost and latency: about 75 times cheaper than Terra, and about 8 times cheaper and 30 times faster than Luna, the cheapest model near its score. It does not match the strongest models. Sol and Opus 5 lead by five to six points. The evals also report no calibration metric, so the property the category is named for is not what this benchmark measures. The usual caution applies to a vendor grading itself against references it chose.
Who the category threatens, and who it helps
LLM providers lose the least glamorous traffic first: classification, routing and moderation calls now served by their cheapest tiers with JSON output. A provider could respond by exposing calibrated probabilities for enumerated outputs. As of September 2026 none of the major providers documents such an endpoint.
Classifier vendors and in-house ML teams face the sharper threat. A zero-shot model that returns a calibrated distribution over labels written in English erodes the case for training a bespoke model. That holds most where labels change often or data is scarce. It holds least where volume is huge and labels are stable, because a self-hosted encoder then has no per-token fee at all.
Guardrail and evaluation tooling is both threatened and helped. TypeSafe publishes a guardrails cookbook that screens “every message going into and out of an LLM app with one TypeSafe request”, and its launch post lists “Score, judge, verify, guardrail, and detect jailbreaks” as core uses. That puts a System One model directly into the LLM-as-judge role at a far lower per-call cost. A cheap scorer inside an existing guardrail or eval product also improves that product’s margins.
Infrastructure stands to gain. A new model type with its own endpoint is a routing, key-management and logging problem. Gateways have moved to carry it. Bifrost, the open-source AI gateway, exposes TypeSafe’s native API alongside a /v1/decisions endpoint, covered in the Bifrost integration story. Reasoning models gain too, because a front-line System One filter sends them only the hard, low-confidence cases, where their extra cost pays.
Other entrants, and investor interest
As of September 2026, no other commercial vendor has documented a model sold as a System One or equivalent decision model. The imitators so far are open-weight. Laya, from Convai Innovations under Apache 2.0, calls itself a “non-autoregressive System 1 decision engine”. It is built on ModernBERT-large with 421 million parameters, plus a 322-million-parameter multilingual variant, and exposes typed choice, score and yes/no decisions. Von, also Apache 2.0 and built on ModernBERT-large, describes itself as “an open-source, non-autoregressive System One decision model” answering Choice, Noul and Score questions. Both appeared within two weeks of Jev’s launch, and both are the promptable classifier described above, built on an existing encoder.
Laya’s README is also the only published outside ECE figure for Jev. It reports 0.246 for Jev against 0.081 for Laya after temperature refitting. It also concedes that Laya’s own base checkpoints are “over-confident as shipped”, with an ECE of 0.466 before refitting. That is Guo et al.’s 2017 finding repeated in 2026. These are a competitor’s self-run numbers on its own benchmark and should be treated as such. They do show that calibration claims are now being contested in public.
Investors have been less cautious. TypeSafe’s launch release names DCVC as lead of a US$40 million seed. DCVC general partner James Hardiman is quoted describing the problem as “turning increasingly capable models into technology that developers can reliably build into products at scale.” Bloomberg reported, citing the Financial Times, that the seed valued the company at about US$200 million and that investors have since approached TypeSafe with offers valuing it at more than US$10 billion. Neither valuation has been confirmed by the company, and the reported round had not closed at the time of writing.
The case against the category
Structured outputs are good enough. For many teams the pain Jev removes, malformed JSON and invented enum values, was already removed by constrained decoding. If the task needs only a label, a small LLM with a schema is a known quantity from a vendor the team already uses. The counter is cost and latency at volume, which matters only above a certain traffic level.
Fine-tuned small classifiers are cheap. An encoder such as ModernBERT, trained with a few hundred labelled examples, runs in milliseconds on commodity hardware with no per-call fee. Calibration can be tuned with temperature scaling on held-out data, a standard technique. Teams with stable labels and in-house ML capacity can own the whole stack. The existence of Laya and Von shows how quickly that stack can be assembled.
Lock-in to one vendor’s proprietary model. Jev’s weights are closed, one set serves every customer, and there is no fine-tuning. A team that tunes thresholds against jev-1.13.0 depends on TypeSafe’s pricing, capacity and roadmap. The models page warns that rate limits “can change without notice” while demand is high. The launch post concedes that TypeSafe “can’t prove” its pricing is not subsidised. Two things soften the risk. TypeSafe publishes an MIT-licensed adapter that serves the same system_one API from OpenAI, Anthropic or Gemini models. And open-weight models now copy the primitives. The interface may be more portable than the model.
Calibration is unproven. This is the strongest objection. As of September 2026, the defining property of the category has been asserted by its vendor and measured by nobody independent. The vendor’s benchmark does not report it, and the only outside figure comes from a competitor. Calibration is also distribution-specific. A number measured on anyone’s benchmark says little about a given team’s tickets, claims or logs. Until independent reliability diagrams exist, “calibrated” is a hypothesis each customer must test.
What would have to be true for System One to stick
A category outlives its first vendor when buyers can compare suppliers and switch between them. Four conditions look necessary.
- Calibration holds up under independent measurement. That means published reliability diagrams and ECE on public datasets, run by third parties, across domains and languages. Without it the category collapses into “a fast classifier”, which the market already has.
- The unit economics survive. Free output tokens and US$42 per billion input tokens are a launch price from a company that says it cannot yet prove the price is sustainable. The category needs the price to hold, or fall, as volume grows.
- The interface becomes common. Choice, Score and Noul are a small, sensible vocabulary. If open-weight models, adapters and gateways adopt it, as they have begun to, buyers get the portability that makes a category rather than a single product.
- The big labs decline to compete, or compete and validate the idea. A calibrated-classification endpoint from a major provider would crowd TypeSafe’s market. It would also confirm that System One is a category.
What to do now
For teams deciding whether any of this matters to them:
- Inventory your decision calls. Count the LLM requests whose output is a label, a score or a yes/no. If that traffic is large, the cost gap in Table 2 is worth testing. If it is small, the integration cost may outweigh the savings.
- Measure calibration on your own data before trusting a threshold. Label a few hundred cases, bin the model’s probabilities, and draw the reliability diagram. Do this for Jev, for an open-weight alternative and for your current LLM.
- Gate on probabilities you have validated, not on
confidencealone. Confidence describes the shape of a distribution. It is not the chance of being right. - Keep the fallback path. Escalate low-confidence cases to a reasoning model or a person. That limits the damage if calibration is worse than claimed, and the lock-in if terms change.
- Watch for three signals. An independent calibration study, a second commercial supplier, and a frontier lab shipping enumerated-probability outputs. Any one of them would settle more about the category than another funding headline.
The idea behind System One models is sound and overdue. Much of what software asks of AI is a decision, not an essay, and a decision needs a probability you can trust more than it needs fluent prose. What is unsettled is whether trustworthy probabilities are a property of one company’s training method or an interface anyone can implement. As of September 2026 the evidence points both ways, and the calibration numbers that would decide it have not been published.
Sources
- Jev, an AI Model That Can't Chat, Takes On Bigger Rivals Bloomberg
- Introducing System One Models & Jev TypeSafe AI
- System One (concept documentation) TypeSafe AI
- AI primer (RLHF, RLVR and RLCD) TypeSafe AI
- Confidence (documentation) TypeSafe AI
- Models, pricing and rate limits TypeSafe AI
- Jev 1.13 jaggedness TypeSafe AI
- Workflow evals TypeSafe AI
- TypeSafe AI Emerges From Stealth With $40M in Funding With New Model for Composable AI Business Wire
- Kahneman, "Maps of Bounded Rationality: A Perspective on Intuitive Judgment and Choice" (Nobel Prize lecture) Nobel Foundation
- Bengio, "From System 1 Deep Learning to System 2 Deep Learning" (NeurIPS 2019 invited talk) NeurIPS
- Li et al., "From System 1 to System 2: A Survey of Reasoning Large Language Models" arXiv
- Guo et al., "On Calibration of Modern Neural Networks" arXiv (ICML 2017)
- OpenAI, "GPT-4 Technical Report" arXiv
- Kadavath et al., "Language Models (Mostly) Know What They Know" arXiv
- Xiong et al., "Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs" arXiv
- Tian et al., "Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback" arXiv
- Structured model outputs guide OpenAI
- Structured outputs documentation Anthropic
- Claude API pricing Anthropic
- Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" arXiv
- Ouyang et al., "Training language models to follow instructions with human feedback" arXiv
- Gu et al., "Non-Autoregressive Neural Machine Translation" arXiv
- Tunstall et al., "Efficient Few-Shot Learning Without Prompts" (SetFit) arXiv
- Warner et al., "Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder" (ModernBERT) arXiv
- Laya repository Convai Innovations (GitHub)
- Von repository GitHub
- System One adapter for Python TypeSafe AI (GitHub)
Frequently asked questions
What is a System One model?
It is TypeSafe's term for "a new class of frontier models built to make fast, structured decisions that software can use directly." A System One model reads unstructured state, such as a support ticket or a JSON record, answers a set of typed questions about it in one pass, and returns choices, scores or yes/no probabilities with a probability distribution. It does not generate text.
Why is it called System One?
TypeSafe says the name comes from Daniel Kahneman's Thinking, Fast and Slow, which popularised the distinction between fast, automatic System 1 thinking and slow, deliberate System 2 reasoning. The AI field had already used System 2 as shorthand for reasoning models. TypeSafe's naming stakes out the fast half as a product category, though it claims its models are more reliable than the psychological System 1 would suggest.
Is Jev the only System One model?
Jev is the first and, as of September 2026, the only commercial model sold under the name. Two open-weight projects, Laya from Convai Innovations and Von, appeared within two weeks of launch. Both describe themselves as System 1 or System One decision models built on the ModernBERT-large encoder, and both copy the same three question types. No other commercial vendor has documented a comparable product.
How is a System One model different from an LLM with structured outputs?
Structured outputs use constrained decoding to guarantee that an autoregressive LLM's text matches a JSON schema, so the format is safe but the model still generates token by token and gives no calibrated probability for its choice. A System One model returns the full probability distribution over the allowed answers directly, in one pass, and is billed only on input tokens.
What does calibrated mean for an AI model?
A model is calibrated when its stated probabilities match observed frequencies. Across all the answers it gives 0.8, about 80 per cent should be correct. Guo et al. showed in 2017 that modern deep networks tend to be overconfident and that a one-parameter fix, temperature scaling, often corrects this. Calibration is a property of groups of predictions, not a guarantee for any single answer.
How much does Jev cost compared with an LLM?
As of September 2026, Jev costs US$0.042 per million input tokens with output free. Claude Haiku 4.5 lists at US$1 per million input tokens and US$5 per million output tokens. On TypeSafe's published workflow evals Jev averaged US$0.0004 per case against US$0.0304 for the Terra model at the same agreement score, though those evals are TypeSafe's own.
Are TypeSafe's calibration claims independently verified?
Not as of September 2026. TypeSafe's workflow evals report agreement with reference models, cost and latency, not calibration error. The one outside expected calibration error figure for Jev appears in the repository of Laya, an open-weight competitor, which reports its own measurement. Teams should measure calibration on their own labelled data before relying on the thresholds.


