News

TypeSafe AI launches Jev, a $40M bet on AI that decides instead of chats

Jev takes text state and typed questions and returns calibrated probabilities in one pass, at $0.042 per million input tokens. What launched, who built it, and what the evidence shows so far.

Illustration of an API call card labelled POST /v1/systemone, where a text state document flows through an arrow into three typed answers (Choice, Score, Noul) with probability bars, beside a crossed-out chat bubble reading no chat and a gradient panel reading $40M seed round, 15 Sep 2026.
Illustration of an API call card labelled POST /v1/systemone, where a text state document flows through an arrow into three typed answers (Choice, Score, Noul) with probability bars, beside a crossed-out chat bubble reading no chat and a gradient panel reading $40M seed round, 15 Sep 2026.

TypeSafe AI came out of stealth on 15 September 2026 with a $40 million seed round led by DCVC and a model that cannot hold a conversation. Jev, which the company calls its first “System One” model, reads a block of text and a set of questions whose possible answers the developer has already defined, and returns a probability for every answer in one pass. TypeSafe prices it at $0.042 per million input tokens, with output free, and quotes end-to-end responses of 70 to 500 milliseconds.

The launch mattered for anyone who puts a model inside a loop: support-ticket routing, content moderation, agent tool selection, guardrails, any place where code needs a decision rather than a paragraph. Within two weeks Jev was on Vercel’s and Cloudflare’s AI platforms, Vercel called it the fastest-adopted model in its gateway’s history, and Bloomberg, citing the Financial Times, reported investor offers valuing the company above $10 billion.

This piece sets out what shipped, who built it, how the API works as documented, and how TypeSafe’s claims compare with the independent measurements published so far. Where a number comes from TypeSafe it is attributed; where it comes from someone else, that is said too.

What TypeSafe released on 15 September

Three things went out on the same day. A launch post by founder Diogo Almeida introduced the “System One Model” category and Jev. A Business Wire release announced the funding. And TypeSafe opened early access, moving developers off a waitlist through its console.

The launch post frames the product with one question: “Models have been superhuman at chat for years, so where is all the automation?” Its answer is that chat models are trained for the wrong job. TypeSafe describes a new stack “entirely focused on automation”: a new model architecture, a parallel sampler and a training method it calls Reinforcement Learning for Calibrated Decisions, or RLCD. The post’s shortest description of Jev is “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.”

The company made four headline claims at launch:

  • Jev reaches “similar levels of intelligence on System One tasks compared to existing LLMs, while being two orders of magnitude faster and more efficient.”
  • It “can’t hallucinate”, in the narrow sense that it can only return answers from a set the developer defined.
  • The model “never makes type errors”, because the output schema is fixed before the call.
  • Every answer carries calibrated probabilities, meaning, in TypeSafe’s definition, that higher confidence should correspond to higher accuracy across many predictions.

Alongside the post, TypeSafe published a side-by-side demo against GPT-5.6 Terra, a set of four “workflow evals”, a Doom-playing bot and a Wikiracing demo. The company’s homepage claim of “193.6x faster, 444.6x cheaper” comes from those workflow evals.

The name comes from William Stanley Jevons. The launch post says TypeSafe expects machine intelligence “to follow a similar path to coal”, where efficiency gains raised total demand rather than lowering it. The model class name borrows Daniel Kahneman’s split between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning.

Who is behind TypeSafe

TypeSafe was founded in 2024 and is based in San Francisco, according to the press release. Its team page lists three founders.

Diogo Almeida, CEO. TypeSafe and DCVC describe him as a co-inventor of RLHF and InstructGPT. The verifiable part is that he is one of 20 authors on OpenAI’s 2022 paper Training language models to follow instructions with human feedback, the InstructGPT paper that set out the recipe behind ChatGPT. Reinforcement learning from human preferences as a technique predates that paper, so “co-inventor” is the company’s framing of his role in applying it to language models. The team page also lists prior work at Google Brain. In the launch post he writes that at OpenAI he “helped build the methods that made language models useful at following instructions and talking with people”, and that TypeSafe spent “two years in stealth”.

Erik Gafni, CTO. The team page describes him as a repeat founder (Ravel, a multimodal AI company for DNA sequencing) and an early employee at Invitae and Freenome.

Sasha Sheng, COO. Described as a former research engineer at Meta and FAIR who worked on News Feed, AI experiences and AI research, with publications at NeurIPS and ECCV.

Those biographies are self-reported. None of them is unusual for a well-funded AI start-up, and the InstructGPT authorship, the one most relevant to the product’s thesis, checks out on arXiv. The thesis itself follows from it: the person who helped tune models to please human raters now argues that pleasing human raters is the wrong objective for software.

The money and the reported valuation interest

The funding facts are consistent across primary sources. The press release, DCVC’s announcement by general partner James Hardiman, and a deal note from Wilson Sonsini, which advised the company, all give $40 million in seed funding led by DCVC and announced on 15 September 2026. DCVC calls it a “Series Seed”. No other investors are named in those three documents.

Valuation is where the sourcing thins out. None of TypeSafe’s own announcements states a valuation. Bloomberg reported on 25 September that, according to the Financial Times, the seed round valued the company at $200 million, and that TypeSafe has since been approached by investors with offers valuing it at more than $10 billion. Bloomberg also reported that the launch video had been viewed about 40 million times on X, and compared the reaction to the market’s response to DeepSeek in 2025.

Two cautions apply. An approach is not a term sheet, and neither Bloomberg nor the FT, as Bloomberg reports it, says a round at that price has closed. And a fifty-fold jump in two weeks would be priced on adoption signals and demos rather than revenue, since no revenue or customer figures have been disclosed. For readers deciding whether to build on Jev, the relevance of the valuation is indirect: a large round would pay for the GPU capacity that TypeSafe’s own documentation says it is waiting on.

How Jev’s API works

The TypeSafe documentation describes a single endpoint, POST https://api.typesafe.ai/v1/systemone, that takes three fields:

  • state: the content to evaluate. A string, a JSON object or an array of text values.
  • model: jev-latest or a pinned version such as jev-1.13.0.
  • questions: a map of named, typed questions.

There are three question types, which TypeSafe calls primitives.

PrimitiveWhat it asksWhat comes back
ChoicePick one option from a set you define (up to 255)choice, a probability per option, confidence
ScoreRate the state on an ordered rubric (2 to 10 levels)a probability-weighted score, a probability per level, confidence
NoulIs this statement true?noul, a probability from 0 to 1

The quick-start example sends a support message (“I’ve been trying to connect my Stripe account for 3 days and the integration keeps failing. I’m losing sales.”) with three questions: which team should handle it, how frustrated the customer is, and whether it is urgent. The documented response routes it to technical with probability 0.85 and confidence 0.78, puts frustration at level 1 (“Frustrated but civil”) with confidence 1.0, and returns a noul of 1.0 for urgency. The usage block reports 392 input tokens.

A few design choices stand out in the API reference.

Questions are evaluated in isolation. The docs say every question “is evaluated in parallel and in isolation against the same state in one go”, so adding questions “barely changes the response time” and does not cause context rot. The question’s key name is not sent to the model.

Confidence is separate from probability. Choice and Score answers carry a confidence value derived from the shape of the probability distribution. A Choice that splits 0.5 and 0.5 has low confidence even though one option wins. TypeSafe’s patterns section builds on this, with “confidence-gated routing” that acts automatically above a threshold and escalates to a person or a reasoning model below it.

Atomic questions, composed in code. The docs tell developers to decompose judgements: rather than “rate this startup pitch”, ask separately about market size, technical feasibility and differentiation, then weight the results in code. When priorities shift, the developer changes a coefficient rather than a prompt.

Errors follow ordinary HTTP conventions: 401 for a bad key, 422 for a malformed request, 429 for rate limits and 529 when TypeSafe is overloaded. Official SDKs exist for Python (typesafe-sdk) and JavaScript (@typesafe-ai/sdk), and TypeSafe ships an agent skill for Claude Code and other coding agents.

Two lanes. Top lane, an LLM call: prompt plus schema, then tokens generated one by one, then parse and validate, then retry or give up with no calibrated probability. Bottom lane, a Jev call: state plus Choice, Score and Noul questions, then one parallel pass, then typed answers restricted to listed options, then probabilities plus confidence

Figure 1: The difference is not only speed. An LLM produces a string that code must check; Jev can only return options the developer listed, each with a probability.

Price, limits and availability as of September 2026

The Models page lists one current model, Jev 1.13 (jev-1.13.0), with these terms:

ItemDocumented value
Price$0.042 per million input tokens ($42 per billion); output tokens free
Rate limits250,000 tokens per second and 1,200 requests per minute
Context64k tokens per request; 32k for the state plus the longest single question
InputText only: string, JSON object or array of text. No image, audio or video
Choice optionsUp to 255 per question
Score levelsAt least 2, up to 10
LanguagesEnglish primary; others, including CJK scripts, “handled but not equally well”
CustomisationNone. No fine-tuning or LoRA; same weights for every account
DataNot trained on customer requests; zero data retention for enterprise customers

Two warnings on that page matter more than the table. Rate limits “are adjusting dynamically” and “can change without notice” while TypeSafe serves “a very large volume of demand” ahead of “upcoming large GPU deals”. And the jev-latest alias moves when a new release ships, so “the answers behind it can change without a change on your side”. TypeSafe advises pinning a versioned ID once confidence thresholds have been tuned against it.

To put the price in concrete terms: the quick-start ticket used 392 input tokens. At $0.042 per million, that is about $0.000016 per call, or roughly $16.50 for a million tickets, each answered with three separate judgements. The same 392 tokens in, plus around 60 tokens of JSON out, at the $0.10 input and $0.50 output list price of GPT-6 Luna reported in GenAI Brief’s September price survey, would cost about $0.00007, roughly four times as much, before any reasoning tokens or retries. That is a much smaller gap than the headline multiples, which compare Jev with frontier models on longer workflows.

TypeSafe itself flags the obvious question about its price. The launch post says: “We can’t prove it isn’t subsidized; we’ll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up).”

Jev is also reachable outside TypeSafe’s own console. Vercel added it to its AI Gateway on 16 September, as typesafe-ai/jev through an experimental evaluate API in AI SDK 7. Cloudflare lists it in Workers AI as typesafe/jev, version jev-1.13.0, with a 32,000-token context at the same $0.042 input price. Bifrost, the open-source gateway, exposes TypeSafe’s native API under a /typesafe prefix so the official SDKs work by changing a base URL, with Bifrost virtual keys, budgets and logging applied per request. That integration is covered in detail in a separate GenAI Brief report.

How the launch was received

Adoption

The fastest signal came from Vercel. In a blog post on 18 September, the company said that within 24 hours of launching on its AI Gateway, Jev “reached more than twice as many paid teams as any previous model launch”, and that by hour 24 nearly 13 per cent of paid teams were using it: twice the share of the GPT-5.6 family and more than six times that of Fable 5.1. Vercel added its own caveat: “the next test is whether that early adoption lasts.” Free promotional access on some gateways in the first week will have inflated trial numbers.

Developers

The Register described Jev on 23 September as “a classifier with brains” and catalogued what developers built in its first week. FPV Ventures partner Nikunj Kothari set up a site called Jevable to collect prototypes. They included an “Urgency” column for spreadsheets, a virtual clothes try-on that The Register timed at about 620 ms and $0.0011 per decision, and game bots for Doom, Tetris and chess, where Jev lost to the open-weight GLM 5.3 model but cost far less to run. Someone coined “JevOps” for running code on Jev-based emulation, which The Register noted was “still only a meme, not a discipline”.

The same article collected the early reaction from people with large followings. Adam Jacob, chief executive of Swamp Club, wrote on X: “It’s time to move past the idea that what we need is smarter frontier models. What we need is smarter systems.” Andrej Karpathy, identified by The Register as an Anthropic researcher, wrote that Jev “revealed latent demand […] that was under-invested into because of a race to higher intelligence”, describing that demand as a single-token model with low latency and “acceptable intelligence”.

The sceptical line was put most plainly by YouTuber Mo Bitar: “I know it’s fast. I know it’s cheap, but is it good?”

Clones

Within days, developers were reproducing the interface. The Register mentions open-source classifiers in the same vein, named Jeff and Nimble. None had been independently evaluated at the time of writing, and matching Jev’s request and response shape is not the same as matching its quality.

Claims versus independent evidence

“Is it good?” is the right question, and TypeSafe deliberately declined to answer it with public benchmarks. Its launch FAQ says the company “deliberately chose not to publish performance against public benchmarks” and plans “only one-off evals when we make product updates”, urging users to build their own evals instead. That position has merit, given how easily leaderboard numbers are gamed (see GenAI Brief’s editorial on benchmarks), but it leaves outsiders to do the measuring.

What TypeSafe’s own evals show, and how it qualifies them

The launch post is unusually candid about its own evidence. On the workflow evals behind the 193.6x and 444.6x figures, it says:

  • the reference answer is the average of GPT-6 Astra and Fable 5.1, not a human-labelled ground truth, which “biases answers towards OpenAI and Anthropic’s models”;
  • the workflows “were made by individuals on our model capabilities team, so some bias could exist”, although TypeSafe says they are outside its training distribution;
  • the multiples are “on the higher end of real world gains”;
  • the LLMs were run through TypeSafe’s own wrapper that constrains them to structured decisions, which it says is the most accurate way to get decisions from LLMs but slower and costlier.

On hallucination, the post concedes that Jev’s 0 per cent type-error figure “is not empirical”: it follows from the output being restricted to a fixed schema, while the comparison numbers for LLMs come from OpenRouter traffic, which carries its own bias. On speed, it notes that published evals “are generally run from our laptops on the West Coast”, near where the service is hosted.

Even TypeSafe’s latency figures vary by document. The launch post says 70 to 500 ms. The press release and DCVC’s post say “less than 100 milliseconds”. The Register, citing TypeSafe, gives “as little as 150 ms”. These are not contradictions so much as different best cases, and they show why teams should time calls from their own region.

What outsiders have measured

Three pieces of independent work stand out in the first fortnight.

A black-box probe. Engineer Archer Hume published Jev’s Architecture Unmasked on 17 September after roughly 10,000 API calls. With one question, median latency rose from about 57 ms at a 360-token state to about 218 ms at 29,835 tokens. With a short state, it rose from about 87 ms for one question to about 610 ms for 1,500 questions. On 1,200 MMLU-Pro items he measured an expected calibration error (ECE) of 0.0313, with most predictions in the 0.9 to 1.0 confidence band. He inferred a causal transformer backbone that reads probabilities directly from internal representations, a shared state encoding with isolated question branches, and a tokenizer that matched none of 192 public tokenizers exactly but agreed most with Qwen’s (348 of 415 probes). He labels much of the architecture analysis “quite speculative”. It is still the strongest evidence against the FAQ line that Jev is “neither small nor an LLM”: the backbone looks like a language model, with the text-generation head replaced.

A calibration audit. An independent GitHub audit of about 7,000 API calls, costing under a dollar, tested what Jev’s probabilities mean. It found no option-order bias (zero argmax flips in 400 trials), negligible interference when 16 questions were bundled instead of one, and a p50 of 280 ms and p95 of 397 ms for a single question called from Seoul. It also found weaknesses. Complementary questions produced probabilities summing to anywhere from 0.71 to 1.42. Removing an “unknown” option on an ambiguous-question dataset sent ECE from 0.023 to 0.793. And 50 identical requests gave 15 distinct answers. The audit also notes that earlier independent tests disagree: near-calibrated on MMLU, over-confident on post-cutoff phishing data.

A preprint. On 24 September Delip Rao and Chris Callison-Burch posted JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places to arXiv. They compared Jev with three flash-tier LLM judges on nine panels drawn from seven benchmarks. Jev’s accuracy differed significantly from an LLM judge’s in only 8 of 27 paired comparisons, ahead mostly on binary criteria and behind on graded ones. Summed across panels, the LLM judges cost 29 to 325 times as much and took 30 to 220 times as long. The catch is in the title. The LLM judges repeat nearly all of Jev’s most confident errors, so a cascade that sends Jev’s uncertain cases to an LLM gained at most 1.5 points over the best single judge.

TypeSafe’s own documentation is franker than its launch copy about failure modes. The Jev 1.13 jaggedness page, last reviewed on 17 September, lists nine: literal reading, arithmetic and counting, date comparison, multi-hop indirection, large states full of irrelevant detail, adversarial content in the state, contradictory instructions, structural invariants that do not hold (a Noul and a yes/no Choice on the same question can disagree), and generation. It says plainly that state “is data, and jev-1.13 does not treat it as hostile by default”, so injected instructions can move an answer.

Three columns. TypeSafe says: 70 to 500 ms per call, up to 193.6x faster and 444.6x cheaper, zero type errors, calibrated probabilities. Others measured: about 57 to 610 ms in Hume's probes, ECE 0.031 on MMLU-Pro, LLM judges 29 to 325 times costlier, errors shared with LLM judges. Still open: accuracy on real workloads, whether pricing is subsidised, architecture and model size, rate limits at scale

Figure 2: After two weeks, outside measurements broadly confirm the speed and cost claims. Whether Jev is as good as the LLMs it replaces on a given workload is still for each team to test.

Where Jev fits against structured outputs and classifiers

Jev enters a space that already has two incumbents.

LLM structured outputs

Every major API now offers constrained decoding. OpenAI’s structured outputs guide says the feature ensures responses adhere to a supplied JSON schema, and documents that a refusal can come back in place of a schema-valid answer. So schema adherence alone is no longer unique to Jev. What differs is the rest of the call:

  • Probabilities. An LLM returns one sampled answer. Getting a distribution means asking the model to state a confidence, which TypeSafe’s launch post argues is “overconfident and inconsistent”, or reading token log-probabilities, which reflect the text rather than the decision. The GPT-4 technical report showed calibration on multiple-choice questions was good in the pre-trained model and got worse after post-training. RLCD is TypeSafe’s attempt to train for calibration directly.
  • Cost shape. LLM output tokens are priced higher than input and reasoning models add hidden tokens. Jev charges only for input and returns all answers at once.
  • Latency shape. An LLM’s latency grows with the length of what it writes. Jev’s grows mainly with state size and, slowly, with the number of questions.
  • Flexibility. An LLM can explain itself, handle an unforeseen answer, do arithmetic and reason over several steps. Jev cannot, by design.

Trained classifiers

A fine-tuned encoder classifier trained on a team’s own labels is still likely to be cheaper per call and more accurate on a fixed, high-volume task. Jev’s advantage is that it needs no training data and no retraining when labels change. Adding a routing category is a one-line change to a Choice’s criteria, not a new labelling project. The trade-off is lock-in to a hosted model whose weights change behind an alias, and TypeSafe does not allow fine-tuning.

A practical reading: Jev sits between a rules engine and an LLM. It suits decisions with a bounded answer and fuzzy input, where the labels move too often to train a classifier and the volume or latency budget makes an LLM call painful. Worked use cases, with an escalation pattern, are covered in GenAI Brief’s guide to what Jev is good for.

What to watch next

Evidence on real workloads. The independent work so far uses academic benchmarks and synthetic probes. The first published production comparisons, with accuracy against human labels rather than against other models, will say more than any launch demo.

Pricing durability. TypeSafe says it cannot prove its price is unsubsidised. Watch whether the $0.042 rate and free output survive the move from early access to general availability, and whether gateway resellers keep the same price.

Capacity. Rate limits are officially provisional pending GPU deals. The reported valuation interest, if it becomes a priced round, would mostly go to capacity.

Model versioning. jev-preview pointed at the same model as jev-latest at the time of writing. The first alias move will show how much confidence thresholds drift between versions, which matters for anyone gating actions on them.

Modalities and safety. Image input is marked “not supported (yet)”. Prompt-injection resistance is listed as a known weakness that TypeSafe expects to improve.

Clones. If an open-weight model can match Jev’s interface and most of its quality on commodity hardware, the moat shifts to TypeSafe’s data and training method, which the company describes as its core: “TypeSafe is primarily a data research lab.”

The takeaway for builders

Jev is a real product with a documented API, public pricing and a candid list of failure modes, and the independent measurements published in its first two weeks support its speed and cost claims. They do not yet support treating it as a drop-in replacement for LLM judgement. The best evidence so far suggests it is roughly as right as flash-tier LLMs on many decision tasks and wrong in the same places.

For a team with a decision step in production, the sensible move is cheap and quick: take a few hundred labelled cases from the existing pipeline, run them through Jev with the version pinned, compare accuracy and calibration against the current approach, and time the calls from the region where the service will run. At $0.042 per million input tokens, the test will cost less than the meeting to discuss it. Treat the probabilities as a signal to calibrate against local data, not as a guarantee. Keep arithmetic, dates and anything adversarial in code. And set thresholds that assume the model behind jev-latest will change.

Sources

  1. Introducing System One Models & Jev TypeSafe AI
  2. TypeSafe AI Emerges From Stealth With $40M in Funding With New Model for Composable AI Business Wire
  3. TypeSafe emerges from stealth with a new way of doing AI DCVC
  4. Wilson Sonsini advises TypeSafe AI on $40 million seed round Wilson Sonsini
  5. Team TypeSafe AI
  6. Introduction TypeSafe AI docs
  7. Models TypeSafe AI docs
  8. API reference TypeSafe AI docs
  9. Jev 1.13 jaggedness TypeSafe AI docs
  10. AI primer TypeSafe AI docs
  11. Training language models to follow instructions with human feedback arXiv
  12. JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places arXiv
  13. Jev's Architecture Unmasked Archer Hume
  14. jev-calibration-audit GitHub
  15. Jev is the fastest-adopted model in AI Gateway history Vercel
  16. TypeSafe AI's Jev now available on AI Gateway Vercel
  17. Jev (typesafe) Cloudflare
  18. TypeSafe SDK Bifrost docs
  19. Structured model outputs OpenAI
  20. GPT-4 Technical Report arXiv
  21. Jev, an AI Model That Can't Chat, Takes On Bigger Rivals Bloomberg
  22. Shut up and calculate: Jev's new AI primitives for coders The Register

Frequently asked questions

What is Jev from TypeSafe AI?

Jev is TypeSafe AI's first public model, released in early access on 15 September 2026. It reads text state and a set of typed questions and returns structured answers with probabilities in a single request. It does not chat, write text or generate code.

What does System One model mean?

It is TypeSafe's name for a class of models built to make fast, structured decisions for software, borrowing Daniel Kahneman's System 1 (fast, intuitive) versus System 2 (slow, deliberate) distinction. You define the possible answers in advance; the model returns a probability for each.

How much does Jev cost?

As of September 2026 TypeSafe charges $0.042 per million input tokens for jev-1.13.0, and output tokens are free. TypeSafe says it cannot yet prove the price is not subsidised.

Is Jev an LLM?

TypeSafe says Jev is neither small nor an LLM. An independent black-box study by Archer Hume inferred a causal transformer backbone that reads probabilities directly from internal representations rather than generating tokens, and found its tokenizer closest to Qwen's. TypeSafe has not published the architecture.

How fast is Jev?

TypeSafe's launch post quotes 70 to 500 milliseconds end to end. Independent probes measured roughly 57 to 218 ms as the state grows to about 30,000 tokens, and a p95 of 397 ms for a single question called from Seoul.

Who founded TypeSafe AI and who invested?

TypeSafe was founded in 2024 in San Francisco by CEO Diogo Almeida, an author of OpenAI's InstructGPT paper, with CTO Erik Gafni and COO Sasha Sheng. DCVC led the $40 million seed round announced on 15 September 2026.

Can Jev replace structured outputs from an LLM?

For bounded decisions such as routing, classification, scoring and guardrails, it is designed to. For anything that needs generated text, extended reasoning, arithmetic or date comparison, TypeSafe's own documentation says to use code or a generative model instead.

All news →