— Pilot Plan

We Are Testing Jev, a Non Generative Model, for CMS Content Model Scoring. Here Is the Plan.

|

TypeSafe AI released Jev on 15 September 2026. We plan to move content model assignment in CMS migrations onto it, while SEO, AEO and Visits scoring and the migration verbs stay unchanged. This post sets out the design, the early evidence and the limits.

Jev is a non generative decision model from TypeSafe AI that reads a block of text and returns typed answers with a calibrated confidence and no prose. In a CMS migration, that makes it a candidate for assigning each URL to a content model, which is a classification job.

Jev, released by TypeSafe AI on 15 September 2026, is a decision model that returns choices, scores or yes/no probabilities with a confidence value and no generated text. Dotfusion plans to move content model assignment in CMS migrations onto Jev, leaving SEO, AEO and Visits scoring and migration verbs unchanged.

What is Jev, and what does it return?

Jev is a decision model that TypeSafe AI released on 15 September 2026, priced at $0.042 per million input tokens with free output. Users send a block of state plus typed questions. Jev answers each one as a Choice from options the user defines, a Score on a bounded numeric scale, or a yes/no probability, which TypeSafe calls a Noul. Each answer carries a confidence that TypeSafe describes as calibrated, with no text and no rationale.

Founder Diogo Almeida is a former OpenAI researcher who co-authored the 2022 paper that established reinforcement learning from human feedback as a key LLM training method, TechTarget reports.

TypeSafe reports 70 to 500 milliseconds end to end latency. Access is through an API and the Vercel AI Gateway, where nearly 13% of paid teams used Jev within 24 hours of launching on the gateway, more than twice as many as any previous model launch there, Vercel reported on 18 September 2026. Jev has shipped in early access, behind a waitlist, hosted only.

Where does the work go in a migration?

On a migration with 20,000 URLs, the work goes into assigning a content model and a component inventory to every one of them. Per URL we ask which content model it belongs to, and whether a hero, FAQ block, embedded form, author byline, video and related content are present. These are classification questions.

Today we run that work as agent batches on generative models and parse their answers. Our discovery method proposes candidate content models, writes falsification rules that can reject each one, and aims for a 95% fit target across the inventory. On WordPress sources the wp_body_class value is ground truth for the template, and the H2 sequence acts as a template fingerprint.

The groundwork is in our guides to running a content audit first, preparing your content for a new platform enterprise content migration best practices and content operations for enterprise.

How will we use Jev in content model discovery?

The plan below is a design for a pilot that has not run yet. Per URL we will send Jev the title, H1, H2 sequence, a body excerpt and the body class. We will ask one Choice question over the candidate content models and six yes/no questions: hero, FAQ block, embedded form, author byline, video and related content. We will accept answers at a confidence of 0.9 or higher, send the rest to human review, and log every probability in our Decision Log.

TypeSafe's docs, last reviewed 2 October 2026, say accuracy falls as the state fills with irrelevant detail and that Jev reads questions literally. So we will send an excerpt and not the page, and keep questions short and exact. Option order can bias a Choice, so we will reorder the options and check that the answer holds. Counting and comparing stay in code.

This is our own arithmetic at published limits, not a measured run. TypeSafe lists a limit of 100,000 tokens per second and says limits can change without notice.

We assume 1,500 to 5,000 tokens per URL; Lindfors's example request used 4,995. A 20,000 URL run is then 30 to 100 million tokens: about 5 to 17 minutes, and $1.26 to $4.20 of input at $0.042 per million. After a rule change we will rerun the whole inventory and report the measured time in the follow-up.

We plan to pilot Jev on a live Contentful migration this month. If agreement exceeds 95% with a review pile below 10%, Jev becomes the standard classifier in our content model discovery. We will publish the results in a follow-up post.

What does Jev not decide?

In our plan, Jev assigns a content model and component flags to each URL and makes no other call. SEO, AEO and Visits scores stay arithmetic over Google Search Console, GA4 and logged prompt runs, and the subjective score belongs to the client. The migration verbs (keep, consolidate, rewrite, redirect, retire, create) stay rules a person can repeat in a client meeting.

Tom's Hardware says the developer, not Jev, is responsible for acting on the confidence factors, and that Jev can still misclassify, fall victim to adversarial attacks, or answer literal wording rather than meaning.

What are the honest caveats?

Jev is three weeks old, in early access behind a waitlist, and hosted only. In TypeSafe's own four workflow evaluations, Jev scored 67.8%, GPT 5.6 Terra 67.9%, Claude Opus 5 73.1% and GPT 5.6 Sol 74.1%, the highest of the four. TypeSafe's team wrote those workflows and scored them against the average of GPT 6 Astra and Fable 5.1.

One small independent test used 24 Norwegian hearing responses, run by Emil Lindfors on 18 September 2026 and summarized by APIMaster. Jev tied DeepSeek V4.1 Flash with reasoning off on stance, 20 of 24 each, and beat it on the ordered scale. It trailed slightly on respondent type, 21 of 23 against 22 of 23, and on 192 yes/no argument judgments, 0.86 against 0.89.

Of those 192, the 43 that Jev scored at 0.9 to 1.0 agreed with the reference 98% of the time. On Choice questions at 0.9 or higher, 34 of 35 agreed, but 12 of 47 Choice answers (26%) fell below 0.9, against a review pile bar of 10% in our pilot (our sums from the Lindfors tables). Lindfors calls 24 documents a first look and not a benchmark.

Jev cost $0.22 per 1,000 documents, a sixth of DeepSeek's cheapest run
Measure Jev 1.13 DeepSeek V4.1 Flash, reasoning off DeepSeek V4.1 Flash, reasoning on
Cost per 1,000 documents $0.22 $1.31 $3.08
Median latency 0.32 s 2.7 s 26 s
Substance, exact level 19 of 24 14 of 24 14 of 24
Stance, 4 options 20 of 24 20 of 24 22 of 24
Respondent type, 6 options 21 of 23 22 of 23 23 of 23
Arguments, 192 yes/no 0.86 0.89 0.88

Source: Emil Lindfors, 18 September 2026, lindfors.no/blog/a-first-look-at-typesafes-jev.

Output is not deterministic on rerun. The format is fixed and the content is not, so the same input can return a different answer or confidence, TechTarget reports. Laurie Voss, head of developer relations at Arize, told TechTarget he expects the major model labs to release decision models very quickly, and Hackaday reported on 6 October 2026 that others are building decision models too, naming Kev and Nimble as two examples.

So we design the classifier step to be swapped out. The Jev call will sit behind one interface, we will pin the model version in every call because TypeSafe's docs say aliases move when new releases ship, and we will log probabilities so a replacement can be scored against the same labels.

Method note: figures here are vendor reported or come from one 24 document test run on 18 September 2026 against jev-1.13.0. Prices and limits are as published on 8 October 2026. Agreement is with labels from a frontier model, which is not proof of correctness. All figures are as of 8 October 2026.

What do readers ask about Jev in a migration?

What is Jev?

Jev is a non generative decision model from TypeSafe AI, released on 15 September 2026. It reads a block of text and typed questions, then returns choices, scores or yes/no probabilities with a calibrated confidence and no prose. Input costs $0.042 per million tokens.

Is Jev a replacement for LLMs?

Jev is not a replacement for LLMs. It returns decisions and not text, and TypeSafe's docs send generation to a generative model. TypeSafe positions it for classifying, scoring and routing, and we plan to test it on assigning URLs to content models.

How does a decision model help a CMS migration?

A decision model assigns every URL a content model and component flags, each with a confidence. At TypeSafe's list price, our arithmetic puts a full rerun of 20,000 URLs at $1.26 to $4.20 of input, and cases below our 0.9 threshold will go to people.

Does Jev decide what content gets retired?

Jev does not decide retirement. Retire is a migration verb set by rules a person can repeat in a client meeting, and by SEO, AEO and Visits scores computed from Search Console, GA4 and logged prompt runs. Jev supplies only the classification.

This post was conceived and written by Chris Bryce with our stack of AI research agents doing the heavy lifting on sources and data. If you want to talk about enterprise web strategy, AEO, or content operations, talk to us. We love this stuff.

Sources

  1. TypeSafe AI, Introducing System One Models & Jev (launch post, 15 September 2026)
  2. TypeSafe docs, Models
  3. TypeSafe docs, Jev 1.13 jaggedness (last reviewed 2 October 2026)
  4. Vercel, Jev is the fastest-adopted model in AI Gateway history (18 September 2026)
  5. Emil Lindfors, An early-access test of TypeSafe's Jev: calibrated judgments for half a cent (18 September 2026)
  6. TechTarget, Jev decision model touted as quicker, cheaper LLM alternative (22 September 2026)
  7. DataCamp, Jev: TypeSafe's System One Model Explained (16 September 2026)
  8. Tom's Hardware, TypeSafe AI's Jev offers an alternative to LLMs that claims to be 193x faster and 445x cheaper (21 September 2026)
  9. APIMaster, Jev vs LLMs: Where a Decision Model Beats Prompting, and Where It Doesn't (20 September 2026)
  10. Hackaday, A New Type Of LLM On The Block: Decision-Making Models (6 October 2026)