Notes on Jev, a model that won't write you a sentence
TypeSafe's Jev answers with a typed value and a probability, never text. What it's actually for, where it fits in an agent loop, and which launch numbers deserve an asterisk.
Watch an agent work for a few minutes and count how many of its model calls produce actual writing. It's fewer than you'd think. Most of the loop is the model answering small questions for the harness. Which tool next? Did the test pass? Is this command safe to run? Are we done yet?
Each of those goes to a model that could write a novel. We wait a second or two, fish one word out of the reply, and hope it didn't wrap that word in a friendly sentence. It works. It's also a strange way to build software, and if you build agents for a living you've probably felt it.
On 15 September a startup called TypeSafe AI came out of two years in stealth with a model that only does the small questions. It's called Jev. You tell it what the possible answers are, and it picks one and tells you how sure it is. It never writes a sentence back. That sounds like a limitation until you remember how much of your harness exists to dig answers out of chat replies.
These are my notes after going through TypeSafe's docs, their evals, and the first independent paper on it.
What you send, what you get back
A Jev call has two parts. The state is whatever you want judged: a support ticket, a tool call your agent is about to make, a search result, a blob of JSON from your app. The questions are what you want to know about it, and each one arrives with its possible answers already written out.
There are three kinds of question:
- A Choice picks one option from a list you give it, up to 255 options.
- A Score places the state on a scale you define, like calm, frustrated, very angry.
- A Noul is TypeSafe's name for a yes/no question. You get back a number between 0 and 1.
You can mix all three in one request. They're evaluated in parallel and independently, so asking five questions takes about as long as asking one, and one answer never leaks into another. Click through these:
Send a state, get typed answers
State
Help! My payouts have been failing for 3 days.
Questions, and what came back
department · choice
confidence 0.81Which team should handle this?
is_urgent · noul
Does this convey urgency?
frustration · score
confidence 0.92How frustrated is the customer?
score 1.05 (weighted across levels)
What your code does with it
Route it to billing and bump the priority. Confidence is over 0.8, so nobody has to look at it first.
The numbers for this one come from TypeSafe's API docs.
The ticket example is lifted straight from TypeSafe's API docs. The other two are mine and the numbers are made up, so read them as the shape of a response rather than a benchmark.
The boring output is the point
If you've ever told an LLM to "respond only with JSON," you know how this goes. It mostly works. Then one day it wraps the JSON in a code fence, or opens with "Sure! Here's the JSON:", or invents a fourth department called "general" that you never mentioned. So you write a parser. Then a retry. Then a validator for the parser.
None of that can happen here. The options you define are the entire output space. Ask Jev to choose between billing, technical and sales, and you get one of those three strings plus a probability for each. There's nothing to parse because nothing was generated.
TypeSafe's launch post says Jev "can't hallucinate." I'd phrase it more carefully. It can't make up an answer outside your list. It can absolutely pick the wrong one from inside it, and it'll hand you a number while it does. The number is what makes that livable.
The number next to the answer
Choice and Score answers come with a probability for every option and a single confidence value between 0 and 1. The claim is that the probabilities are calibrated: take all the answers Jev gave a probability of 0.8, and about 80% of them should turn out right.
That sounds like a footnote. I think it's the whole product. The launch post puts it well: "If a model can do a task 95% of the time but doesn't say when it's in the 5%, it can't automate that task." You'd have to check everything, which defeats the point. A model that can say "not sure about this one" lets your code do the sensible thing: act on the confident answers, and send the rest to a bigger model or a person.
Drag the threshold, then flip to the overconfident model.
Where do you draw the line?
- Handled automatically
- 43 / 60
- Wrong, and nobody checked
- 3 (7%)
- Sent to a person or bigger model
- 17
60 simulated decisions. In the calibrated batch, answers marked 0.8 are right about 80% of the time. The overconfident model reports 0.88 or higher on everything, right or wrong.
With the calibrated batch, raising the bar does what you'd hope. Fewer mistakes get through, and you pay for it by escalating more. With the overconfident one, the slider does nothing until you get close to 0.9, because every answer claims to be at least 0.88 sure. Ask a chat model to rate its own confidence and you get something like the second picture. It says 90% a lot.
Why an RLHF co-inventor built a model that doesn't try to please you(optional)
TypeSafe's co-founder Diogo Almeida co-invented RLHF, the training method behind InstructGPT and ChatGPT. (I walked through how it works in my alignment post.) RLHF rewards the answers people prefer. That's great for a chatbot and bad for a decision system, because people prefer answers that sound sure of themselves.
Jev is trained with what TypeSafe calls RLCD, reinforcement learning for calibrated decisions. Instead of chasing what a human would like to read, training pushes its probabilities to match how often it's actually right. Their docs are careful to add that calibration describes a batch of predictions. It's not a promise about any single answer.
Where it goes in an agent loop
The mental model I've landed on is short. An LLM writes, Jev decides, code acts.
If a step produces words for a person (a reply, a summary, a commit message), it's an LLM's job. If it picks from options, scores something, or answers yes or no, it's a Jev question. And if it follows an exact rule, like date math, counting, permissions or anything a regex can do, it belongs in plain code with no model anywhere near it. People on X were calling this "Jev engineering" within days of the launch. The name is new. The idea underneath it isn't: keep the model on a short leash and let code own the control flow.
Try sorting a few real agent steps:
LLM, Jev, or plain code?
Pick which of 40 tools the agent should call next
Write the commit message for the change
Check whether an invoice is more than 30 days overdue
Judge whether a shell command would delete files
Explain a failing test log to the user
Count how many changed files live in src/auth
Decide if a retrieved passage actually answers the question
If it writes words, it's an LLM. If it picks, scores or says yes/no, it's Jev. If there's an exact rule, it's code.
If I were adding Jev to an agent tomorrow, these are the four places I'd start.
Picking the tool. Instead of handing the LLM forty tool schemas and asking which one it fancies, ask Jev a Choice over the tool names and let the LLM write arguments only for the winner. If Jev isn't confident, fall back to the normal LLM step. TypeSafe's function-calling cookbook goes a step further and has Jev fill in any argument that comes from a fixed list, like a ticker symbol or a time window.
Guarding the tool call. Before a command runs, ask a few yes/no questions about it. Does it delete anything? Does it reach outside the project? Is it actually what the task asked for? Then let your code decide what happens at 0.95 versus 0.4.
Checking the work. "Is the task done?" and "is the agent going in circles?" are judgment calls over a trace, and an agent asks them constantly. That's exactly where cheap and fast matter most.
Routing. Decide which model a request goes to, or whether it needs a model at all. Some requests are a database lookup wearing a trench coat.
Here's what the guard looks like with TypeSafe's TypeScript SDK:
import { noul, TypeSafeClient } from "@typesafe-ai/sdk";
const typesafe = new TypeSafeClient(); // reads TYPESAFE_API_KEY
export async function screenCommand(task: string, command: string) {
const { answers } = await typesafe.systemOne({
state: { task, command },
questions: {
deletes_data: noul("Would running `command` delete or overwrite files?"),
leaves_project: noul("Does `command` touch paths outside the project directory?"),
on_task: noul("Is `command` a reasonable step toward `task`?"),
},
});
if (answers.leaves_project.noul > 0.2) return "ask_human";
if (answers.on_task.noul < 0.5) return "reject";
if (answers.deletes_data.noul > 0.5 && answers.on_task.noul < 0.9) return "ask_human";
return "run";
}Jev never decides whether the command runs. It hands back three numbers, and four lines of ordinary code make the call. You can read that policy in a code review, which is more than you can say for a paragraph buried in a system prompt.
TypeSafe also ships an agent skill for Claude Code, Codex and friends, which tells you who they expect the first users to be. There's even a page in the docs explaining that no, you can't set Jev as the model inside Cursor. I assume they got that question a lot.
Asking good questions
The API is small. Almost all the work is in writing the questions, and TypeSafe's docs are refreshingly blunt about how to do that well.
The big one is to ask one thing per question. "Is this ticket spam?" hides three separate judgments. Does it ask for a password? Does the sender's name match the email domain? Does it announce a prize nobody asked for? Each of those is easy on its own, and you can weight them in code however you like. When the weights turn out to be wrong, you change a number instead of rewriting a prompt.
Send less. Accuracy drops as the state fills up with detail that has nothing to do with the question, so filter in code before you call.
Ask everything at once. Because questions run in parallel, tacking on a few you might not need is cheaper than a second round trip later. TypeSafe calls this speculative fan-out, which is a very serious name for "ask while you're there."
Write the edge cases into the options. Each option can carry a description of what it covers, what it's not for, and a couple of examples. The docs' own triage example does this for every option, and it reads a lot like writing good test cases.
Let thresholds follow the stakes. Showing someone the wrong screen at 0.6 is fine. Approving a bank transfer at 0.6 is not.
And once you've tuned those thresholds, pin the model version. jev-latest moves when a new release ships, and your carefully chosen 0.85 might not mean the same thing next month.
Where it falls over
My favorite page in the docs is the one listing the model's known failure modes. More labs should publish one. For jev-1.13, the current version, it says:
- It reads literally. It answers the question you wrote, not the one you meant.
- It's not a calculator, and it can't count. Keep arithmetic in code.
- It reads dates as text, so "which came first?" is unreliable. Pull out the parts with Jev, compare them in code.
- Instructions planted in the content can move its answers. It treats state as data, but not as hostile data.
- A question and its opposite don't have to add up. In one of their own examples, "is this a refund request?" came back 0.72 and "is this something other than a refund request?" came back 0.47. That's 119%.
The fourth one matters if you're planning to use Jev as a safety guard, which is one of the most obvious things to do with it. A guard that can be argued with needs other defenses behind it: an allowlist, a sandbox, a person for anything you can't undo.
Then there are the plain limits. It's text only, so no images. English works best and other languages less so. A request tops out at 64k tokens, with 32k for the state plus your longest question. A Choice maxes out at 255 options and a Score at 10 levels.
About those numbers
Depending on which TypeSafe page you're reading, Jev is 40x to 200x faster than frontier LLMs, or "two orders of magnitude" faster, or, in the homepage demo, exactly 193.6x faster and 444.6x cheaper. Pick your favorite.
The number I trust most is the price, because it's just arithmetic. Jev costs $0.042 per million input tokens, and output is free. Here's what that looks like at your volume:
What does a decision cost?
Decisions per day
LLM input price, per million tokens
94x cheaper. Now the gap is real money. This is the kind of volume Jev was priced for.
Jev: $0.042 per million input tokens, output free. LLM: output priced at 5x input (TypeSafe's own rule of thumb) and 50 output tokens per decision (my guess). 30-day month.
The name, by the way, is a nod to William Stanley Jevons, who pointed out in 1865 that more efficient steam engines made Britain burn more coal, not less. TypeSafe's bet has the same shape: make decisions cheap enough and people will make a lot more of them. Which is a polite way of saying your bill might not actually go down.
The gap is enormous against frontier models and much smaller against cheap ones. Against a model that charges $0.20 per million input tokens, it shrinks to somewhere around 5 to 10x, and at that point you're choosing on accuracy and calibration, not cost.
TypeSafe says openly that it can't prove its pricing isn't subsidized, and that its published evals were mostly run from laptops on the US West Coast, which is also where the service lives. I like that they said it. It also means you should measure latency from wherever your servers actually are.
Their workflow eval scores every model against the averaged answers of two big reasoning models (GPT-6 Astra and Claude Fable 5.1, both on high thinking), not against human labels. Across their four example tasks, it also found that every model, not just Jev, got more accurate, cheaper and faster once the task was broken into narrow questions. So part of the win is Jev, and part of it is simply asking better questions. You get that second part with whatever model you already use.
Rate limits sit at 1,200 requests a minute for now, and the docs warn they can change without notice while TypeSafe adds GPUs. It's still early access.
The independent evidence so far is encouraging, with caveats attached. A paper on an agent design called REFLEX used Jev for the bounded decisions and handed off to a strong LLM whenever Jev wasn't confident or something had to be written. On the authors' 100-task benchmark it hit 95% success with 72.7% fewer calls to the expensive model. On public function-calling benchmarks, where routing was already easy, it didn't do much better than a cheap LLM cascade. They also found reliability depended on how many actions were on the table, and on how close a plausible wrong option sat to a permission boundary. That second case is the tool guard from earlier, so test it hard.
A separate hands-on test by Manjunath Janardhan put the same 200 decisions to Jev and five frontier LLMs. Two of the LLMs beat it on accuracy, at many times the price.
So, should you use it?
If you build agents, I think Jev is worth an afternoon. Not as a replacement for your main model (it can't be one) but as the thing that answers all the small questions around it. Pick one decision you already make with an LLM call. Run a few hundred real examples through both. Plot confidence against accuracy, and choose your thresholds from what you see, not from anyone's launch post. Including this one.
I wouldn't use it for anything that needs writing, math or dates, and I wouldn't make it the only thing standing between user input and an action you can't undo.
One practical warning. A lot of unofficial "Jev" sites popped up in the first two weeks, some offering their own API keys and MCP endpoints. Some of them are probably fine. Either way, the real docs are at docs.typesafe.ai, keys come from console.typesafe.ai, and I wouldn't paste a key anywhere else.
Sources
- Introducing System One Models and Jev, TypeSafe's launch post
- TypeSafe docs, especially Confidence, Patterns and Jev 1.13 jaggedness
- Models, pricing and limits
- Workflow evals
- REFLEX with Jev for Efficient Selective Control in LLM Agents, arXiv
- I Tested TypeSafe's Jev Against Claude, GPT-6, Kimi, MiniMax and DeepSeek, Manjunath Janardhan