TL;DR:
On 15 September a San Francisco startup called TypeSafe came out of two years of stealth with a model called Jev. It doesn't write text. At all. You hand it some information and a list of questions, and it hands back the answers with a probability attached to each one.
The founders are Diogo Almeida, Erik Gafni and Sasha Sheng, and they raised a $40m seed led by DCVC before shipping a thing. Almeida is the name to know: at OpenAI he co-invented RLHF, the training method that turned raw language models into every chatbot you've ever used.
Developers lost their minds over it. Vercel said it was the fastest adoption in the history of their AI gateway — 13% of their paying teams inside 24 hours.
Hashi's take: the speed got the attention, but the real story is in how Jev answers. It doesn't write. It decides, and it tells you how sure it is. That's a vital piece of making AI-generated answers reliable.
STAT WORTH SHARING
TypeSafe says Jev is 193.6x faster than GPT-5.6 Terra. Independent developer testing puts the median at 7x.
Know someone who's been talking about Jev? Send them this issue. Share your personal link and when they join, you earn rewards: an AI Toolkit Cheat Sheet, an exclusive Deep Dive report, and a 1-on-1 AI strategy chat.
What Jev Actually Is
Start with what every AI you've used does today. You ask it something and it writes you an answer, one word at a time. Even when you ask for a tidy answer in a fixed format, it's still writing — it's just writing the punctuation too.
Jev doesn't write anything.
You give it two things. The information the decision is about, which TypeSafe calls the state. And your questions, with the answers it's allowed to give. It hands back the picked answers and a number next to each one, between 0 and 1, saying how confident it is.
No sentences. No paragraphs. Picked answers and confidence scores.
Three kinds of question, and that's the whole menu:
Noul — a yes or no, with a probability. Is this a refund request?
Choice — one option from a list you define, up to 255 of them. Which team handles this: billing, technical, or sales?
Score — a rating on a scale you define. How urgent is this: low, medium, high, critical?
TypeSafe teaches those three in its own playground, using deliberately daft questions. Is a hotdog a sandwich. What colour is the sky. Can monkeys create art.
Run the hotdog one and Jev answers 77% true. Not a verdict. Not a paragraph weighing it up. A number saying how much of a sandwich a hotdog is. The whole thing took under a quarter of a second.

The question on the left, the answer and its probability on the right. Credit: TypeSafe AI
Look at the left-hand side and you can see the other half of the job, which is yours rather than the model's. You write the criteria. What counts as true, what counts as false. The model doesn't decide what a sandwich is — you do, and it tells you how well this thing matches.
Scroll past the jokes in that playground and the use cases TypeSafe lists are the tell: screening a CV, auditing a support agent's chat session, catching jailbreak attempts against a chatbot. Judgement calls with consequences, all of them.
Run a real one through it. A customer emails to say they've been charged twice this month and would like it fixed.
You send Jev the email, plus that customer's recent charges. You send all three questions at once, in a single request. Back comes: refund request, 0.90. Billing team, 0.85. Urgency, high.
A person skimming that email would have known all three instantly, without thinking about it. That's the kind of work Jev is built for.

Fast Thinking and Slow Thinking
TypeSafe calls this a System One model, and the name is borrowed from Daniel Kahneman's Thinking, Fast and Slow.
Kahneman split human thinking in two. System one is fast and automatic — ask me what two times two is and the four arrives before I've decided to work it out. System two is slow and deliberate. Ask me what seventeen times twenty-four is and I have to actually sit down with it. It's 408, and I had to check.
Nearly every AI model you've heard of is built for system two work, or is at least pretending to be. A chatbot writes its answer out token by token. A reasoning model goes further and writes out its thinking first, so you can watch it work — which is about as close to deliberate thought as AI currently gets.
But most of the decisions inside a business aren't system two. They're quick calls. Is this invoice a duplicate? Does this message need a manager? Is this claim straightforward or complicated? Nobody sits down with a cup of tea to decide those.
We have been using the slow, expensive, deliberate machine to do fast, automatic work. That's the gap TypeSafe went after.
Why Everyone's Talking About It
Three reasons.
It's faster and cheaper by a lot. Responses come back in 70 to 500 milliseconds. Input costs $0.042 per million tokens and output tokens are free, which is possible because there is no output text to charge for. TypeSafe's own benchmark claims 193.6x faster and 444.6x cheaper than GPT-5.6 Terra. Independent numbers are smaller — more on that below — but even the conservative ones are a step change.
Developers actually adopted it. Vercel reported the fastest uptake in the history of their AI gateway: 13% of paying teams within a day of launch. That's not a press release number, it's a platform reporting on its own traffic.
And they're building real things with it. Made With Jev, a directory run by a developer with no connection to TypeSafe, already lists more than 750 projects — support triage, security checks, document sorting, product search, trading tools. One of them classified 500 emails for three and a half cents.
And then there's Almeida. RLHF is his own work, and what it does to a model is the reason Jev exists. You train a model by having people pick which of two answers they prefer. People prefer answers that sound confident. So the model learns to sound sure of itself, whether or not it is. Almeida's own verdict on what he helped build: RLHF chatbots suffer from "mode dropping, overconfidence, and an overall lack of reliability."
His line on why a chatbot is the wrong shape for software is the best sentence written about this all month. LLMs generate sequentially, "which is great for a conversation, but totally useless for computers."
ATTIO
Quick note: this issue is supported by Attio. A model that decides in 70 milliseconds is only as good as the state you hand it — and that's harder when your customer record is spread across four systems than when it lives in one. Their CRM is built for AI-native teams.
The agentic era needs a different CRM. That’s Attio.
Teams like Parallel, Turbopuffer, and Wordsmith are already setting the pace on Attio. Get an always-on revenue engine, with agents and workflows that build pipeline, chase every buying signal, and move deals forward with your team. Whether you're working in your browser, inbox, or favorite agent, connect to your customer data in real-time through Attio's web app, MCP, API, and SDK.
Does Jev Help With AI Safety?
Yes — though in a narrower, more practical way than that word usually suggests.
You can ask an ordinary LLM for a tidy structured answer today. Plenty of teams do. You can even ask it how confident it is, and it'll give you a number.
That number means very little. It's the model writing a number because you asked for one, not the model reporting what it actually calculated.
Jev's numbers are different, because of how it's trained. TypeSafe calls the method RLCD — reinforcement learning for calibrated decisions. In plain terms, the model is rewarded for its probabilities turning out to be right, not just its answers. Calibrated means what it sounds like: when it says 80%, it should be right about 80% of the time.
Which turns a confidence score into something you can write a rule on.
Back to that refund email. Set a threshold at 0.9. Anything above it goes straight to the refund queue, automatically. Anything below 0.1 isn't a refund request, so nothing happens. Everything in between — where the model isn't sure — goes to a person.
Because the numbers are calibrated, that threshold also tells you roughly how often the automatic path will be wrong. The more a mistake costs you, the higher you set the number. That's a governance decision, not a technical one.

Last issue I argued that the way to run agents safely is to decide which actions can't be taken back, and make those ones stop and ask. The question I left open was how the software knows when it's unsure enough to stop.
This is that. A calibrated probability is the trigger. It's the first time I've seen the stop built into the model rather than bolted on afterwards.
Before You Believe the Numbers
Let me set a little reality check before you go running and building things in your business around this.
The headline benchmarks are TypeSafe's own. The 193.6x figure comes from their internal workflow evals, and the company says openly that it can't prove the pricing isn't subsidised. Independent testing is a lot more modest: a Vercel engineer measured 5 to 18 times faster, and an analysis of nearly 13,000 developer posts found a median 7x speedup and 30x cost saving. Genuinely good. Not two hundred times.
"Zero hallucinations" doesn't mean what it sounds like. TypeSafe can guarantee the shape of the answer — it will always be one of the options you defined. It cannot guarantee the answer is right. It will hand you a confidently wrong valid option, same as anything else. If someone quotes that phrase at you in a vendor meeting, this is the question to ask.
It's narrow on purpose. Text in only, no images yet. It's unreliable at counting, arithmetic and comparing dates — TypeSafe's own documentation tells you to keep maths in your own code. Accuracy drops on large, noisy inputs. And like any model reading data from outside, it can be fooled by instructions hidden in that data.
Nobody outside TypeSafe knows how it works. No published architecture, no weights, no parameter count, no self-hosting. It's early access behind a waitlist, and you're trusting a two-year-old company with $40m of seed money and one product.
None of this disqualifies Jev. Focused is what it's built to be. But it's new, and I'd want to watch it mature before putting anything that matters on top of it.
Final Thoughts
Jev is named after William Stanley Jevons, who noticed something in 1865 that's still the most useful idea in technology economics. As steam engines got more efficient, Britain burned more coal, not less. Cheaper meant more, not cheaper.
Jev isn't going to replace your chatbot. But every business runs on thousands of small judgement calls a day, and nearly all of them wait. Not because they're hard. Because somebody has to get to them.
Hand those to something that decides in 70 milliseconds, for a fraction of a penny, with a number telling you when to step in. Then scale that across a company. Across a working day. Across the hundred small decisions you make outside work too.
How much faster could we move?
Keep reading!
If this issue was useful, pass it on. One colleague joining counts, and your rewards start from there.
Copy, paste and send to someone on your team:
I've been reading The Context Window, a 5-minute weekly on practical AI for business leaders. This issue on Jev is a good place to start. You can join here: {{rp_refer_url}}
Your link is also in the box below, with your reward progress.





