← Back to writing

Jev vs OpenAI's Decisions API: where Jev still wins, and what nobody has proven yet

September 30, 2026 Versão em português

Jev vs OpenAI's Decisions API: where Jev still wins, and what nobody has proven yet

Eight days apart. TypeSafe released Jev on September 21, and on September 29, at DevDay, OpenAI announced the Decisions API. Both do the same basic thing. You describe a question and the answers you'll accept, and you get back a decision your code can branch on, instead of a paragraph you have to parse.

Earlier this month I benchmarked Jev at Battleship, so I wanted to know where it stands now that OpenAI has its own version.

Today, Jev is still ahead. Part of the reason is simply that OpenAI's version isn't really out yet, so I want to be clear about which parts of this are measured and which parts are just what each company has published.

Same category, different machine

The two products share an interface idea and not much else.

Jev is its own model. TypeSafe calls it the first "System One" model, built from the start to evaluate typed questions against a state and return probabilities. It doesn't generate text at all. Simon Willison prefers the term "decision model", and I'll use his.

The Decisions API is an endpoint in front of a specialised version of GPT-6 Luna, OpenAI's general model. You still define the question and the allowed answers, and OpenAI keeps the output inside those answers. Behind the endpoint, though, sits a frontier model that can also write you a sonnet.

That difference explains most of what follows. It's why OpenAI can accept images on day one. It's also why its headline speed claim, roughly 150 ms and about ten times faster, is measured against asking GPT-6 Luna the same thing through the regular API. The comparison is Luna against Luna. OpenAI hasn't put it next to Jev.

Where Jev wins today

You can actually use it

The Decisions API is in limited preview. Most of us can read the announcement and that's about it.

Jev is public. My Battleship benchmark made 6,000 calls to it and cost 77 cents, and the code is open source so anyone can rerun it. I can't run that benchmark against OpenAI, and neither can you unless you're in the preview.

The price is public

Jev charges $0.042 per million input tokens and nothing for output. Simon pointed out that's below even GPT-5 Nano's input price. OpenAI hasn't published pricing for the Decisions API yet. With Luna underneath I'd be surprised if it came in cheaper, but that's a guess, and I'll treat it as one until there's a pricing page.

More kinds of question

Jev has three question types. A Choice picks from a list and returns a probability for each option. A Score rates something on a numeric range you describe. A Noul returns how likely a statement is to be true, from 0 to 1. You can mix all three in a single call, and each question is evaluated independently and in parallel, so ten questions take about as long as one, and adding more doesn't pollute what the others see.

That changes how you build with it. The TypeSafe docs push you to break a big judgement into small ones and combine the answers in your own code. A support ticket becomes "does this message ask for a refund?", "how frustrated is the customer, from 0 to 10?" and "which team owns this?", all asked against the same state, in one request.

Everything OpenAI has published so far describes one shape: a question with predefined answers. That covers classification and routing, which is most of what people will use it for. For a yes/no probability or a numeric score, you'd have to fake it with buckets of answers.

People are already building on it

In nine days Jev picked up a plugin for Simon's llm tool, so this works straight from the terminal:

llm -m jev 'Please refund my last payment.' -s 'Does this message explicitly request a refund?'

Josh Rosen wrote about using it inside ThruWire as a checkpoint for agent work. The agent does whatever it wants between checkpoints, and at each one Jev judges whether the evidence actually supports the claim before the work moves on. TypeSafe's CEO shared notes on coding agents with ideas like asking a yes/no question about every chunk of context to decide what the next turn really needs to see. There's even an open-weight imitation already, called Kev.

This is the part I find most interesting. The use cases people reach for first are plumbing inside agents, where one task needs hundreds of small judgements and each one has to be cheap and fast. A small dedicated model should beat a frontier model behind an endpoint at exactly that job.

Where OpenAI could win

I don't want to write the version of this post where Jev wins everything, because the data doesn't say that.

Images are the obvious one. Jev takes text. If your decision depends on a screenshot, a scanned invoice or a photo of a damaged parcel, OpenAI is the only one of the two that can look at it (in preview, for now).

Update, 6 October 2026: a day after this went out, Cloudflare released Clef, an open-weight decision model that uses the same API as Jev and accepts up to four images per request. A dedicated decision model can now look at images too. I've since benchmarked Clef against Jev.

The other is the model behind it. Jev has documented weak spots, and TypeSafe's own model jaggedness page lists numbers, dates and adversarial content. In my Battleship run it rounded probabilities to two decimal places, so across 90 options about half came back as zero and most of the rest tied at 0.01, leaving a usable ranking over roughly a dozen cells. Given the raw board, it lost to a heuristic you could write from memory. Luna has far more general knowledge behind it, and handed a long, messy context it might simply read it better. Nobody has shown that yet, but it's the most plausible way OpenAI wins on quality.

Correction, 6 October 2026: an earlier version of this paragraph said the distribution came back mostly as zeros. A later probe showed about 45 of 90 options come back non-zero and the probabilities still add up to 0.996. What the rounding loses is the order of the tail.

And some teams will pick OpenAI because they already pay OpenAI. One vendor, one bill, one set of compliance paperwork. It's a boring reason and a real one.

What nobody has proven yet

The biggest open question is whether either "confidence" means anything. The Hugging Face guide to the Decisions API says it plainly: don't assume a field called confidence is a calibrated probability. That applies to Jev as well. Simon asked it to rate Bay Area cities and got Cupertino at the top and East Palo Alto at the bottom, which is the kind of pattern you don't want sitting quietly inside a hiring or lending decision.

Then there's the simplest question of all, which one is better at the same job. I haven't found a single head-to-head benchmark. Every comparison table I read, including the ones in the guides above, compares features and availability. None of them compares results.

And the Decisions API is a preview. Endpoint names, limits, pricing and what the scores mean can all change before it ships to everyone.

What I'd do right now

If I were building on a decision model this week, I'd build on Jev, because it's the one I can call, measure and pay for. I'd keep the call behind a small interface in my own code, so swapping in the Decisions API later is a change in one file.

And I'd carry over the one lesson from Battleship that has nothing to do with which vendor wins. Jev only looked good once my code had already cut the problem down to 16 described options, and even then it tied with plain code. Build the baseline ladder first (random, a naive heuristic, the best code-only solution you can be bothered to write) and hold whichever model you pick to it. That test works the same whether the model behind the endpoint is Jev or Luna.

The Battleship harness is open source. If you're in the Decisions API preview, pointing it at OpenAI would give us the first head-to-head I've seen. I'd love to see the numbers.