← Back to writing

Benchmarking Cloudflare's Clef against Jev at Battleship: two new models, the same tie

October 6, 2026 Versão em português

Benchmarking Cloudflare's Clef against Jev at Battleship: two new models, the same tie

On October 1, Cloudflare released Clef, two decision models from the Workers AI team. Clef has 27 billion parameters and Clef-flash has 9. Both follow the same "System One" API as TypeSafe's Jev: you send a state and some typed questions, and you get a probability back for every allowed answer. Cloudflare says moving from Jev is a direct migration.

A week earlier I'd written that a small dedicated model should beat a frontier model behind an endpoint at the hundreds of tiny judgements agents need. Clef is a second dedicated model, with open weights, image support and some very fast latency claims. Same API meant my Battleship benchmark could talk to it with almost no changes, so I ran it.

Sixty games per player, four new runs, $7.18 all in.

Keeping the comparison honest

Every player used the same 60 fleet layouts as the September run, drawn from the mixed family so no baseline gets to play on home ground. I checked the pairing instead of assuming it: the fleets are byte-identical, layout by layout, so every comparison below is paired, including the ones against the old Jev and baseline numbers.

All three models went through the same endpoint, OpenRouter's Decisions API. Reaching Jev one way and Clef through Workers AI would have measured two networks and two regions along with two models, with no way to tell afterwards which one made the difference. One endpoint leaves the model as the only thing that changes.

The players are the same two as before. The pure player gets the raw board and every untried cell as an option, about 90 of them mid-game. The hybrid player gets a shortlist of 16 cells that a density solver has already ranked and described in plain words.

The scoreboard

Lower is better. Seventeen shots is a perfect game, 100 is the worst you can do.

PlayerMean shotsCost
Jev hybrid46.0$0.12
Density solver (code)48.3
Clef hybrid48.6$0.60
Clef-flash hybrid49.9$0.23
Hunt / Target (code)51.9
Clef pure76.2$2.95
Jev pure85.5$0.65
Clef-flash pure87.0$1.21
Random95.3

With a shortlist, everything ties

Once the code has ranked the top 16 and described them, all three models land between 46.0 and 49.9 shots, and the density solver sits at 48.3 right in the middle of them. No model beats the solver. No model beats another model.

Clef against the solver is the closest result in the whole project: 0.3 shots apart over 60 boards, p = 0.85. Settling it would take about 12,345 games. The other pairs would need somewhere between 210 and 1,047.

I ran 36 comparisons, so the corrected threshold for calling anything significant is 0.0014. Jev beating Clef-flash at the hybrid task came in at p = 0.030, which clears the usual 5% bar and nothing else. I'm treating it as a tie.

So the conclusion from September survives two more models and gets a bit stronger. The setup where a model looks good is the setup where code has already done the work, and in that setup the model adds nothing I can measure.

On a raw board, Clef is better

This is the one place where any model separates from any other. Clef's pure player needs 76.2 shots against Jev's 85.5, a gap of 9.3 shots with p = 0.0002, comfortably inside the corrected threshold. It also beats Clef-flash by 10.8.

The reason was predictable before a single game ran. I wrote a small probe that sends a 90-option board and looks at what comes back.

ModelDecimalsNon-zero of 90Top option
Jev244.90.28
Clef490.00.05
Clef-flash490.00.05

Jev returns two decimal places. Across 90 options, about 45 come back as exactly zero and another 33 tie at 0.01, which leaves a usable ranking over roughly a dozen cells. Clef returns four decimals and ranks all ninety. The pure player is exactly the task where that matters. With only 16 options in the hybrid player, Jev's dozen is plenty and the advantage disappears.

And it doesn't help much. At 76.2 shots, Clef's pure player is still 28 shots behind the density solver and 24 behind the hunt-and-target heuristic. Being the best of three models at a task where all three lose badly to a hundred lines of plain code isn't something I'd build on.

A correction to my own article

That probe also caught a mistake in my September write-up. I said only six cells came back with any value and that most of the distribution rounded to zero. Both were wrong. Jev keeps 0.996 of the probability, and about 45 of 90 options are non-zero, measured at 6, 8 and 12 board positions with a stable answer each time.

What the rounding destroys is the order of the tail. My heatmap looked empty because 0.01 barely shows on the colour ramp, not because the values were zero. The conclusion that Jev's pure player has very little usable signal still stands, but the number and the mechanism I gave for it didn't, so I've corrected the original article with a dated note.

How much each model listens to the code

I replayed every hybrid game, recomputed the solver's ranking at each shot, and checked where the model's pick landed in it.

Hybrid playerMean rank picked (of 16)Picks the top cellMean shots
Jev4.6826.9%46.0
Clef5.9227.1%48.6
Clef-flash6.5420.0%49.9
Chance7.506.25%

The columns line up. The further a model drifts from the code's ranking, the worse it plays. All three are well clear of chance, so they're all reading the shortlist. They just differ in how much they trust it.

That also explains why Clef is so uneven. It has the widest spread of any model player and three games above 75 shots, where Jev has none. On the worst one, layout 1031, Clef took 99 shots and Jev took 44. Clef's average pick was 8.41 out of 16, it chose the solver's worst candidate more often than its best, and 58 of its 99 shots missed while a real ship cell was sitting in the shortlist. The code was doing its job and the model kept throwing it away.

Clef-flash is the opposite: the smallest spread of any model player, no disasters, and never particularly good.

My latency numbers came out upside down

Cloudflare's announcement quotes median latencies of 38.8 ms for Clef-flash, 209.3 ms for Clef and 524.1 ms for Jev. I measured 40 interleaved calls per model from the same machine, through the same endpoint, end to end including the network.

Model16 options, median90 options, median
Jev300 ms277 ms
Clef-flash350 ms425 ms
Clef524 ms713 ms

Jev was the fastest at both sizes, and its median barely moved as the input grew from about 760 to 2,300 tokens. Clef's went up by a third.

I don't think either set of numbers is wrong. 38.8 ms is barely a round trip to anywhere, so Cloudflare's figures almost certainly leave the network out. Mine describe the path through OpenRouter, which makes them comparable to each other and to nothing else. If you're choosing between these models for latency, measure from where your code actually runs.

Cost

Output is free on both Clef models, so cost comes down to input. Clef charges $0.240 per million input tokens, 5.7 times Jev's $0.042. The four runs in the table cost $4.99, and discarded runs, probes and smoke tests added another $2.19. The Clef pure run cost $2.95 to finish 28 shots behind a solver that costs nothing to run.

What Battleship can't tell you about Clef

There are things Clef does that this benchmark never touches.

It takes images, up to four per request, and Jev only takes text. If your decision depends on a screenshot, a scanned document or a photo of a damaged parcel, Clef can look at it and Jev can't. Of the decision APIs I'd looked at before Clef, only OpenAI's took images, and that one is still in limited preview.

The weights are open, under Apache 2.0 on Hugging Face. You can run Clef on your own hardware, which matters if the data you're deciding on can't leave your network. Self-hosting also means a silent model update can't change your results under you. Through the hosted API, though, Clef reports no version string, so I can't pin these numbers to a build the way I can pin Jev's to jev-1.13-20260917.

None of that shows up in a game of Battleship. It's a text-only benchmark on one game and one layout family, and it says nothing about how Clef does on the work Cloudflare built it for, like support triage or trust and safety scoring.

What I'd take from this

Two new models went into the harness and the main result didn't move. When code has built a good shortlist, the model on top of it ties with the code, whichever model it is. When the model has to do the hard part alone, the best of the three is still far behind a few lines of plain code.

The thing worth copying is the setup more than any of the numbers. One endpoint for every model, so the network isn't a hidden variable. Paired layouts, checked rather than assumed. A corrected significance threshold when you run dozens of comparisons. And a probe that tests the mechanism before you run the games, which in my case predicted the one real result and caught my own mistake from September.

The full results, including everything these numbers don't establish, are in results-clef.md. The code is at github.com/ickas/battleship-vs-jev, and the commands to reproduce every table above are at the end of that file.