Jev vs Clef: which decision model for the hot path
Sergio Cardenas ·

Intro
TypeSafe opened Jev to early access on September 15. Two weeks later, on October 1, Cloudflare shipped Clef, a 27B model, and Clef-flash, a 9B one.
Two launches in two weeks, both calling themselves the same new thing: a decision model, a model that returns probabilities over a fixed set of answers and cannot write a sentence at all.
A good opportunity to build a use case, run all three through it, and keep the answers, the bill and the stopwatch.
TLDR
What to expect, from five returns and fifteen timed calls per model.
- All three agree on the action. Fifteen decisions, every one landed the
same across Jev, Clef and Clef-flash. The probabilities moved, the
payorreviewnever did. - Boolean and category are solid. Fourteen of fifteen category picks were right. The boolean swung on one return where the wording was ambiguous.
- Score is the shakiest. The 0 to 4 scale drifted 1.2 points between models on the same record. Fine behind a threshold, not for a ranking.
- Jev and Clef-flash tie on speed. Both median 249 ms from a laptop. Jev has the tighter band, 198 to 321 ms. Clef is twice that at 506 ms and spreads to 939 ms.
- Input size barely matters. From 33 words to 8,000, Jev and Clef-flash moved under 120 ms. Most of a call is the trip to the model and back.
Context
The test hands each decision model one record and a fixed list of questions, and reads back a probability for every allowed answer. Three question types are on offer:
- boolean, a yes or no with a probability for yes
- category, a pick from a named list, with a probability for each name
- score, a position on an ordered scale, with a probability for each step
To exercise all three in one call the record has to carry a yes or no, a pick from a list, and a scale. A warranty return does: does the fault match the report, what caused it, how strong is the case to refuse. So the use case is a warranty desk.
Five returns, written for the test, each one built to push on a different corner.
Objective
See how each answer type behaves on the same record across three models, and what one call costs in money and in time from outside their network.
Scope
Requirements
- Three questions per record, one of each type.
- Three models, same questions, same policy.
- A stopwatch around every HTTP call.
- Standard library only, no SDK.
Below the line
- A real operation. The desk exists to carry the three question types.
- Using these models as a router between different paths.
Cost breakdown
Three question types on a 33-word record cost 569 input tokens on Jev and 493 on Clef. Output tokens are free on all three.
The following table shows the bill at the list price, at a thousand returns a month and at a hundred thousand.
Implementation
Data flow
The following diagram shows the path of one record.
The request to the model carries the record and the three questions, not the thresholds. The policy function takes the three answers, not the record.
The scenarios
The following table lists the five returns, what each side said, and the action a person would take.
The questions
Three questions, one per type. Each has a short instruction and a criteria block, which is the set of examples the model anchors on.
The following code is the boolean question as it goes on the wire; the category and score questions have the same shape.
{
ID: "match", Kind: KindYesNo,
Instructions: "Does the fault the technician found match the fault the customer reported?",
Criteria: map[string]string{
"true": "customer says it stopped heating, technician finds a dead heating element",
"false": "customer says it stopped heating, technician finds it heats fine but the door is broken",
},
},
The category question lists factory, misuse, shipping and none. The
score question lists five levels, from "clearly pay" at 0 to "clearly refuse"
at 4.
The policy
This is where the decision lives. Three thresholds, checked in order. Any miss means a person reads the return.
The following diagram shows the three checks in order.
The following code is the three thresholds as they sit in the file.
const (
payMatchAt = 0.70 // fault reported and fault found agree
payCauseAt = 0.70 // "factory" chosen with this much probability
payRefuseAt = 1.0 // score at or under "probably pay"
)
If any answer is missing or has the wrong type, the policy returns review.
It fails closed. A bad response costs a read, never a payment.
How to run
The code is public at
github.com/warike/warranty-returns.
The measured numbers are in data/results.json, next to the script that draws
the tables and charts from them.
The following commands run the tests and then the five returns.
cp .env.example .env # add the TypeSafe key and the Cloudflare account and token
go test -race ./... # policy tests, no network
go run . # five returns, three models, one table
Results
Answers
Five returns, three models, fifteen decisions. Two of the three questions came back the same from every model, so they do not need a table:
- Cause. Same pick on fourteen of fifteen. The one miss is the fridge,
where Clef-flash said
shippingat 0.46 instead ofmisuse. The low number is the tell. - Refuse. The 0 to 4 score stayed within half a point on three returns, drifted 0.8 on the fridge and 1.2 on the oven (2.7 to 3.9). Same side of the line every time, but not a number to rank by.
The following table shows where the models split: the probability of "yes, the report matches the inspection" per model, with the action the policy took.
On a clear record all three agree. On the washing machine
and the oven, where the customer's story and the technician's note overlap but
do not match, Jev leans "yes" and both Clef models lean "no". The record is
ambiguous and the probability shows it. The threshold at 0.70 is what turns that
into one safe answer, review, instead of two different ones.
Latency
The answers held up. Next, time per call. Fifteen calls per model on the dishwasher record, from a laptop, stopwatch around the HTTP call only. One Jev call came back with an HTTP 520 and was dropped.
The following chart shows the median per model with the min to max band.
Jev and Clef-flash land on the same median. Jev stays inside a 123 ms band. Clef-flash has the fastest single call at 121 ms and the wider band. Clef is twice as slow and spreads from 291 to 939 ms on the same input.
Input size
Cloudflare's launch post lists Clef-flash at 39 ms, Clef at 209 ms and Jev at 524 ms, a different order from the chart above. One candidate reason: the test records are tiny. Thirty-three words and three questions is not a workload, so the dishwasher record was padded with service notes and run fifteen times per size.
The following chart shows the median at three record sizes.
Two things in there.
The small models barely move. From 33 words to 8,000, Jev goes 249 to 365 ms and Clef-flash 249 to 275 ms. Reading the text costs a few milliseconds either way. Most of the measured time is the fixed cost around the model: the hop, the queue, the response.
Clef's spread is not the input. On the same 33 words, fifteen calls went from 291 to 939 ms. The model did the same work fifteen times. The extra half second is the request waiting somewhere between the laptop and the GPU.
The vendor table does not say how it was measured. The likely read: model time on their hardware, without the trip there and back. That number matters when choosing a GPU; the end-to-end number is the one a product feels.
Cloudflare also reports Clef ahead of Jev on a classification benchmark. Five records cannot test that. On these five both Clef models gave lower probabilities on every category pick, and the one wrong pick came from Clef-flash.
Conclusions
All three models read a messy record and returned the same fifteen actions. Boolean and category answers are ready to use behind a threshold today. The score answer drifted 1.2 points and needs a wider margin than the policy gave it.
For a hot path the pick is Jev: same 249 ms as Clef-flash, the tighter band, category picks at 0.98 or above, and the lowest price. Clef-flash is the one to use if the rest of the stack already runs on Cloudflare. Clef did nothing the small one did not, at 506 ms against 249.
The rule about money is three constants and one switch, in a file a finance
person can open and argue with. A threshold is defended in a code review; a
prompt is defended in an incident.
What is next
The warranty desk is done. What stays is the call itself: one record in, a category and a boolean out, no text to parse. That is a router.
The following diagram shows the router in front of three agents.
The router reads the request and picks the agent: a category question with a probability per agent, which is what a decision model returns. That is the next test, on harder input, since a router gets whatever the user typed, including requests that fit two agents or none. Three things to measure: how often the pick matches a person's, what the probability does when nothing fits, and whether 249 ms holds on a paragraph of real text.
Frequently asked questions
- What is a decision model?
- A model that reads one record and a fixed list of questions, and returns a probability for every allowed answer. It writes no free text, so it cannot invent an answer outside the list.
- Why not let the model decide the refund?
- A probability is an input. The threshold, the pay or review rule, and the fallback when an answer is missing live in code, where a person can read them, test them, and change them.