# Jev vs Clef: which decision model for the hot path

Canonical: https://warike.tech/blog/jev-vs-clef-decision-models
By Sergio Cardenas · 2026-10-05
Tags: decision models, ai models, evaluation

<img src="/blog/jev-vs-clef-hero-960.webp" srcset="/blog/jev-vs-clef-hero-640.webp 640w, /blog/jev-vs-clef-hero-960.webp 960w, /blog/jev-vs-clef-hero.webp 1600w" sizes="(min-width: 1024px) 960px, calc(100vw - 48px)" width="960" height="541" loading="eager" fetchpriority="high" alt="Jev vs Clef, decision models in practice. The TypeSafe AI and Cloudflare logos either side of a record going through a decision into three outcomes.">

## Intro

[TypeSafe](https://typesafe.ai) opened Jev to early access on September 15.
Two weeks later, on October 1, Cloudflare shipped
[Clef](https://blog.cloudflare.com/clef-decision-models/), a 27B model, and
Clef-flash, a 9B one.

Two launches in two weeks, both calling themselves the same new thing: a
decision model, a model that returns probabilities over a fixed set of answers
and cannot write a sentence at all.

A good opportunity to build a use case, run all three through it, and keep
the answers, the bill and the stopwatch.

## TLDR

What to expect, from five returns and fifteen timed calls per model.

1. **All three agree on the action.** Fifteen decisions, every one landed the
   same across Jev, Clef and Clef-flash. The probabilities moved, the `pay` or
   `review` never did.
2. **Boolean and category are solid.** Fourteen of fifteen category picks were
   right. The boolean swung on one return where the wording was ambiguous.
3. **Score is the shakiest.** The 0 to 4 scale drifted 1.2 points between
   models on the same record. Fine behind a threshold, not for a ranking.
4. **Jev and Clef-flash tie on speed.** Both median 249 ms from a laptop. Jev
   has the tighter band, 198 to 321 ms. Clef is twice that at 506 ms and
   spreads to 939 ms.
5. **Input size barely matters.** From 33 words to 8,000, Jev and Clef-flash
   moved under 120 ms. Most of a call is the trip to the model and back.

## Context

The test hands each decision model one record and a fixed list of questions,
and reads back a probability for every allowed answer. Three question types
are on offer:

- **boolean**, a yes or no with a probability for yes
- **category**, a pick from a named list, with a probability for each name
- **score**, a position on an ordered scale, with a probability for each step

To exercise all three in one call the record has to carry a yes or no, a pick
from a list, and a scale. A warranty return does: does the fault match the
report, what caused it, how strong is the case to refuse. So the use case is a
warranty desk.

Five returns, written for the test, each one built to push on a different
corner.

## Objective

See how each answer type behaves on the same record across three models, and
what one call costs in money and in time from outside their network.

## Scope

### Requirements

- Three questions per record, one of each type.
- Three models, same questions, same policy.
- A stopwatch around every HTTP call.
- Standard library only, no SDK.

### Below the line

- A real operation. The desk exists to carry the three question types.
- Using these models as a router between different paths.

## Cost breakdown

Three question types on a 33-word record cost 569 input tokens on Jev and 493
on Clef. Output tokens are free on all three.

The following table shows the bill at the list price, at a thousand returns
a month and at a hundred thousand.

<img class="only-light" src="/blog/decision-model-cost-light.svg" width="760" height="150" loading="lazy" decoding="async" alt="Cost per month in USD. Jev 0.042 per million input tokens, 0.02 for 1,000 returns, 2.39 for 100,000. Clef-flash 0.09, 0.04, 4.44. Clef 0.24, 0.12, 11.83.">
<img class="only-dark" src="/blog/decision-model-cost-dark.svg" width="760" height="150" loading="lazy" decoding="async" alt="Cost per month in USD. Jev 0.042 per million input tokens, 0.02 for 1,000 returns, 2.39 for 100,000. Clef-flash 0.09, 0.04, 4.44. Clef 0.24, 0.12, 11.83.">

## Implementation

### Data flow

The following diagram shows the path of one record.

<img class="only-light" src="/blog/decision-model-flow-light.svg" width="936" height="225" loading="lazy" decoding="async" alt="One record goes to the model. The model returns a boolean, a category and a score. The policy turns them into pay or review.">
<img class="only-dark" src="/blog/decision-model-flow-dark.svg" width="936" height="225" loading="lazy" decoding="async" alt="One record goes to the model. The model returns a boolean, a category and a score. The policy turns them into pay or review.">

The request to the model carries the record and the three questions, not the
thresholds. The policy function takes the three answers, not the record.

### The scenarios

The following table lists the five returns, what each side said, and the
action a person would take.

<img class="only-light" src="/blog/decision-model-scenarios-light.svg" width="760" height="218" loading="lazy" decoding="async" alt="Five returns: dishwasher, washing machine, fridge, oven, laptop. What the customer said, what the technician found, and the expected action: pay for the dishwasher, review for the other four.">
<img class="only-dark" src="/blog/decision-model-scenarios-dark.svg" width="760" height="218" loading="lazy" decoding="async" alt="Five returns: dishwasher, washing machine, fridge, oven, laptop. What the customer said, what the technician found, and the expected action: pay for the dishwasher, review for the other four.">

### The questions

Three questions, one per type. Each has a short instruction and a criteria
block, which is the set of examples the model anchors on.

The following code is the boolean question as it goes on the wire; the
category and score questions have the same shape.

```go
{
    ID: "match", Kind: KindYesNo,
    Instructions: "Does the fault the technician found match the fault the customer reported?",
    Criteria: map[string]string{
        "true":  "customer says it stopped heating, technician finds a dead heating element",
        "false": "customer says it stopped heating, technician finds it heats fine but the door is broken",
    },
},
```

The category question lists `factory`, `misuse`, `shipping` and `none`. The
score question lists five levels, from "clearly pay" at 0 to "clearly refuse"
at 4.

### The policy

This is where the decision lives. Three thresholds, checked in order. Any miss
means a person reads the return.

The following diagram shows the three checks in order.

<img class="only-light" src="/blog/decision-model-policy-light.svg" width="1208" height="305" loading="lazy" decoding="async" alt="Three checks in a row. Factory at 0.70 or more, match at 0.70 or more, refuse score at 1.0 or less. Any miss goes to review, all three pass goes to pay.">
<img class="only-dark" src="/blog/decision-model-policy-dark.svg" width="1208" height="305" loading="lazy" decoding="async" alt="Three checks in a row. Factory at 0.70 or more, match at 0.70 or more, refuse score at 1.0 or less. Any miss goes to review, all three pass goes to pay.">

The following code is the three thresholds as they sit in the file.

```go
const (
    payMatchAt  = 0.70 // fault reported and fault found agree
    payCauseAt  = 0.70 // "factory" chosen with this much probability
    payRefuseAt = 1.0  // score at or under "probably pay"
)
```

If any answer is missing or has the wrong type, the policy returns `review`.
It fails closed. A bad response costs a read, never a payment.

### How to run

The code is public at
[github.com/warike/warranty-returns](https://github.com/warike/warranty-returns).
The measured numbers are in `data/results.json`, next to the script that draws
the tables and charts from them.

The following commands run the tests and then the five returns.

```sh
cp .env.example .env   # add the TypeSafe key and the Cloudflare account and token
go test -race ./...    # policy tests, no network
go run .               # five returns, three models, one table
```

## Results

### Answers

Five returns, three models, fifteen decisions. Two of the three questions came
back the same from every model, so they do not need a table:

- **Cause.** Same pick on fourteen of fifteen. The one miss is the fridge,
  where Clef-flash said `shipping` at 0.46 instead of `misuse`. The low number
  is the tell.
- **Refuse.** The 0 to 4 score stayed within half a point on three returns,
  drifted 0.8 on the fridge and 1.2 on the oven (2.7 to 3.9). Same side of the
  line every time, but not a number to rank by.

The following table shows where the models split: the probability of "yes,
the report matches the inspection" per model, with the action the policy took.

<img class="only-light" src="/blog/decision-model-answers-light.svg" width="760" height="218" loading="lazy" decoding="async" alt="Story matches inspection, Jev, Clef, Clef-flash, outcome. Dishwasher 0.97, 0.99, 0.99, pay. Washing machine 0.78, 0.50, 0.35, review. Fridge 0.30, 0.12, 0.11, review. Oven 0.83, 0.17, 0.08, review. Laptop 0.07, 0.17, 0.35, review.">
<img class="only-dark" src="/blog/decision-model-answers-dark.svg" width="760" height="218" loading="lazy" decoding="async" alt="Story matches inspection, Jev, Clef, Clef-flash, outcome. Dishwasher 0.97, 0.99, 0.99, pay. Washing machine 0.78, 0.50, 0.35, review. Fridge 0.30, 0.12, 0.11, review. Oven 0.83, 0.17, 0.08, review. Laptop 0.07, 0.17, 0.35, review.">

On a clear record all three agree. On the washing machine
and the oven, where the customer's story and the technician's note overlap but
do not match, Jev leans "yes" and both Clef models lean "no". The record is
ambiguous and the probability shows it. The threshold at 0.70 is what turns that
into one safe answer, `review`, instead of two different ones.

### Latency

The answers held up. Next, time per call. Fifteen calls per model on the
dishwasher record, from a laptop, stopwatch around the HTTP call only. One Jev
call came back with an HTTP 520 and was dropped.

The following chart shows the median per model with the min to max band.

<img class="only-light" src="/blog/decision-model-latency-light.svg" width="760" height="228" loading="lazy" decoding="async" alt="Horizontal bars. Jev median 249 ms, range 198 to 321. Clef-flash median 249 ms, range 121 to 412. Clef median 506 ms, range 291 to 939.">
<img class="only-dark" src="/blog/decision-model-latency-dark.svg" width="760" height="228" loading="lazy" decoding="async" alt="Horizontal bars. Jev median 249 ms, range 198 to 321. Clef-flash median 249 ms, range 121 to 412. Clef median 506 ms, range 291 to 939.">

Jev and Clef-flash land on the same median. Jev stays inside a 123 ms band.
Clef-flash has the fastest single call at 121 ms and the wider band. Clef is
twice as slow and spreads from 291 to 939 ms on the same input.

### Input size

Cloudflare's launch post lists Clef-flash at 39 ms, Clef at 209 ms and Jev at
524 ms, a different order from the chart above. One candidate reason: the test
records are tiny. Thirty-three words and three questions is not a workload,
so the dishwasher record was padded with service notes and run fifteen times
per size.

The following chart shows the median at three record sizes.

<img class="only-light" src="/blog/decision-model-scale-light.svg" width="760" height="290" loading="lazy" decoding="async" alt="Line chart, median latency against words in the record. Jev 249, 244, 365 ms. Clef-flash 249, 335, 275 ms. Clef 506, 755, 999 ms.">
<img class="only-dark" src="/blog/decision-model-scale-dark.svg" width="760" height="290" loading="lazy" decoding="async" alt="Line chart, median latency against words in the record. Jev 249, 244, 365 ms. Clef-flash 249, 335, 275 ms. Clef 506, 755, 999 ms.">

Two things in there.

**The small models barely move.** From 33 words to 8,000, Jev goes 249 to 365
ms and Clef-flash 249 to 275 ms. Reading the text costs a few milliseconds
either way. Most of the measured time is the fixed cost around the model: the
hop, the queue, the response.

**Clef's spread is not the input.** On the same 33 words, fifteen calls went
from 291 to 939 ms. The model did the same work fifteen times. The extra half
second is the request waiting somewhere between the laptop and the GPU.

The vendor table does not say how it was measured. The likely read: model time
on their hardware, without the trip there and back. That number matters when
choosing a GPU; the end-to-end number is the one a product feels.

Cloudflare also reports Clef ahead of Jev on a classification benchmark. Five
records cannot test that. On these five both Clef models gave lower
probabilities on every category pick, and the one wrong pick came from
Clef-flash.

## Conclusions

All three models read a messy record and returned the same fifteen actions.
Boolean and category answers are ready to use behind a threshold today. The
score answer drifted 1.2 points and needs a wider margin than the policy gave
it.

For a hot path the pick is Jev: same 249 ms as Clef-flash, the tighter band,
category picks at 0.98 or above, and the lowest price. Clef-flash is the one to
use if the rest of the stack already runs on Cloudflare. Clef did nothing the
small one did not, at 506 ms against 249.

The rule about money is three constants and one `switch`, in a file a finance
person can open and argue with. A threshold is defended in a code review; a
prompt is defended in an incident.

### What is next

The warranty desk is done. What stays is the call itself: one record in, a
category and a boolean out, no text to parse. That is a router.

The following diagram shows the router in front of three agents.

<img class="only-light" src="/blog/decision-model-router-light.svg" width="936" height="225" loading="lazy" decoding="async" alt="A user sends a request to a router. The router sends it to agent 1, 2 or 3. Each agent takes actions.">
<img class="only-dark" src="/blog/decision-model-router-dark.svg" width="936" height="225" loading="lazy" decoding="async" alt="A user sends a request to a router. The router sends it to agent 1, 2 or 3. Each agent takes actions.">

The router reads the request and picks the agent: a category question with a
probability per agent, which is what a decision model returns. That is the
next test, on harder input, since a router gets whatever the user typed,
including requests that fit two agents or none. Three things to measure: how
often the pick matches a person's, what the probability does when nothing
fits, and whether 249 ms holds on a paragraph of real text.

## Frequently asked questions

### What is a decision model?

A model that reads one record and a fixed list of questions, and returns a probability for every allowed answer. It writes no free text, so it cannot invent an answer outside the list.

### Why not let the model decide the refund?

A probability is an input. The threshold, the pay or review rule, and the fallback when an answer is missing live in code, where a person can read them, test them, and change them.
