← Insights

Jev vs Clef: which decision model for the hot path

Sergio Cardenas ·

Jev vs Clef, decision models in practice. The TypeSafe AI and Cloudflare logos either side of a record going through a decision into three outcomes.

Intro

TypeSafe opened Jev to early access on September 15. Two weeks later, on October 1, Cloudflare shipped Clef, a 27B model, and Clef-flash, a 9B one.

Two launches in two weeks, both calling themselves the same new thing: a decision model, a model that returns probabilities over a fixed set of answers and cannot write a sentence at all.

A good opportunity to build a use case, run all three through it, and keep the answers, the bill and the stopwatch.

TLDR

What to expect, from five returns and fifteen timed calls per model.

  1. All three agree on the action. Fifteen decisions, every one landed the same across Jev, Clef and Clef-flash. The probabilities moved, the pay or review never did.
  2. Boolean and category are solid. Fourteen of fifteen category picks were right. The boolean swung on one return where the wording was ambiguous.
  3. Score is the shakiest. The 0 to 4 scale drifted 1.2 points between models on the same record. Fine behind a threshold, not for a ranking.
  4. Jev and Clef-flash tie on speed. Both median 249 ms from a laptop. Jev has the tighter band, 198 to 321 ms. Clef is twice that at 506 ms and spreads to 939 ms.
  5. Input size barely matters. From 33 words to 8,000, Jev and Clef-flash moved under 120 ms. Most of a call is the trip to the model and back.

Context

The test hands each decision model one record and a fixed list of questions, and reads back a probability for every allowed answer. Three question types are on offer:

  • boolean, a yes or no with a probability for yes
  • category, a pick from a named list, with a probability for each name
  • score, a position on an ordered scale, with a probability for each step

To exercise all three in one call the record has to carry a yes or no, a pick from a list, and a scale. A warranty return does: does the fault match the report, what caused it, how strong is the case to refuse. So the use case is a warranty desk.

Five returns, written for the test, each one built to push on a different corner.

Objective

See how each answer type behaves on the same record across three models, and what one call costs in money and in time from outside their network.

Scope

Requirements

  • Three questions per record, one of each type.
  • Three models, same questions, same policy.
  • A stopwatch around every HTTP call.
  • Standard library only, no SDK.

Below the line

  • A real operation. The desk exists to carry the three question types.
  • Using these models as a router between different paths.

Cost breakdown

Three question types on a 33-word record cost 569 input tokens on Jev and 493 on Clef. Output tokens are free on all three.

The following table shows the bill at the list price, at a thousand returns a month and at a hundred thousand.

Cost per month in USD. Jev 0.042 per million input tokens, 0.02 for 1,000 returns, 2.39 for 100,000. Clef-flash 0.09, 0.04, 4.44. Clef 0.24, 0.12, 11.83. Cost per month in USD. Jev 0.042 per million input tokens, 0.02 for 1,000 returns, 2.39 for 100,000. Clef-flash 0.09, 0.04, 4.44. Clef 0.24, 0.12, 11.83.

Implementation

Data flow

The following diagram shows the path of one record.

One record goes to the model. The model returns a boolean, a category and a score. The policy turns them into pay or review. One record goes to the model. The model returns a boolean, a category and a score. The policy turns them into pay or review.

The request to the model carries the record and the three questions, not the thresholds. The policy function takes the three answers, not the record.

The scenarios

The following table lists the five returns, what each side said, and the action a person would take.

Five returns: dishwasher, washing machine, fridge, oven, laptop. What the customer said, what the technician found, and the expected action: pay for the dishwasher, review for the other four. Five returns: dishwasher, washing machine, fridge, oven, laptop. What the customer said, what the technician found, and the expected action: pay for the dishwasher, review for the other four.

The questions

Three questions, one per type. Each has a short instruction and a criteria block, which is the set of examples the model anchors on.

The following code is the boolean question as it goes on the wire; the category and score questions have the same shape.

{
    ID: "match", Kind: KindYesNo,
    Instructions: "Does the fault the technician found match the fault the customer reported?",
    Criteria: map[string]string{
        "true":  "customer says it stopped heating, technician finds a dead heating element",
        "false": "customer says it stopped heating, technician finds it heats fine but the door is broken",
    },
},

The category question lists factory, misuse, shipping and none. The score question lists five levels, from "clearly pay" at 0 to "clearly refuse" at 4.

The policy

This is where the decision lives. Three thresholds, checked in order. Any miss means a person reads the return.

The following diagram shows the three checks in order.

Three checks in a row. Factory at 0.70 or more, match at 0.70 or more, refuse score at 1.0 or less. Any miss goes to review, all three pass goes to pay. Three checks in a row. Factory at 0.70 or more, match at 0.70 or more, refuse score at 1.0 or less. Any miss goes to review, all three pass goes to pay.

The following code is the three thresholds as they sit in the file.

const (
    payMatchAt  = 0.70 // fault reported and fault found agree
    payCauseAt  = 0.70 // "factory" chosen with this much probability
    payRefuseAt = 1.0  // score at or under "probably pay"
)

If any answer is missing or has the wrong type, the policy returns review. It fails closed. A bad response costs a read, never a payment.

How to run

The code is public at github.com/warike/warranty-returns. The measured numbers are in data/results.json, next to the script that draws the tables and charts from them.

The following commands run the tests and then the five returns.

cp .env.example .env   # add the TypeSafe key and the Cloudflare account and token
go test -race ./...    # policy tests, no network
go run .               # five returns, three models, one table

Results

Answers

Five returns, three models, fifteen decisions. Two of the three questions came back the same from every model, so they do not need a table:

  • Cause. Same pick on fourteen of fifteen. The one miss is the fridge, where Clef-flash said shipping at 0.46 instead of misuse. The low number is the tell.
  • Refuse. The 0 to 4 score stayed within half a point on three returns, drifted 0.8 on the fridge and 1.2 on the oven (2.7 to 3.9). Same side of the line every time, but not a number to rank by.

The following table shows where the models split: the probability of "yes, the report matches the inspection" per model, with the action the policy took.

Story matches inspection, Jev, Clef, Clef-flash, outcome. Dishwasher 0.97, 0.99, 0.99, pay. Washing machine 0.78, 0.50, 0.35, review. Fridge 0.30, 0.12, 0.11, review. Oven 0.83, 0.17, 0.08, review. Laptop 0.07, 0.17, 0.35, review. Story matches inspection, Jev, Clef, Clef-flash, outcome. Dishwasher 0.97, 0.99, 0.99, pay. Washing machine 0.78, 0.50, 0.35, review. Fridge 0.30, 0.12, 0.11, review. Oven 0.83, 0.17, 0.08, review. Laptop 0.07, 0.17, 0.35, review.

On a clear record all three agree. On the washing machine and the oven, where the customer's story and the technician's note overlap but do not match, Jev leans "yes" and both Clef models lean "no". The record is ambiguous and the probability shows it. The threshold at 0.70 is what turns that into one safe answer, review, instead of two different ones.

Latency

The answers held up. Next, time per call. Fifteen calls per model on the dishwasher record, from a laptop, stopwatch around the HTTP call only. One Jev call came back with an HTTP 520 and was dropped.

The following chart shows the median per model with the min to max band.

Horizontal bars. Jev median 249 ms, range 198 to 321. Clef-flash median 249 ms, range 121 to 412. Clef median 506 ms, range 291 to 939. Horizontal bars. Jev median 249 ms, range 198 to 321. Clef-flash median 249 ms, range 121 to 412. Clef median 506 ms, range 291 to 939.

Jev and Clef-flash land on the same median. Jev stays inside a 123 ms band. Clef-flash has the fastest single call at 121 ms and the wider band. Clef is twice as slow and spreads from 291 to 939 ms on the same input.

Input size

Cloudflare's launch post lists Clef-flash at 39 ms, Clef at 209 ms and Jev at 524 ms, a different order from the chart above. One candidate reason: the test records are tiny. Thirty-three words and three questions is not a workload, so the dishwasher record was padded with service notes and run fifteen times per size.

The following chart shows the median at three record sizes.

Line chart, median latency against words in the record. Jev 249, 244, 365 ms. Clef-flash 249, 335, 275 ms. Clef 506, 755, 999 ms. Line chart, median latency against words in the record. Jev 249, 244, 365 ms. Clef-flash 249, 335, 275 ms. Clef 506, 755, 999 ms.

Two things in there.

The small models barely move. From 33 words to 8,000, Jev goes 249 to 365 ms and Clef-flash 249 to 275 ms. Reading the text costs a few milliseconds either way. Most of the measured time is the fixed cost around the model: the hop, the queue, the response.

Clef's spread is not the input. On the same 33 words, fifteen calls went from 291 to 939 ms. The model did the same work fifteen times. The extra half second is the request waiting somewhere between the laptop and the GPU.

The vendor table does not say how it was measured. The likely read: model time on their hardware, without the trip there and back. That number matters when choosing a GPU; the end-to-end number is the one a product feels.

Cloudflare also reports Clef ahead of Jev on a classification benchmark. Five records cannot test that. On these five both Clef models gave lower probabilities on every category pick, and the one wrong pick came from Clef-flash.

Conclusions

All three models read a messy record and returned the same fifteen actions. Boolean and category answers are ready to use behind a threshold today. The score answer drifted 1.2 points and needs a wider margin than the policy gave it.

For a hot path the pick is Jev: same 249 ms as Clef-flash, the tighter band, category picks at 0.98 or above, and the lowest price. Clef-flash is the one to use if the rest of the stack already runs on Cloudflare. Clef did nothing the small one did not, at 506 ms against 249.

The rule about money is three constants and one switch, in a file a finance person can open and argue with. A threshold is defended in a code review; a prompt is defended in an incident.

What is next

The warranty desk is done. What stays is the call itself: one record in, a category and a boolean out, no text to parse. That is a router.

The following diagram shows the router in front of three agents.

A user sends a request to a router. The router sends it to agent 1, 2 or 3. Each agent takes actions. A user sends a request to a router. The router sends it to agent 1, 2 or 3. Each agent takes actions.

The router reads the request and picks the agent: a category question with a probability per agent, which is what a decision model returns. That is the next test, on harder input, since a router gets whatever the user typed, including requests that fit two agents or none. Three things to measure: how often the pick matches a person's, what the probability does when nothing fits, and whether 249 ms holds on a paragraph of real text.

Frequently asked questions

What is a decision model?
A model that reads one record and a fixed list of questions, and returns a probability for every allowed answer. It writes no free text, so it cannot invent an answer outside the list.
Why not let the model decide the refund?
A probability is an input. The threshold, the pay or review rule, and the fallback when an answer is missing live in code, where a person can read them, test them, and change them.