Skip to content

Experiment report · Decision models

Jev calibration, tested on 1,232 banking support messages

On 1,232 support messages, Jev's 1.0 answers were 97.5% right and its 0.6 to 0.7 answers 44%. One temperature fit on 50 labels cut the error by a third.

  • Run 30 Sep 2026
  • jev-1.13-20260917
  • 1,232 messages · 77 intents
  • $0.094 in API calls

On our test, 41% of Jev’s answers came with a confidence of 1.0, and those answers were right 97.5% of the time. Answers it scored between 0.6 and 0.7 were right 44% of the time. Both numbers come from the same model, the same 1,232 banking support messages and the same run on 30 September 2026. Together they are the most useful thing I can tell you about Jev calibration: the top of the scale held up, and the middle did not.

Calibration is the match between a model’s stated probabilities and how often it turns out right. The scikit-learn documentation puts it concretely: a calibrated classifier is one where, “among the samples to which it gave a predict_proba value close to, say, 0.8, approximately 80% actually belong to the positive class.”[1] Jev returns “the selected option, a probability for each option, and a confidence value” for every choice question,[2] so whether those numbers mean what they say decides whether you can route work on them.

I am a data engineer at Verne, and routing is where this stops being a debate, so we measured it.

Is Jev calibrated? It depends on the task

TypeSafe says that “All answers are accompanied with calibrated probabilities and confidence scores.”[3] The same launch post names its yardstick: “use the predictions of the largest, smartest, and most expensive external models as reference probabilities.”[3] That reference is agreement with bigger models, not ground truth, and the page shows no calibration metrics. The speed claims are vendor numbers too: TypeSafe claims Jev is 193.6x faster and 444.6x cheaper, and MarkTechPost notes that those figures “come from TypeSafe’s own workflow evals.”[4]

Independent tests disagree, and the disagreement is useful. A fair-die stress test, posted as a GitHub issue, reports: “The main fair-die run contains 400 trials. Choice selected face 1 on every trial with 82.9% mean reported probability and 19.0% observed accuracy.”[5] It is a stress test, not a benchmark of Jev overall.

An audit by Rafe and Das, on Texas police crash narratives scored against human labels, found the opposite problem. “The calibration slope is 1.63, above one, so the probabilities are not extreme enough,” they write. Out of fold, Platt scaling or isotonic regression cut their pooled calibration error from 0.0231 to 0.0069, “a factor of 3.3”.[6] Their own caution applies to every study here, ours included: “Transfer is a hypothesis rather than a finding, because this study covers one state, one model version and one schema”.[6]

Alex Molas summed up the pattern in one line: “The same model can be calibrated on one dataset but not on another.” His practical advice is to “treat Jev’s outputs as good scores (they rank examples well) rather than good probabilities.”[7] I think he is right, and our numbers show why.

Three tests of Jev’s confidence, side by side
TestTask and ground truthWhat it found
Fair-die stress test[5]400 rolls of a fair die; the true odds are 1 in 682.9% mean stated probability, 19.0% observed accuracy
Rafe and Das audit[6]Texas police crash narratives; human reference labelsUnderconfident (slope 1.63); error cut by a factor of 3.3 out of fold
Verne test (this post)1,232 Banking77 support messages, 77 intents; the dataset’s labelsOverconfident in the middle range; median held-out ECE over 50 splits 0.092, then 0.058 after one temperature

Accuracy is a separate question from calibration. On text annotation, Ibrahim and Zaki evaluated what they call the first commercial decision model and found that “The decision model trails the per-task best LLM on 14 of 15 evaluation tasks, with a median deficit of 11.6 macro-F1 points, at a median 44 times lower measured cost.”[8] Products already build on it; LangWatch writes that “Instant Evals runs on Jev, so we wanted to know how far open models are from it before recommending either.”[9]

Self-reported confidence has a history. Kadavath and colleagues studied “asking models to first propose answers, and then to evaluate the probability ‘P(True)’ that their answers are correct.”[10] Xiong and colleagues found that “LLMs, when verbalizing their confidence, tend to be overconfident, potentially imitating human patterns of expressing confidence.”[11] Jev returns a number per option rather than a sentence, so neither result settles it. Jev calibration needs its own test.

What we measured and how

For our Jev calibration test we used Banking77, a public dataset of online banking queries labelled with 77 intents, released by PolyAI in 2020 under CC BY 4.0.[12] With 77 options and several near-duplicate intents, it is a hard routing test. The run had six steps:

  1. Sample. Sixteen messages per intent from the test split, drawn with seed 7: 1,232 messages in all.
  2. Ask. One choice question per message to typesafe/jev-1.13 through OpenRouter’s Decisions API (alpha), the 77 intent names as options. All calls ran on 30 September 2026, served by jev-1.13-20260917.
  3. Record. The chosen intent, the probability for every option and the confidence value. We used the probability on the chosen intent as Jev’s stated confidence; the separate confidence value never differed from it by more than 0.02.
  4. Score. Accuracy against the dataset’s label, and top-label expected calibration error over 15 equal-width bins.
  5. Recalibrate. Fit one temperature on half the messages, score the other half, and repeat over 50 random splits.
  6. Repeat. Re-send the first 100 messages to check stability.

Expected calibration error (ECE) is the average gap between stated confidence and actual accuracy, weighted by how many answers fall in each confidence bin. Zero is perfect.

The whole run cost $0.094 for 1,332 calls, about $0.07 per 1,000 decisions. TypeSafe’s models page says Jev is “Charged per input token. Output tokens are free.”[13] That was the pricing on 1 October 2026. Our calls averaged 1,689 input tokens each.

How the test ran · one real message

Banking77 test split, 16 messages for each of 77 intents, drawn with seed 7.

text: "I transferred some money but the receiver can't pick it up for some reason."
label: transfer_not_received_by_recipient

Stage 1 of 6: Sample, 1,232 messages. Banking77 test split, 16 messages for each of 77 intents, drawn with seed 7. text: "I transferred some money but the receiver can't pick it up for some reason.". label: transfer_not_received_by_recipient

Figure 1. The run, stage by stage, following one real message from the test set. Every value comes from Verne’s run on 30 September 2026. Select a stage, use Next, or play the run (the play button is hidden when your system asks for reduced motion).

Here is the core of the script we ran, without the retries, parallel workers and spend guard.

Pythonjev_calib.py, simplified: ask Jev one choice question and score calibration
import json
import os
import urllib.request

import numpy as np

URL = "https://openrouter.ai/api/alpha/decisions"  # OpenRouter Decisions API (alpha)
KEY = os.environ["OPENROUTER_API_KEY"]


def ask_jev(message: str, intents: list[str]) -> dict:
    """One choice question: which of the 77 intents does this message express?"""
    body = {
        "model": "typesafe/jev-1.13",
        "state": message,
        "questions": {
            "intent": {
                "type": "choice",
                "instructions": "Which banking support intent does this customer message express?",
                # Option text is the intent name with underscores replaced by spaces.
                "criteria": {i: i.replace("_", " ") for i in intents},
            }
        },
    }
    req = urllib.request.Request(
        URL,
        data=json.dumps(body).encode(),
        headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
    )
    with urllib.request.urlopen(req, timeout=90) as r:
        # {"choice": ..., "probabilities": {intent: p, ...}, "confidence": ...}
        return json.load(r)["answers"]["intent"]


def ece(conf: np.ndarray, correct: np.ndarray, bins: int = 15) -> float:
    """Top-label expected calibration error over equal-width confidence bins."""
    edges = np.linspace(0, 1, bins + 1)
    idx = np.clip(np.digitize(conf, edges[1:-1]), 0, bins - 1)
    total = 0.0
    for b in range(bins):
        in_bin = idx == b
        if in_bin.any():
            # Weight each bin's |accuracy - stated confidence| gap by its share of answers.
            total += in_bin.mean() * abs(correct[in_bin].mean() - conf[in_bin].mean())
    return total


# conf = probability on the chosen intent, correct = choice matched the label
# print(ece(conf, correct))  # 0.086 on our 1,232 messages

Where Jev’s confidence holds, and where it breaks

Jev’s top choice matched the dataset’s intent on 979 of 1,232 messages, an accuracy of 79.5%. Its mean stated confidence was 88.1%, so on average it overstated its hit rate by about 8.6 points, and ECE across the full set was 0.086 (the same with equal-mass bins). The averages hide the shape of Jev calibration, which is what matters for routing.

Reliability diagram · real dataVerne's measurement, 30 Sep 2026 · Jev (jev-1.13-20260917) on 1,232 Banking77 test messages, 77 intents
Confidence shown
Accuracy
79.5%
979 of 1,232; scaling never changes the top choice
Mean stated confidence
88.1%
8.6 points above accuracy
Held-out ECE (median, 50 splits)
0.092
15 bins; 0 means perfectly calibrated
0%25%50%75%100%Share rightn810.00.20.40.60.81.0Jev's stated confidence
  • share right
  • mean stated confidence
  • gap
  • perfect calibration
  • messages in band (n)

Band 0.6 to 0.7 · 81 messages

Jev said 0.65 on average; 44% were right.

21 points overconfident.

Show the numbers for every band
Reliability by confidence band, raw and recalibrated
BandJev’s own numbersRecalibrated
nStatedRightnStatedRight
0.0 to 0.10··0··
0.1 to 0.20··10.170%
0.2 to 0.310.240%50.2740%
0.3 to 0.4150.3633%410.3534%
0.4 to 0.5440.4536%870.4639%
0.5 to 0.6810.5440%1010.5549%
0.6 to 0.7810.6544%920.6466%
0.7 to 0.8800.7568%1180.7569%
0.8 to 0.91030.8570%1970.8585%
0.9 to 1.08270.9892%5900.9297%
Figure 2. Verne’s measurement, 30 September 2026, method as above. Bars show how often Jev was right in each confidence band, the pale tick is the confidence it stated, and the hatched block is the gap. Hover, tap or use the arrow keys on a band for its numbers. The recalibrated view uses one temperature fitted on all 1,232 messages, so it is in-sample; held-out figures are in the next section.

Above 0.9, where 827 of the 1,232 answers sat, Jev stated 0.98 on average and was right 92% of the time. The 509 answers at a flat 1.0 did best: 496 of them were right, 97.5%. Between 0.5 and 0.9 the stated number ran 8 to 21 points ahead of the hit rate, and answers stated between 0.6 and 0.7 (mean 0.65) were right 44% of the time.

Stated confidence against actual accuracy, by band (Verne’s test, 30 Sep 2026)
Jev’s confidenceMessagesMean statedRightGap (points)
Below 0.5600.4235%7
0.5 to 0.6810.5440%15
0.6 to 0.7810.6544%21
0.7 to 0.8800.7568%8
0.8 to 0.91030.8570%16
0.9 to 1.08270.9892%6
Exactly 1.0 (part of 0.9 to 1.0)5091.0097.5%3

Zeros that no threshold can rescue

Jev’s probabilities are coarse. On our run, 96.9% of all per-option probabilities were exactly 0, and only 101 distinct values appeared: a two-decimal grid. In 76 messages (6.2%) the correct intent got a probability of exactly 0. No fallback to Jev’s second choice can recover those, and recalibration cannot either: every zero gets the same floor value, so there is no ranking left.

The same message, twice

Jev was not deterministic in our run. Re-sending the first 100 messages returned the same top choice 97 times and identical probabilities 55 times. For an audit trail, store the probabilities you acted on and the model version; you cannot count on regenerating them.

Misses cluster where intents overlap. The most frequent confusions were get_physical_card read as change_pin (9 times), order_physical_card read as get_physical_card (9), and top_up_by_bank_transfer_charge read as transfer_fee_charged or transfer_into_account (6 each). Some confident misses look more like label disputes than model errors; the last three messages in the explorer below are worth reading closely.

Message explorer · 12 real messagesVerne's measurement, 30 Sep 2026 · Jev (jev-1.13-20260917) on 1,232 Banking77 test messages, 77 intents
Show
Probabilities
Pick a message

Jev chose transfer not received by recipient at 0.74; the dataset label is transfer not received by recipient, so it was right.

“I transferred some money but the receiver can't pick it up for some reason.”

Dataset label transfer not received by recipientJev chose transfer not received by recipient right

  • transfer not received by recipientchosenlabel0.74
  • receiving money0.25
  • pending transfer0.01

74 other intents: exactly 0.

The 242 answers between 0.5 and 0.8 were right 122 times, about half. The runner-up often holds most of the rest of the probability, which makes these good candidates for a person with two options on screen.

Figure 3. Twelve real messages from the run, drawn at random from four groups: sure and right (1.0), sure and wrong (0.9 or more), the middle range (0.5 to 0.8) and messages where the dataset’s intent got 0. Switch to recalibrated to see what one temperature does to each message’s probabilities.

How many labels does it take to fix Jev calibration?

Temperature scaling divides a model’s log-probabilities by a single number, T, and renormalises. A T above 1 softens overconfident probabilities; below 1 it sharpens underconfident ones. Studying neural networks, Guo and colleagues found that on most datasets temperature scaling, “a single-parameter variant of Platt Scaling”, is “surprisingly effective at calibrating predictions.”[14] It needs no access to the model, only the probabilities Jev already returns.

Fitted on all 1,232 messages, the temperature came out at about 1.34, confirming the overconfidence, but the held-out numbers are what count. Over 50 random splits, fitting on one half and scoring the other, median ECE fell from 0.092 raw to 0.058, a reduction of about 37%.

The practical question is how many labels it takes to repair Jev calibration. We refit the temperature on 25 to 616 randomly chosen labels from the calibration half and scored the untouched half each time. Most of the gain arrived by about 50 labels: median held-out ECE was 0.059 at 50, 0.057 at 200 and 0.058 at 400. At 25 labels the median was 0.067, but the middle half of splits ran from 0.055 to 0.095: an unlucky sample can leave you close to where you started.

Label budget · real dataVerne's measurement, 30 Sep 2026 · Jev (jev-1.13-20260917) on 1,232 Banking77 test messages, 77 intents
Labels used to fit the temperature
0.0000.0250.0500.0750.1002550100200400616Labels used to fit one temperature (log scale)Held-out ECEno recalibration · 0.09225 labels: median 0.067, middle half 0.055 to 0.09550 labels: median 0.059, middle half 0.052 to 0.066100 labels: median 0.059, middle half 0.052 to 0.066200 labels: median 0.057, middle half 0.051 to 0.061400 labels: median 0.058, middle half 0.053 to 0.064616 labels: median 0.058, middle half 0.053 to 0.0620.059
  • median, one temperature
  • middle half of 50 splits
  • no recalibration
Median held-out ECE
0.059
with 50 labels; raw is 0.092
Middle half of splits
0.052 to 0.066
spread 0.014
Error removed
36%
median against median, same 50 splits

50 labels: median held-out ECE 0.059, middle half 0.052 to 0.066. Most of the gain has arrived, and the spread between splits has tightened.

Figure 4. Verne’s measurement, 30 September 2026. For each budget, one temperature was fitted on that many random labels from one half of the messages, and ECE (15 bins) was scored on the other half, over 50 random splits. The shaded band is the middle half of those splits. Pick a budget, or select a point.

One parameter is deliberate. More flexible methods need more data: Niculescu-Mizil and Caruana found that “A learning curve analysis shows that Isotonic Regression is more prone to overfitting, and thus performs worse than Platt Scaling, when data is scarce.”[15] With a few hundred labels, I would start with one temperature and reach for isotonic regression only when the reliability diagram shows a shape one parameter cannot fix.

Recalibration also changes the scale. After scaling, no message scored above 0.93, because a stated 1.0 with 76 floored zeros becomes 0.927. A threshold of 0.95 that made sense on the raw scale selects nothing on the recalibrated one, so thresholds have to be chosen again after every refit.

Pythonanalyze_jev.py, simplified: fit one temperature, then pick a threshold from measured precision
EPS = 1e-4  # Jev returns many exact zeros; floor them before taking powers


def temp_probs(P: np.ndarray, T: float) -> np.ndarray:
    """Temperature scaling on probabilities: p ** (1 / T), renormalised per message."""
    Q = np.power(np.maximum(P, EPS), 1.0 / T)
    return Q / Q.sum(axis=1, keepdims=True)


def fit_temperature(P: np.ndarray, y: np.ndarray) -> float:
    """The T that maximises the log-likelihood of the true labels (a simple grid search)."""
    grid = np.exp(np.linspace(np.log(0.2), np.log(12), 240))
    nll = [-np.log(temp_probs(P, T)[np.arange(len(y)), y]).mean() for T in grid]
    return float(grid[int(np.argmin(nll))])


def pick_threshold(conf, correct, target=0.95, min_accepted=30):
    """Lowest threshold whose measured precision on labelled data meets the target."""
    for t in np.linspace(0.30, 0.99, 139):
        accepted = conf >= t
        if accepted.sum() >= min_accepted and correct[accepted].mean() >= target:
            return float(t)
    return None  # nothing meets the target: every ticket goes to a person


# P: (messages, 77) probabilities from Jev; y: index of the true intent per message.
# Fit on one half, check on the other. Never score on the rows you fitted on.
T = fit_temperature(P[cal], y[cal])  # about 1.34 on our data
Q_cal, Q_test = temp_probs(P[cal], T), temp_probs(P[test], T)
threshold = pick_threshold(Q_cal.max(1), Q_cal.argmax(1) == y[cal])  # about 0.84
precision = (Q_test.argmax(1) == y[test])[Q_test.max(1) >= threshold].mean()

Setting the routing threshold from measured precision

Confidence threshold routing is selective classification: the system decides the cases it is sure about and hands the rest to a person. Geifman and El-Yaniv framed it as a promise about risk: “Our method allows a user to set a desired risk level. At test time, the classifier rejects instances as needed, to grant the desired risk (with high probability).”[16] In a support queue, the risk level is the precision you demand of tickets that skip a human.

We set a 95% precision target and compared two rules on the held-out halves. Trusting Jev’s own number, auto-accepting at confidence 0.95 or more, accepted 59.7% of messages at 95.1% precision. Choosing the threshold from measured precision after recalibration accepted 58.0% at 95.3%, with the threshold landing around 0.84. Both are medians over the same 50 splits.

So on this dataset the raw number happened to work at the top of the scale, and only just: on all 1,232 messages the same 0.95 rule lands at 94.8%, a hair under the target. I would not build on that luck. The fair-die test shows a stated 82.9% meeting 19.0% accuracy,[5] and in our own middle range, answers stated between 0.6 and 0.7 were right 44% of the time. Even inside the top band, answers stated from 0.90 up to 0.95 were right 65 times out of 90 (from the routing curve).

In Figure 5 the two curves lie almost on top of each other. Temperature scaling changed what the numbers say and barely touched which messages rank above which, so the second rule does not buy more coverage. What it buys is a threshold backed by a measurement you can repeat when the model changes. That is Molas’s point about good scores and good probabilities.

Routing threshold · real dataVerne's measurement, 30 Sep 2026 · Jev (jev-1.13-20260917) on 1,232 Banking77 test messages, 77 intents
Route on
80%85%90%95%100%40%50%60%70%80%90%100%95% targetShare of messages auto-acceptedPrecision of auto-acceptedJev's own confidence ≥ 0.30: 99.9% auto-accepted, 79.5% precisionJev's own confidence ≥ 0.35: 99.6% auto-accepted, 79.7% precisionJev's own confidence ≥ 0.40: 98.7% auto-accepted, 80.1% precisionJev's own confidence ≥ 0.45: 97.2% auto-accepted, 80.5% precisionJev's own confidence ≥ 0.50: 95.1% auto-accepted, 81.7% precisionJev's own confidence ≥ 0.55: 91.4% auto-accepted, 83.6% precisionJev's own confidence ≥ 0.60: 88.8% auto-accepted, 84.8% precisionJev's own confidence ≥ 0.65: 85.7% auto-accepted, 85.9% precisionJev's own confidence ≥ 0.70: 82.6% auto-accepted, 87.7% precisionJev's own confidence ≥ 0.75: 79.1% auto-accepted, 88.7% precisionJev's own confidence ≥ 0.80: 75.5% auto-accepted, 89.9% precisionJev's own confidence ≥ 0.85: 72.5% auto-accepted, 90.8% precisionJev's own confidence ≥ 0.90: 67.1% auto-accepted, 92.4% precisionJev's own confidence ≥ 0.95: 59.8% auto-accepted, 94.8% precisionJev's own confidence ≥ 1.00: 41.3% auto-accepted, 97.5% precisionrecalibrated confidence ≥ 0.30: 99.5% auto-accepted, 79.7% precisionrecalibrated confidence ≥ 0.35: 98.0% auto-accepted, 80.5% precisionrecalibrated confidence ≥ 0.40: 96.2% auto-accepted, 81.3% precisionrecalibrated confidence ≥ 0.45: 93.3% auto-accepted, 82.7% precisionrecalibrated confidence ≥ 0.50: 89.1% auto-accepted, 84.6% precisionrecalibrated confidence ≥ 0.55: 84.9% auto-accepted, 86.6% precisionrecalibrated confidence ≥ 0.60: 80.9% auto-accepted, 88.3% precisionrecalibrated confidence ≥ 0.65: 76.1% auto-accepted, 89.8% precisionrecalibrated confidence ≥ 0.70: 73.5% auto-accepted, 90.5% precisionrecalibrated confidence ≥ 0.75: 69.0% auto-accepted, 91.5% precisionrecalibrated confidence ≥ 0.80: 63.9% auto-accepted, 93.8% precisionrecalibrated confidence ≥ 0.85: 56.1% auto-accepted, 96.0% precisionrecalibrated confidence ≥ 0.90: 47.9% auto-accepted, 96.8% precision≥ 0.95
  • Jev’s own confidence
  • recalibrated
  • 95% precision target
Auto-accepted
59.8%
737 of 1,232 messages
Precision of those
94.8%
below the 95% target · 699 right
To a person, per 1,000 tickets
402
about 6.7 reviewer-hours at your assumed triage time
Wrong but auto-accepted, per 1,000
31
errors no person sees at this threshold
Human triage time per ticket (your assumption)

Jev itself cost about $0.07 per 1,000 decisions on this run. The threshold decides the size of the human queue, and the queue is the cost.

Threshold 0.95 on the Jev's own confidence: 59.8% auto-accepted at 94.8% precision, missing the 95% target. 402 of every 1,000 tickets go to a person. The 228 messages scoring from 0.95 up to 1.00 were right 89% of the time (203 of 228). Lowering the line to here adds them to the auto-accepted pile.

Figure 5. Verne’s measurement, 30 September 2026. Each point is a threshold from 0.30 to 1.00. Curves are computed on all 1,232 messages, and the recalibrated curve uses a temperature fitted on the same messages, so it is in-sample; held-out results for the 95% target are in the text. The triage time is your assumption, not a measurement.

At a 95% target, roughly 400 to 420 of every 1,000 tickets still reach a person. At about $0.07 per 1,000 decisions, the model is the cheap part of this system: the human queue is the cost, and the threshold sets its size. Our guide to human-in-the-loop approval covers what that queue needs so review does not turn into a rubber stamp.

The rule I would ship

  1. Label a few hundred messages drawn at random from your own traffic, not only the easy ones.
  2. Fit one temperature on half of them and draw the reliability diagram on the other half.
  3. On the calibration half, pick the lowest threshold whose precision meets your target, with at least 30 accepted messages behind the estimate.
  4. Confirm that precision on the held-out half before anything goes live.
  5. Send everything below the threshold to a person, and log the probabilities, the threshold and the model version with every decision.
  6. Measure again whenever the model version, the option list or your traffic changes.

The same harness applies to the next decision endpoint. OpenAI’s DevDay community post says: “Decisions API is in limited preview. It uses Luna to classify inputs, route requests, or choose an action from predefined answers.”[17] When it opens up, I would run this exact test on it before comparing prices. Our piece on why enterprise RAG fails quietly makes the same point: averages hide the failures that matter.

Honest limits

This is one public dataset. Banking77 was released in 2020 and may have been in Jev’s training data, which would flatter accuracy and could flatter calibration too. Its labels are not perfect either, which cuts the other way: in at least two of the three zero-probability examples above, I would side with Jev over the label. It is also one model version (jev-1.13-20260917), one question design (77 options named after the intents, with no descriptions) and one day of calls.

Recalibration has a limit of its own. Studying models under dataset shift, Ovadia and colleagues wrote: “We find that traditional post-hoc calibration does indeed fall short, as do several other previous methods.”[18] A temperature fitted on September’s tickets can drift when a product launch changes what customers write about, which is why the last step of the rule exists.

The result says nothing about Jev calibration on your tasks. What it gives you is a way to find out, for under ten cents of API calls and a few hundred labels.

Frequently asked questions

Is Jev calibrated?

It depends on the task. On our Banking77 test (1,232 messages, 30 September 2026), Jev's answers at 0.9 or above were right 92% of the time against a stated average of 0.98, while answers between 0.5 and 0.9 were overconfident by 8 to 21 points. An independent audit on police crash narratives found Jev underconfident instead.

How do you measure Jev calibration?

Label a random sample of your own inputs, send each to Jev as a choice question, and compare the probability on the chosen option with whether it was right, band by band. Expected calibration error (ECE) summarises the gap in one number, and a reliability diagram shows where it sits.

How many labels do you need to recalibrate Jev?

On our test, one temperature fitted on 50 labels brought the median held-out ECE from 0.092 to 0.059, the same as 400 labels (0.058). With 25 labels the result varied widely between random samples.

What confidence threshold should I use to route with Jev?

Choose it from measured precision, not from the stated number. On our data a 95% precision target meant a recalibrated threshold of about 0.84, which auto-accepted 58% of messages; the right value for your task comes from your own labelled sample.

How much does Jev cost per decision?

Our run cost $0.094 for 1,332 calls averaging 1,689 input tokens, about $0.07 per 1,000 decisions. TypeSafe charges per input token and does not charge for output tokens.

Does temperature scaling fix answers where Jev gives the right intent 0?

No. In 6.2% of our messages the correct intent had a probability of exactly 0, and temperature scaling gives every zero the same small value, so there is no ranking left to recover.

Sources

  1. Probability calibration (opens in a new tab) scikit-learn user guide. Definition of a calibrated classifier, stated for the binary case.
  2. Jev Documentation: Using the TypeSafe Decision Model on OpenRouter (opens in a new tab) OpenRouter documentation. What a choice question returns: the selected option, a probability for each option and a confidence value. Accessed 1 October 2026.
  3. Introducing System One models and Jev (opens in a new tab) TypeSafe, 2026-09-15. Vendor claim. Reference probabilities come from larger external models, not ground truth; no calibration metrics are shown.
  4. TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text (opens in a new tab) MarkTechPost, 2026-09-19. Reports the 193.6x and 444.6x figures as TypeSafe's own workflow evals.
  5. [Project]: Jev Does Not Play Dice: probability calibration evaluation (issue #86) (opens in a new tab) awesome-jev-projects, GitHub, 2026-09-23. Third-party fair-die stress test posted as a GitHub issue; not a benchmark of Jev overall.
  6. Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev) (opens in a new tab) Rafe and Das (arXiv 2609.24052), 2026-09-21. Texas crash narratives against a human reference, jev-1.13.0. Quotes from PDF pages 16, 17 and 29.
  7. Jev can't be calibrated (opens in a new tab) Alex Molas, 2026-09-23. Opinion essay.
  8. Evaluating Decision Models for Text Annotation in Computational Social Science (opens in a new tab) Ibrahim and Zaki (arXiv 2609.24574), 2026-09-21. Decision models against the per-task best LLM on 15 annotation tasks.
  9. Open models caught up with Jev (opens in a new tab) LangWatch. Vendor benchmark from a company that builds on Jev. Accessed 1 October 2026.
  10. Language Models (Mostly) Know What They Know (opens in a new tab) Kadavath et al., Anthropic (arXiv), 2022-07.
  11. Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs (opens in a new tab) Xiong et al., ICLR 2024 (arXiv), 2023-06. About verbalised LLM confidence, not Jev.
  12. Banking77 (task-specific-datasets) (opens in a new tab) PolyAI, 2020. Public benchmark, CC BY 4.0; used for our test.
  13. Models (opens in a new tab) TypeSafe documentation. Pricing: charged per input token, output tokens free. Accessed 1 October 2026; re-check on the day you read this.
  14. On Calibration of Modern Neural Networks (opens in a new tab) Guo et al., ICML 2017 (arXiv), 2017-06. Temperature scaling; neural networks, most datasets tested.
  15. Predicting Good Probabilities With Supervised Learning (opens in a new tab) Niculescu-Mizil and Caruana, ICML 2005, 2005. Isotonic regression overfits when calibration data is scarce.
  16. Selective Classification for Deep Neural Networks (opens in a new tab) Geifman and El-Yaniv, NeurIPS 2017 (arXiv), 2017-05. Rejecting inputs to meet a chosen risk level. Selective classification, not calibration.
  17. DevDay 2026 announcements and developer resources (opens in a new tab) OpenAI Developer Community, 2026-09-29. Decisions API in limited preview.
  18. Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift (opens in a new tab) Ovadia et al., NeurIPS 2019 (arXiv), 2019-06. Post-hoc calibration falls short under dataset shift.