Skip to content

Field experiment · AI evaluation

LLM regression testing: the same eval, ten runs, a 4.7-point swing

We ran one model, one prompt and one setting (temperature 0.7) ten times on 150 math problems. Accuracy moved 4.7 points and 22% of run pairs differed by 3 points or more. A paired test on the items flagged none. How to tell a real change from noise.

We ran the same eval ten times on the same day: same model (Llama 3.1 8B Instruct), same prompts, same settings (temperature 0.7), same provider, on 150 GSM8K math problems graded by exact numeric match. Accuracy landed anywhere from 81.3% to 86.0%, a 4.7-point swing with nothing changed. Of the 45 ways to compare two of those runs, 10 differed by 3 points or more, a gap that a simple 3-point rule would report as a regression. A paired test on the individual problems flagged none of them.

That experiment is an A/A test. An A/A test runs two identical arms against each other, so any difference it measures is noise by construction, and the size of that noise tells you how big a real change has to be before your eval can reliably see it. After every model release someone asks me whether the model got worse; this post answers with a measurement.

The argument is older than the latest model releases. In 2023, Chen, Zaharia and Zou reported that “GPT-4 (March 2023) was reasonable at identifying prime vs. composite numbers (84% accuracy) but GPT-4 (June 2023) was poor on these same questions (51% accuracy).”[1] Narayanan and Kapoor replied: “A model that has a capability may or may not display that capability in response to a particular prompt.”[2] Three years on, the same fight restarts whenever a new release draws a ‘nerfed’ thread. Part of the reason it drags on is that the noise is rarely measured first.

Can a pinned model change under you?

The vendors’ position is clear. Anthropic’s documentation says: “Anthropic does not update the weights or configuration of an existing model ID. When an updated version is available, it ships under a new model ID.”[3] OpenAI writes that “Snapshots let you lock in a specific version of the model so that performance and behavior remain consistent.”[4] Google says “Stable models usually don’t change.”[5]

The same Anthropic page adds the part that matters for testing: “Model weights are fixed for a given ID, but the serving infrastructure around the model can change over time. This infrastructure includes components such as the request router, safety classifiers, and sampling logic.”[3] (Its fixed-snapshot guarantee covers the 4.6-and-later dateless IDs.) Anthropic’s September 2025 postmortem shows what the serving layer can do. “We never reduce model quality due to demand, time of day, or server load. The problems our users reported were due to infrastructure bugs alone,” it says, and “At the worst impacted hour on August 31, 16% of Sonnet 4 requests were affected,” meaning misrouted during that hour, not degraded across the board.[6]

The line I would pin above every eval dashboard: “The evaluations we ran simply didn’t capture the degradation users were reporting, in part because Claude often recovers well from isolated mistakes.”[6]

Settings move too. “Claude Opus 5.5 defaults to medium” effort, and Anthropic advises: “Run an effort sweep on your own evals rather than carrying settings over from an earlier model.”[7] When a safety classifier declines a request on the models it covers, “you receive a normal response, not an error, with stop_reason: "refusal".”[8] And OpenAI warns about its seed parameter: “Determinism is not guaranteed, and you should refer to the system_fingerprint response parameter to monitor changes in the backend.”[9]

From the outside, little of this can be checked. An August 2026 study of nine first-party API providers and seven third-party hosts found that “no provider in our sample published information allowing an external party to verify that the artifact being served is the same one referred to in this documentation.”[10] We saw a hint of how much hosts differ. One identical call to the same Llama 3.1 8B model took 13.6 seconds on DeepInfra, 2.2 on Novita and 0.3 on Groq (single calls from our server on 30 September 2026, so indicative only). Hosts can differ in quantisation, hardware and batching, so one model name behind two providers can behave like two services.

So “the model ID did not change” and “the output did not change” are different claims. LLM regression testing exists to check the second one, and it can only do that once you know how much your eval moves when nothing changes.

What we ran

We kept the setup deliberately boring, so that the only thing left to vary was the model’s own randomness.

Method · Verne A/A test, 30 September 2026

Tasks
150 problems sampled with seed 11 from the GSM8K test set, OpenAI’s grade-school math benchmark (MIT licence)[11]
Model
meta-llama/llama-3.1-8b-instruct through OpenRouter[12], pinned to Groq with fallbacks disabled
Prompt
One system prompt and one user template for every call
Grading
Exact numeric match on the final “Answer:” line
Runs
Ten full runs at temperature 0.7, five at temperature 0
Volume
2,250 calls, zero errors, $0.056 in total. Every response reported the same model and provider.

Problems with one numeric answer take the grader out of the noise, so every flip belongs to the model, and an open model lets us pin the host. Here is the core of the harness, simplified from our script (which also runs eight requests in parallel and stops if spending passes one dollar).

PythonThe A/A harness: one model, one host, fifteen identical passes over 150 problems
import json, random, re, urllib.request

MODEL = "meta-llama/llama-3.1-8b-instruct"
# Pin the host as well as the model. Without allow_fallbacks=False, OpenRouter
# may send a request to another provider when the first one is busy.
PROVIDER = {"order": ["Groq"], "allow_fallbacks": False}
SYSTEM = "You are a careful math tutor."
SUFFIX = "\n\nSolve step by step, then give the final answer on its own line as: Answer: <number>"


def norm(x):
    """'1,200.' -> '1200', '2.50' -> '2.5'; None if it is not a number."""
    try:
        v = float(x.replace(",", "").rstrip("."))
        return str(int(v)) if v == int(v) else str(v)
    except (AttributeError, ValueError):
        return None


def parse(text):
    # Grade the last "Answer:" line; fall back to the last number in the reply.
    hits = re.findall(r"Answer:\s*\$?\s*(-?\d[\d,]*\.?\d*)", text or "")
    nums = hits or re.findall(r"-?\d[\d,]*\.?\d*", text or "")
    return norm(nums[-1]) if nums else None


def call(question, temperature):
    body = {
        "model": MODEL,
        "messages": [
            {"role": "system", "content": SYSTEM},
            {"role": "user", "content": question + SUFFIX},
        ],
        "temperature": temperature,
        "max_tokens": 512,
        "provider": PROVIDER,
    }
    req = urllib.request.Request(
        "https://openrouter.ai/api/v1/chat/completions",
        data=json.dumps(body).encode(),
        headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
    )
    with urllib.request.urlopen(req, timeout=120) as r:
        return json.load(r)


def run(items, tag, temperature, run_id, out):
    for qid, row in items:
        gold = norm(row["answer"].split("####")[-1].strip())
        res = call(row["question"], temperature)
        pred = parse(res["choices"][0]["message"]["content"])
        out.write(json.dumps({
            "tag": tag, "run": run_id, "qid": qid, "gold": gold, "pred": pred,
            "correct": pred is not None and pred == gold,
            # Keep what each response says it was, so you can show nothing was rerouted.
            "model": res.get("model"), "provider": res.get("provider"),
        }) + "\n")


# 150 GSM8K test problems, sampled once with a fixed seed, then frozen.
rows = [json.loads(line) for line in open("gsm8k_test.jsonl")]
items = random.Random(11).sample(list(enumerate(rows)), 150)

with open("results.jsonl", "w") as out:
    for r in range(10):
        run(items, "t0.7", 0.7, r, out)   # ten identical runs
    for r in range(5):
        run(items, "t0", 0.0, r, out)     # five greedy runs

Two details carry the weight: allow_fallbacks: False stops OpenRouter from quietly routing a request to another host, and each record keeps the model and provider the response reported, so you can show afterwards that nothing was rerouted.

What identical runs look like

Here are all fifteen runs. The ten at temperature 0.7 spread from 81.3% to 86.0%, with a mean of 83.9% and a standard deviation of 1.66 points. Pick any two, call one the old version and the other the new one, and see what a simple rule concludes.

Pick two identical runsVerne measurement · 30 Sep 2026

Click a dot to set the highlighted slot. Every run used the same model, prompts, settings and provider.

Temperature 0.7 · 10 runs

Temperature 0 · 5 runs

Change, B minus A

−4.67 pts

86.0% → 81.3%

Naive rule: a 3-point gap is real

Regression

A false alarm: nothing changed between these runs.

Paired t-test on the 150 items

No change

13 items went right to wrong, 6 wrong to right. |t| = 1.61; a change needs more than 1.98.

Run 4 scored 86.0%, Run 2 scored 81.3%: −4.67 points. The 3-point rule calls this a regression. The paired test on the 150 items detects no change (t = -1.61).

Figure 1. Verne’s A/A test, 30 September 2026. Each dot is one full run of the same 150 GSM8K problems through Llama 3.1 8B Instruct on Groq, with identical prompts and settings. The widest gap, run 2 against run 4 (or run 10), is 4.67 points; the paired test on the items still finds no change.

Across all 45 pairs of temperature-0.7 runs, 26 differ by 2 points or more, 10 (22%) by 3 or more and 4 by 4 or more; none reach 5. A rule such as “a 3-point drop is a regression” would raise a false alarm on 10 of 45 comparisons of identical systems. A paired t-test on per-item correctness, which compares each problem with itself across the two runs, flagged 0 of 45.

How often a gap rule cries wolfVerne measurement · 30 Sep 2026
Compare

Gap rule flags

10/45

22% of comparisons between identical systems

Paired t-test flags

0/45

No false alarms at 5% significance

All 45 pairs of the ten temperature-0.7 runs. Every comparison is between identical systems, so every flag is a false alarm.

With a 3-point rule, 10 of 45 pairs of single runs are flagged (22%). The paired t-test flags 0 of 45.

Figure 2. Every gap between two sides of Verne’s A/A test (30 September 2026), measured on the same 150 problems. Move the threshold to see how often a gap rule fires on identical systems. The five-run view compares averages of five runs against the other five, in all 126 possible splits.

Where the swing comes from

The totals hide how much moves underneath. Across the ten temperature-0.7 runs, 59 of the 150 problems (39%) were right in some runs and wrong in others, and 63 produced more than one distinct answer. Only 86 were right every time, and 5 were wrong every time. The flips mostly cancel, which is why accuracy stays inside a band, but the set of problems that fail changes on every run.

150 problems × every runVerne measurement · 30 Sep 2026
Runs
Sort
  • 86 right every run
  • 5 wrong every run
  • 59 flip between right and wrong
  • 63 gave more than one answer

GSM8K test item 953 · correct answer 452 · right in 2 of 10 runs · 9 distinct answers

  1. R1392 wrong
  2. R2332 wrong
  3. R3197 wrong
  4. R4247 wrong
  5. R5227 wrong
  6. R6452 right
  7. R7342 wrong
  8. R8452 right
  9. R9477 wrong
  10. R10137 wrong
Figure 3. All 150 problems from Verne’s A/A test (30 September 2026), one tile each, one stripe per run. Coral outlines mark problems that were right in some runs and wrong in others. Select a tile, or use the arrow keys, to see every answer the model gave.

Temperature 0 narrows the band without closing it. The five greedy runs scored 83.3% to 84.0%, a range of 0.67 points, and one problem still changed its answer: GSM8K test item 953 came back as 332 in three runs and as 452, the correct answer, in two.

Larger studies see the same thing

On SWE-Bench-Verified, across 60,000 agent trajectories, Bjarnason, Silva and Monperrus found that “single-run pass@1 estimates vary by 2.2 to 6.0 percentage points depending on which run is selected, with standard deviations exceeding 1.5 percentage points even at temperature 0.”[13] Yuan and colleagues report that “DeepSeek-R1-Distill-Qwen-7B can exhibit up to 9% variation in accuracy and 9,000 tokens difference in response length due to differences in GPU count, type, and evaluation batch size.”[14]

Thinking Machines Lab traced much of this to serving: “the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies!” Sampling Qwen3-235B 1,000 times at temperature 0, they got “80 unique completions, with the most common of these occuring 78 times” (their spelling).[15] Greedy decoding gives a narrower band, not a guarantee.

How big must a change be before you believe it?

A single accuracy number from a 150-item eval carries more uncertainty than it shows. Miller’s paper on error bars sets out the frame: “Conceptualizing evaluation questions as having been drawn from an unseen super-population, we present formulas for analyzing evaluation data.”[16] Bowyer, Aitchison and Ivanova warn that on small benchmarks “CLT-based methods perform very poorly, usually dramatically underestimating uncertainty (i.e. producing error bars that are too small).”[17] For chain-of-thought evals, Anthropic recommends “resampling answers from the same model several times, and using the question-level averages as the question scores.”[18]

Our A/A data turns that advice into a number you can plan with. The minimum detectable effect (MDE) is the smallest true change an eval catches 80% of the time at 5% significance. Using the per-item run variance we measured (0.0635), a paired comparison on 150 items with one run per version has an MDE of 8.15 points. Five runs per version bring it to 3.6, 300 items with five runs to 2.6, 500 items with five runs to 2.0, and 500 items with ten runs to 1.4.

Minimum detectable effectFrom measured variance 0.0635 · 30 Sep 2026

Smallest reliable change

8.15 pts

80% power, 5% significance, paired

A real 3-point change

caught 18% of the time

Needs about 1,106 items at 1 run per version.

Calls per comparison

300

150 items × 1 × 2 versions. Our 2,250 calls on the 8B model cost $0.056.

MDE = 2.8 × √(2 × 0.0635 ÷ (1 × 150)) = 8.15 points

150 items with 1 run per version: smallest reliable change 8.15 points. A real 3-point change would be caught 18% of the time. You would need about 1106 items at this number of runs.

Figure 4. Minimum detectable effect for a paired comparison of two versions, computed from the per-item run variance in Verne’s A/A test (0.0635, ten runs at temperature 0.7, 30 September 2026). Your tasks will have their own variance; the shape of the trade-off carries over.

Average the runs, then test the items

Averaging works in our data. We split the ten runs into two groups of five in every possible way (126 splits) and compared the group averages. The largest difference fell to 2.67 points, so the naive 3-point rule flagged none, and the paired test flagged 1 of 126 (0.8%), under the 5% false-alarm rate it allows.

Decision rules on identical systems
SetupComparisonsFlagged as a changeWidest gapMDE, 150 items
One run each, 3-point gap rule4510 (22%)4.67 ptsnot defined
One run each, paired t-test4504.67 pts8.15 pts
Five runs each, 3-point gap rule12602.67 ptsnot defined
Five runs each, paired t-test1261 (0.8%)2.67 pts3.64 pts
Temperature 0, one run each, either rule1000.67 ptsnot measured
Table 1. Every comparison is between identical systems, so every flag is a false alarm. Verne’s A/A test, 30 September 2026: 150 GSM8K problems, Llama 3.1 8B Instruct on Groq, temperature 0.7 unless noted. MDE at 80% power and 5% significance.

A gap rule has no idea how noisy your eval is. The paired test keeps false alarms near the rate you chose, and the MDE tells you what the eval cannot reliably see. Many simple gap-rule setups compute neither.

Livenerf, an open project on GitHub that tracks whether hosted models change, writes its rule down in advance: “A change has to clear a 99% interval in both 10-day windows that follow, be at least 3 points, and not show up in the control arm.”[19] Its interim reading shows how wide the uncertainty can be: “Swapping in Opus 5 was not distinguishable from Opus 5.5 at 99% (−3.8 ± 6.3 points, −23% tokens).”[19] A 3.8-point gap between two different models sat inside the noise: not proof they are the same, only that the test could not yet tell them apart (6 of its 30 days were in when we read it on 29 September).

PythonThe analysis: gap rule against paired test, and the minimum detectable effect
import itertools, json, math
import numpy as np

T_CRIT = 1.976  # two-sided 5% with 149 degrees of freedom

recs = [json.loads(line) for line in open("results.jsonl")]
qids = sorted({r["qid"] for r in recs})
col = {q: i for i, q in enumerate(qids)}

# C[run, item] = 1 if that run got that problem right (temperature 0.7 runs).
C = np.zeros((10, len(qids)))
for r in recs:
    if r["tag"] == "t0.7":
        C[r["run"], col[r["qid"]]] = r["correct"]

acc = C.mean(axis=1) * 100
print(acc.max() - acc.min())   # 4.67 points between identical runs


def paired_t(a, b):
    """Paired t on per-item scores: each problem is compared with itself."""
    d = a - b
    se = d.std(ddof=1) / math.sqrt(len(d))
    return d.mean() / se if se > 0 else 0.0


naive = paired = 0
for i, j in itertools.combinations(range(10), 2):
    naive += abs(acc[i] - acc[j]) >= 3              # "a 3-point gap is real"
    paired += abs(paired_t(C[i], C[j])) > T_CRIT    # the paired test
print(naive, paired)   # 10 and 0, of 45 pairs


# Per-item run variance: how much one problem's score wobbles between runs.
p = C.mean(axis=0)
within = np.mean(p * (1 - p))   # 0.0635


def mde(n_items, k_runs, var=within):
    """Smallest true change caught 80% of the time at 5% significance, in points."""
    # 2.8 = 1.96 (5%, two-sided) + 0.84 (80% power)
    return 2.8 * math.sqrt(2 * var / (k_runs * n_items)) * 100


print(mde(150, 1), mde(150, 5), mde(500, 10))   # 8.15, 3.64, 1.41

The paired test works because it compares each problem with itself. Problems that both versions got right, or both got wrong, contribute a difference of zero, so the verdict rests on the problems that actually changed.

An LLM regression testing setup you can run this week

This is the procedure we recommend, and the one the experiment followed. It takes about a day to build and runs on every model, prompt or provider change.

The regression gate, step by step

Step 1: Freeze the tasks

A fixed set of problems with answers a program can check, versioned with the code.

  • items: 150 (GSM8K test, seed 11)
  • grader: exact match on the Answer: line
  • 12 shown here
Figure 5. The six steps as a gate. The dots are real: the first 12 of our 150 problems, with runs 1 to 5 standing in for the baseline and runs 6 to 10 for the candidate. Because both sides are the same system, the correct verdict is “no change”, and the gate reaches it.
  1. Freeze the task set. At least 150 items drawn from your real traffic, with answers a program can check, versioned next to the code. If you need to see a 3-point change, plan on about 220 items run five times each.
  2. Pin what you can. Model ID, effort or temperature, token limits, system prompt, and the provider with fallbacks off. Store the model and provider that every response reports.
  3. Repeat every item. Check your harness’s default. UK AISI’s Inspect, for example, documents its repeat option as “Number of times to repeat each sample (defaults to 1)”.[20] Set it to five.
  4. Compare items, not totals. Run a paired test on per-item mean scores between baseline and candidate.
  5. Write the decision rule first. For example: a change counts only when the paired test clears 5% and the gap is larger than your MDE. Decide before you see the numbers.
  6. Keep a control arm. Rerun the unchanged baseline beside the candidate. If the control moves, your pipeline or provider moved, and the comparison is void.

None of this needs a platform, only the discipline to hold everything fixed except the thing under test, the same idea behind our other evaluation notes: scoring RAG by slice instead of by average, checking whether a model’s confidence means what it says, and turning reviewer decisions into the evaluation set that gates releases.

If your LLM regression testing finds a real change, you will have evidence instead of a thread. If it finds nothing, you still get the number that settles most arguments early: how big a change would have to be before anyone should believe it.

Honest limits

This was one small open model (Llama 3.1 8B Instruct), one task type (150 GSM8K math problems with deterministic answers), one provider and one day. Frontier models on your tasks will have their own noise level, and open-ended tasks graded by a model add the judge’s noise on top. The method transfers; the numbers do not.

The eval also has a floor. It cannot reliably detect a change smaller than its minimum detectable effect, so a clean result only means no change was large enough to see reliably. It says nothing about whether any specific vendor changed any specific model; answering that would take the same design run over weeks against the vendor’s endpoint, with a control arm beside it.

Frequently asked questions

What is an A/A test for LLMs?

An A/A test runs the same model, prompts and settings against themselves, so any difference in scores is noise by construction. It measures how much your eval moves when nothing changes, which tells you how large a real change must be before you can trust it.

How can I tell if a model really got worse?

Run a frozen task set several times on the old and new setup, compare per-item scores with a paired test, and check the gap against your eval's minimum detectable effect. In our A/A test, a single 150-item run could not reliably detect a change smaller than 8.15 points.

Does temperature 0 make LLM outputs deterministic?

No. In our test, five temperature-0 runs still differed by up to 0.67 points and one problem changed its answer. Thinking Machines Lab names batch size, which varies with server load, as the primary reason nearly all inference endpoints are nondeterministic.

How many test cases do I need for LLM regression testing?

It depends on the change you need to see and on your model's noise. With the per-item variance we measured, catching a 3-point change 80% of the time takes about 220 items run five times each, or about 1,100 items run once, so measure your own variance with an A/A test first.

Do pinned model versions ever change?

Anthropic says it does not update the weights or configuration of an existing model ID, and also that the serving infrastructure around a model, such as the request router, safety classifiers and sampling logic, can change over time. Its September 2025 postmortem traced user-reported quality problems to infrastructure bugs.

What is a minimum detectable effect in an eval?

It is the smallest true change an eval will catch 80% of the time at 5% significance, given its number of items, its repeats and its noise. For our 150-item eval with one run per version it was 8.15 points.

Sources

  1. How is ChatGPT's behavior changing over time? (opens in a new tab) Chen, Zaharia and Zou (arXiv 2307.09009), 2023-07. The 2023 drift claim: GPT-4 prime-number accuracy fell from 84% (March) to 51% (June).
  2. Is GPT-4 getting worse over time? (opens in a new tab) Narayanan and Kapoor (normaltech.ai), 2023-07-19. The critique: capability and behaviour on a particular prompt are different things.
  3. Model IDs and versions (opens in a new tab) Anthropic documentation. Weights fixed per model ID; the serving infrastructure can change. The fixed-snapshot guarantee covers 4.6-and-later dateless IDs.
  4. GPT-5 model documentation (opens in a new tab) OpenAI API documentation. Snapshots lock in a specific model version.
  5. Gemini models (opens in a new tab) Google AI for Developers. Stable models usually don't change.
  6. A postmortem of three recent issues (opens in a new tab) Anthropic Engineering, 2025-09. 16% is the share of Sonnet 4 requests misrouted at the worst hour, not the share degraded overall.
  7. Effort (opens in a new tab) Anthropic documentation. Claude Opus 5.5 defaults to medium effort; run an effort sweep on your own evals.
  8. Refusals and fallback (opens in a new tab) Anthropic documentation. Classifier refusals on the listed models return a normal response with stop_reason refusal.
  9. How to make your completions outputs reproducible with the new seed parameter (opens in a new tab) OpenAI Cookbook. Determinism is not guaranteed; monitor system_fingerprint.
  10. Silent Updates: Measuring and Closing the Post-Deployment Disclosure Gap (opens in a new tab) Abraham and Bucknall (arXiv 2608.11803), 2026-08. Sample: nine first-party API providers and seven third-party hosts.
  11. GSM8K (grade-school-math) (opens in a new tab) OpenAI (GitHub), 2021. Test split, MIT licence; used for our A/A test.
  12. Llama 3.1 8B Instruct on OpenRouter (opens in a new tab) OpenRouter. Model used, pinned to Groq.
  13. On Randomness in Agentic Evals (opens in a new tab) Bjarnason, Silva and Monperrus (arXiv 2602.07150), 2026-02. SWE-Bench-Verified, 60,000 trajectories.
  14. Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference (opens in a new tab) Yuan et al., NeurIPS 2025 (arXiv 2506.09501), 2025-06. One 7B model, bf16, greedy decoding.
  15. Defeating Nondeterminism in LLM Inference (opens in a new tab) Thinking Machines Lab, 2025-09. Qwen3-235B, 1,000 samples at temperature 0.
  16. Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (opens in a new tab) Miller (arXiv 2411.00640), 2024-11.
  17. Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints (opens in a new tab) Bowyer, Aitchison and Ivanova, ICML 2025 (arXiv 2503.01747), 2025-03. Applies to small benchmarks, under a few hundred items.
  18. A statistical approach to model evals (opens in a new tab) Anthropic Research, 2024-11. Resampling recommended for chain-of-thought evals.
  19. Livenerf (opens in a new tab) ninjahawk (GitHub). Live repository, read on 29 September 2026 with 6 of 30 days collected; readings are interim.
  20. Inspect: options reference (opens in a new tab) UK AI Security Institute. Repeats per sample default to 1.