We ran the same eval ten times on the same day: same model (Llama 3.1 8B Instruct), same prompts, same settings (temperature 0.7), same provider, on 150 GSM8K math problems graded by exact numeric match. Accuracy landed anywhere from 81.3% to 86.0%, a 4.7-point swing with nothing changed. Of the 45 ways to compare two of those runs, 10 differed by 3 points or more, a gap that a simple 3-point rule would report as a regression. A paired test on the individual problems flagged none of them.
That experiment is an A/A test. An A/A test runs two identical arms against each other, so any difference it measures is noise by construction, and the size of that noise tells you how big a real change has to be before your eval can reliably see it. After every model release someone asks me whether the model got worse; this post answers with a measurement.
The argument is older than the latest model releases. In 2023, Chen, Zaharia and Zou reported that “GPT-4 (March 2023) was reasonable at identifying prime vs. composite numbers (84% accuracy) but GPT-4 (June 2023) was poor on these same questions (51% accuracy).”[1] Narayanan and Kapoor replied: “A model that has a capability may or may not display that capability in response to a particular prompt.”[2] Three years on, the same fight restarts whenever a new release draws a ‘nerfed’ thread. Part of the reason it drags on is that the noise is rarely measured first.
Can a pinned model change under you?
The vendors’ position is clear. Anthropic’s documentation says: “Anthropic does not update the weights or configuration of an existing model ID. When an updated version is available, it ships under a new model ID.”[3] OpenAI writes that “Snapshots let you lock in a specific version of the model so that performance and behavior remain consistent.”[4] Google says “Stable models usually don’t change.”[5]
The same Anthropic page adds the part that matters for testing: “Model weights are fixed for a given ID, but the serving infrastructure around the model can change over time. This infrastructure includes components such as the request router, safety classifiers, and sampling logic.”[3] (Its fixed-snapshot guarantee covers the 4.6-and-later dateless IDs.) Anthropic’s September 2025 postmortem shows what the serving layer can do. “We never reduce model quality due to demand, time of day, or server load. The problems our users reported were due to infrastructure bugs alone,” it says, and “At the worst impacted hour on August 31, 16% of Sonnet 4 requests were affected,” meaning misrouted during that hour, not degraded across the board.[6]
The line I would pin above every eval dashboard: “The evaluations we ran simply didn’t capture the degradation users were reporting, in part because Claude often recovers well from isolated mistakes.”[6]
Settings move too. “Claude Opus 5.5 defaults to medium” effort, and Anthropic advises: “Run an effort sweep on your own evals rather than carrying settings over from an earlier model.”[7] When a safety classifier declines a request on the models it covers, “you receive a normal response, not an error, with stop_reason: "refusal".”[8] And OpenAI warns about its seed parameter: “Determinism is not guaranteed, and you should refer to the system_fingerprint response parameter to monitor changes in the backend.”[9]
From the outside, little of this can be checked. An August 2026 study of nine first-party API providers and seven third-party hosts found that “no provider in our sample published information allowing an external party to verify that the artifact being served is the same one referred to in this documentation.”[10] We saw a hint of how much hosts differ. One identical call to the same Llama 3.1 8B model took 13.6 seconds on DeepInfra, 2.2 on Novita and 0.3 on Groq (single calls from our server on 30 September 2026, so indicative only). Hosts can differ in quantisation, hardware and batching, so one model name behind two providers can behave like two services.
So “the model ID did not change” and “the output did not change” are different claims. LLM regression testing exists to check the second one, and it can only do that once you know how much your eval moves when nothing changes.
What we ran
We kept the setup deliberately boring, so that the only thing left to vary was the model’s own randomness.
Method · Verne A/A test, 30 September 2026
- Tasks
- 150 problems sampled with seed 11 from the GSM8K test set, OpenAI’s grade-school math benchmark (MIT licence)[11]
- Model
meta-llama/llama-3.1-8b-instructthrough OpenRouter[12], pinned to Groq with fallbacks disabled- Prompt
- One system prompt and one user template for every call
- Grading
- Exact numeric match on the final “Answer:” line
- Runs
- Ten full runs at temperature 0.7, five at temperature 0
- Volume
- 2,250 calls, zero errors, $0.056 in total. Every response reported the same model and provider.
Problems with one numeric answer take the grader out of the noise, so every flip belongs to the model, and an open model lets us pin the host. Here is the core of the harness, simplified from our script (which also runs eight requests in parallel and stops if spending passes one dollar).
import json, random, re, urllib.request
MODEL = "meta-llama/llama-3.1-8b-instruct"
# Pin the host as well as the model. Without allow_fallbacks=False, OpenRouter
# may send a request to another provider when the first one is busy.
PROVIDER = {"order": ["Groq"], "allow_fallbacks": False}
SYSTEM = "You are a careful math tutor."
SUFFIX = "\n\nSolve step by step, then give the final answer on its own line as: Answer: <number>"
def norm(x):
"""'1,200.' -> '1200', '2.50' -> '2.5'; None if it is not a number."""
try:
v = float(x.replace(",", "").rstrip("."))
return str(int(v)) if v == int(v) else str(v)
except (AttributeError, ValueError):
return None
def parse(text):
# Grade the last "Answer:" line; fall back to the last number in the reply.
hits = re.findall(r"Answer:\s*\$?\s*(-?\d[\d,]*\.?\d*)", text or "")
nums = hits or re.findall(r"-?\d[\d,]*\.?\d*", text or "")
return norm(nums[-1]) if nums else None
def call(question, temperature):
body = {
"model": MODEL,
"messages": [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": question + SUFFIX},
],
"temperature": temperature,
"max_tokens": 512,
"provider": PROVIDER,
}
req = urllib.request.Request(
"https://openrouter.ai/api/v1/chat/completions",
data=json.dumps(body).encode(),
headers={"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=120) as r:
return json.load(r)
def run(items, tag, temperature, run_id, out):
for qid, row in items:
gold = norm(row["answer"].split("####")[-1].strip())
res = call(row["question"], temperature)
pred = parse(res["choices"][0]["message"]["content"])
out.write(json.dumps({
"tag": tag, "run": run_id, "qid": qid, "gold": gold, "pred": pred,
"correct": pred is not None and pred == gold,
# Keep what each response says it was, so you can show nothing was rerouted.
"model": res.get("model"), "provider": res.get("provider"),
}) + "\n")
# 150 GSM8K test problems, sampled once with a fixed seed, then frozen.
rows = [json.loads(line) for line in open("gsm8k_test.jsonl")]
items = random.Random(11).sample(list(enumerate(rows)), 150)
with open("results.jsonl", "w") as out:
for r in range(10):
run(items, "t0.7", 0.7, r, out) # ten identical runs
for r in range(5):
run(items, "t0", 0.0, r, out) # five greedy runsTwo details carry the weight: allow_fallbacks: False stops OpenRouter from quietly routing a request to another host, and each record keeps the model and provider the response reported, so you can show afterwards that nothing was rerouted.
What identical runs look like
Here are all fifteen runs. The ten at temperature 0.7 spread from 81.3% to 86.0%, with a mean of 83.9% and a standard deviation of 1.66 points. Pick any two, call one the old version and the other the new one, and see what a simple rule concludes.
Click a dot to set the highlighted slot. Every run used the same model, prompts, settings and provider.
Temperature 0.7 · 10 runs
Temperature 0 · 5 runs
Change, B minus A
−4.67 pts
86.0% → 81.3%
Naive rule: a 3-point gap is real
Regression
A false alarm: nothing changed between these runs.
Paired t-test on the 150 items
No change
13 items went right to wrong, 6 wrong to right. |t| = 1.61; a change needs more than 1.98.
Run 4 scored 86.0%, Run 2 scored 81.3%: −4.67 points. The 3-point rule calls this a regression. The paired test on the 150 items detects no change (t = -1.61).
Across all 45 pairs of temperature-0.7 runs, 26 differ by 2 points or more, 10 (22%) by 3 or more and 4 by 4 or more; none reach 5. A rule such as “a 3-point drop is a regression” would raise a false alarm on 10 of 45 comparisons of identical systems. A paired t-test on per-item correctness, which compares each problem with itself across the two runs, flagged 0 of 45.
Gap rule flags
10/45
22% of comparisons between identical systems
Paired t-test flags
0/45
No false alarms at 5% significance
All 45 pairs of the ten temperature-0.7 runs. Every comparison is between identical systems, so every flag is a false alarm.
With a 3-point rule, 10 of 45 pairs of single runs are flagged (22%). The paired t-test flags 0 of 45.
Where the swing comes from
The totals hide how much moves underneath. Across the ten temperature-0.7 runs, 59 of the 150 problems (39%) were right in some runs and wrong in others, and 63 produced more than one distinct answer. Only 86 were right every time, and 5 were wrong every time. The flips mostly cancel, which is why accuracy stays inside a band, but the set of problems that fail changes on every run.
- 86 right every run
- 5 wrong every run
- 59 flip between right and wrong
- 63 gave more than one answer
GSM8K test item 953 · correct answer 452 · right in 2 of 10 runs · 9 distinct answers
- R1392 wrong
- R2332 wrong
- R3197 wrong
- R4247 wrong
- R5227 wrong
- R6452 right
- R7342 wrong
- R8452 right
- R9477 wrong
- R10137 wrong
Temperature 0 narrows the band without closing it. The five greedy runs scored 83.3% to 84.0%, a range of 0.67 points, and one problem still changed its answer: GSM8K test item 953 came back as 332 in three runs and as 452, the correct answer, in two.
Larger studies see the same thing
On SWE-Bench-Verified, across 60,000 agent trajectories, Bjarnason, Silva and Monperrus found that “single-run pass@1 estimates vary by 2.2 to 6.0 percentage points depending on which run is selected, with standard deviations exceeding 1.5 percentage points even at temperature 0.”[13] Yuan and colleagues report that “DeepSeek-R1-Distill-Qwen-7B can exhibit up to 9% variation in accuracy and 9,000 tokens difference in response length due to differences in GPU count, type, and evaluation batch size.”[14]
Thinking Machines Lab traced much of this to serving: “the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies!” Sampling Qwen3-235B 1,000 times at temperature 0, they got “80 unique completions, with the most common of these occuring 78 times” (their spelling).[15] Greedy decoding gives a narrower band, not a guarantee.
How big must a change be before you believe it?
A single accuracy number from a 150-item eval carries more uncertainty than it shows. Miller’s paper on error bars sets out the frame: “Conceptualizing evaluation questions as having been drawn from an unseen super-population, we present formulas for analyzing evaluation data.”[16] Bowyer, Aitchison and Ivanova warn that on small benchmarks “CLT-based methods perform very poorly, usually dramatically underestimating uncertainty (i.e. producing error bars that are too small).”[17] For chain-of-thought evals, Anthropic recommends “resampling answers from the same model several times, and using the question-level averages as the question scores.”[18]
Our A/A data turns that advice into a number you can plan with. The minimum detectable effect (MDE) is the smallest true change an eval catches 80% of the time at 5% significance. Using the per-item run variance we measured (0.0635), a paired comparison on 150 items with one run per version has an MDE of 8.15 points. Five runs per version bring it to 3.6, 300 items with five runs to 2.6, 500 items with five runs to 2.0, and 500 items with ten runs to 1.4.
Smallest reliable change
8.15 pts
80% power, 5% significance, paired
A real 3-point change
caught 18% of the time
Needs about 1,106 items at 1 run per version.
Calls per comparison
300
150 items × 1 × 2 versions. Our 2,250 calls on the 8B model cost $0.056.
MDE = 2.8 × √(2 × 0.0635 ÷ (1 × 150)) = 8.15 points
150 items with 1 run per version: smallest reliable change 8.15 points. A real 3-point change would be caught 18% of the time. You would need about 1106 items at this number of runs.
Average the runs, then test the items
Averaging works in our data. We split the ten runs into two groups of five in every possible way (126 splits) and compared the group averages. The largest difference fell to 2.67 points, so the naive 3-point rule flagged none, and the paired test flagged 1 of 126 (0.8%), under the 5% false-alarm rate it allows.
| Setup | Comparisons | Flagged as a change | Widest gap | MDE, 150 items |
|---|---|---|---|---|
| One run each, 3-point gap rule | 45 | 10 (22%) | 4.67 pts | not defined |
| One run each, paired t-test | 45 | 0 | 4.67 pts | 8.15 pts |
| Five runs each, 3-point gap rule | 126 | 0 | 2.67 pts | not defined |
| Five runs each, paired t-test | 126 | 1 (0.8%) | 2.67 pts | 3.64 pts |
| Temperature 0, one run each, either rule | 10 | 0 | 0.67 pts | not measured |
A gap rule has no idea how noisy your eval is. The paired test keeps false alarms near the rate you chose, and the MDE tells you what the eval cannot reliably see. Many simple gap-rule setups compute neither.
Livenerf, an open project on GitHub that tracks whether hosted models change, writes its rule down in advance: “A change has to clear a 99% interval in both 10-day windows that follow, be at least 3 points, and not show up in the control arm.”[19] Its interim reading shows how wide the uncertainty can be: “Swapping in Opus 5 was not distinguishable from Opus 5.5 at 99% (−3.8 ± 6.3 points, −23% tokens).”[19] A 3.8-point gap between two different models sat inside the noise: not proof they are the same, only that the test could not yet tell them apart (6 of its 30 days were in when we read it on 29 September).
import itertools, json, math
import numpy as np
T_CRIT = 1.976 # two-sided 5% with 149 degrees of freedom
recs = [json.loads(line) for line in open("results.jsonl")]
qids = sorted({r["qid"] for r in recs})
col = {q: i for i, q in enumerate(qids)}
# C[run, item] = 1 if that run got that problem right (temperature 0.7 runs).
C = np.zeros((10, len(qids)))
for r in recs:
if r["tag"] == "t0.7":
C[r["run"], col[r["qid"]]] = r["correct"]
acc = C.mean(axis=1) * 100
print(acc.max() - acc.min()) # 4.67 points between identical runs
def paired_t(a, b):
"""Paired t on per-item scores: each problem is compared with itself."""
d = a - b
se = d.std(ddof=1) / math.sqrt(len(d))
return d.mean() / se if se > 0 else 0.0
naive = paired = 0
for i, j in itertools.combinations(range(10), 2):
naive += abs(acc[i] - acc[j]) >= 3 # "a 3-point gap is real"
paired += abs(paired_t(C[i], C[j])) > T_CRIT # the paired test
print(naive, paired) # 10 and 0, of 45 pairs
# Per-item run variance: how much one problem's score wobbles between runs.
p = C.mean(axis=0)
within = np.mean(p * (1 - p)) # 0.0635
def mde(n_items, k_runs, var=within):
"""Smallest true change caught 80% of the time at 5% significance, in points."""
# 2.8 = 1.96 (5%, two-sided) + 0.84 (80% power)
return 2.8 * math.sqrt(2 * var / (k_runs * n_items)) * 100
print(mde(150, 1), mde(150, 5), mde(500, 10)) # 8.15, 3.64, 1.41The paired test works because it compares each problem with itself. Problems that both versions got right, or both got wrong, contribute a difference of zero, so the verdict rests on the problems that actually changed.
An LLM regression testing setup you can run this week
This is the procedure we recommend, and the one the experiment followed. It takes about a day to build and runs on every model, prompt or provider change.
Step 1: Freeze the tasks
A fixed set of problems with answers a program can check, versioned with the code.
- items: 150 (GSM8K test, seed 11)
- grader: exact match on the Answer: line
- 12 shown here
- Freeze the task set. At least 150 items drawn from your real traffic, with answers a program can check, versioned next to the code. If you need to see a 3-point change, plan on about 220 items run five times each.
- Pin what you can. Model ID, effort or temperature, token limits, system prompt, and the provider with fallbacks off. Store the model and provider that every response reports.
- Repeat every item. Check your harness’s default. UK AISI’s Inspect, for example, documents its repeat option as “Number of times to repeat each sample (defaults to 1)”.[20] Set it to five.
- Compare items, not totals. Run a paired test on per-item mean scores between baseline and candidate.
- Write the decision rule first. For example: a change counts only when the paired test clears 5% and the gap is larger than your MDE. Decide before you see the numbers.
- Keep a control arm. Rerun the unchanged baseline beside the candidate. If the control moves, your pipeline or provider moved, and the comparison is void.
None of this needs a platform, only the discipline to hold everything fixed except the thing under test, the same idea behind our other evaluation notes: scoring RAG by slice instead of by average, checking whether a model’s confidence means what it says, and turning reviewer decisions into the evaluation set that gates releases.
If your LLM regression testing finds a real change, you will have evidence instead of a thread. If it finds nothing, you still get the number that settles most arguments early: how big a change would have to be before anyone should believe it.
Honest limits
This was one small open model (Llama 3.1 8B Instruct), one task type (150 GSM8K math problems with deterministic answers), one provider and one day. Frontier models on your tasks will have their own noise level, and open-ended tasks graded by a model add the judge’s noise on top. The method transfers; the numbers do not.
The eval also has a floor. It cannot reliably detect a change smaller than its minimum detectable effect, so a clean result only means no change was large enough to see reliably. It says nothing about whether any specific vendor changed any specific model; answering that would take the same design run over weeks against the vendor’s endpoint, with a control arm beside it.
Frequently asked questions
What is an A/A test for LLMs?
An A/A test runs the same model, prompts and settings against themselves, so any difference in scores is noise by construction. It measures how much your eval moves when nothing changes, which tells you how large a real change must be before you can trust it.
How can I tell if a model really got worse?
Run a frozen task set several times on the old and new setup, compare per-item scores with a paired test, and check the gap against your eval's minimum detectable effect. In our A/A test, a single 150-item run could not reliably detect a change smaller than 8.15 points.
Does temperature 0 make LLM outputs deterministic?
No. In our test, five temperature-0 runs still differed by up to 0.67 points and one problem changed its answer. Thinking Machines Lab names batch size, which varies with server load, as the primary reason nearly all inference endpoints are nondeterministic.
How many test cases do I need for LLM regression testing?
It depends on the change you need to see and on your model's noise. With the per-item variance we measured, catching a 3-point change 80% of the time takes about 220 items run five times each, or about 1,100 items run once, so measure your own variance with an A/A test first.
Do pinned model versions ever change?
Anthropic says it does not update the weights or configuration of an existing model ID, and also that the serving infrastructure around a model, such as the request router, safety classifiers and sampling logic, can change over time. Its September 2025 postmortem traced user-reported quality problems to infrastructure bugs.
What is a minimum detectable effect in an eval?
It is the smallest true change an eval will catch 80% of the time at 5% significance, given its number of items, its repeats and its noise. For our 150-item eval with one run per version it was 8.15 points.
Sources
- How is ChatGPT's behavior changing over time? (opens in a new tab) The 2023 drift claim: GPT-4 prime-number accuracy fell from 84% (March) to 51% (June).
- Is GPT-4 getting worse over time? (opens in a new tab) The critique: capability and behaviour on a particular prompt are different things.
- Model IDs and versions (opens in a new tab) Weights fixed per model ID; the serving infrastructure can change. The fixed-snapshot guarantee covers 4.6-and-later dateless IDs.
- GPT-5 model documentation (opens in a new tab) Snapshots lock in a specific model version.
- Gemini models (opens in a new tab) Stable models usually don't change.
- A postmortem of three recent issues (opens in a new tab) 16% is the share of Sonnet 4 requests misrouted at the worst hour, not the share degraded overall.
- Effort (opens in a new tab) Claude Opus 5.5 defaults to medium effort; run an effort sweep on your own evals.
- Refusals and fallback (opens in a new tab) Classifier refusals on the listed models return a normal response with stop_reason refusal.
- How to make your completions outputs reproducible with the new seed parameter (opens in a new tab) Determinism is not guaranteed; monitor system_fingerprint.
- Silent Updates: Measuring and Closing the Post-Deployment Disclosure Gap (opens in a new tab) Sample: nine first-party API providers and seven third-party hosts.
- GSM8K (grade-school-math) (opens in a new tab) Test split, MIT licence; used for our A/A test.
- Llama 3.1 8B Instruct on OpenRouter (opens in a new tab) Model used, pinned to Groq.
- On Randomness in Agentic Evals (opens in a new tab) SWE-Bench-Verified, 60,000 trajectories.
- Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference (opens in a new tab) One 7B model, bf16, greedy decoding.
- Defeating Nondeterminism in LLM Inference (opens in a new tab) Qwen3-235B, 1,000 samples at temperature 0.
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (opens in a new tab)
- Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints (opens in a new tab) Applies to small benchmarks, under a few hundred items.
- A statistical approach to model evals (opens in a new tab) Resampling recommended for chain-of-thought evals.
- Livenerf (opens in a new tab) Live repository, read on 29 September 2026 with 6 of 30 days collected; readings are interim.
- Inspect: options reference (opens in a new tab) Repeats per sample default to 1.