Is δ small? Measuring the premise of a hallucination lower bound on OLMo-2

On the birthday example of Kalai et al., OLMo-2 7B gives δ = +0.028, inside a ±0.05 band set in advance. The result is fragile, and the floor itself is vacuous on this frame.

DOI 10.5281/zenodo.23225543 · PDF · data and code · discussion on alphaXiv · CC BY 4.0 · preprint, 8 Oct 2026 · ORCID 0009-0005-0419-4070
Kalai, Nachum, Vempala and Zhang argue that pretraining keeps the slack term δ of their hallucination bound small, but do not measure it. For OLMo-2 7B base, δ = +0.028 (95% CI [0.017, 0.042]). That run was added after a feasibility probe had scored the people who carry 83% of the weight, and weighting every frame person equally gives +0.156.

Abstract

Kalai, Nachum, Vempala and Zhang prove that a pretrained language model errs on arbitrary facts at least about as often as such facts appear once in training, up to a slack term δ. The term compares the probability the model puts on answers above a threshold with how often the true answer is above it. They argue that pretraining keeps δ small, but do not measure it.

We measure its signed form on their birthday example. OLMo-2's corpus is public, so we weight each of 2,083 people by how often their name and birth date occur together in it. For the 7B base model δ = +0.028 (95% CI [0.017, 0.042]), inside a ±0.05 band set before the probe was built. For the 1B it is inconclusive.

The 7B run was added after part of its data had been seen, and weighting every frame person equally gives +0.156. On this frame the floor itself is vacuous.

The measurement

Pipeline from Wikidata people to the frame, training-corpus counts, strata, scored people, log-probability matrices and the statistic.
The measurement, from Wikidata people to δs. Each box is one directory of the dataset.

The frame is 3,992 Wikidata humans with a day-precise date of birth. For each we count matches of the name within 100 tokens of the date in the OLMo-2 stage-1 corpus, split the frame into six strata by that count, and score 350 per stratum. The model scores all 366 month–day strings, so no sampling is needed. The primary prompt T1 is:

Birthdays (month and day):
Albert Einstein: March 14
Marie Curie: November 7
Abraham Lincoln: February 12
{name}:

Results

modelstatuspromptδs95% CIaccuracy
OLMo-2 1B base, fp32pre-specifiedT1+0.028[−0.001, 0.061]7.2%
T2+0.077[0.021, 0.139]
T3+0.097[0.061, 0.137]
OLMo-2 7B base, Q8_0amendmentT1+0.028[0.017, 0.042]64.6%
T2−0.012[−0.029, 0.006]
T3+0.012[0.001, 0.024]

Count-weighted δs. Accuracy: top-ranked date correct, people with nc ≥ 100, unweighted. The 1B's interval crosses +0.05, so its verdict is inconclusive.

How firm the 7B result is

Delta for the 7B under T1 with 95% intervals: the pre-specified weight at +0.028 and six alternatives, two near +0.05 and frame-uniform weights at +0.156.
δs for the 7B base model under T1, with 95% intervals: the pre-specified weight (blue) and six alternatives computed afterwards. The alternatives reweight or relabel the same scores; they are not new runs. Grey band: ±0.05. Dashed line: the point threshold 0.10 of the harmful row.

Dropping the 50 heaviest people (37% of the weight), or the 63 whose counts the index could only approximate (35.5%), moves δs to about +0.05, with intervals that cross the band.

Where δ sits

By training-count stratum: delta near +0.15 below 100 mentions and near zero above; accuracy at most 7% below 100 mentions and 65% above.
OLMo-2 7B base under T1, by training-count stratum. Left: δs (blue, 95% interval) and its expected value if the true dates were assigned at random within the stratum (orange). Right: share of people whose top-ranked date is correct; dashed line, chance.

For people with nc ≥ 100 the 7B is right 65% of the time and the unweighted δs is +0.012. In each stratum below 100 it is right at most 7% of the time and δs is +0.14 to +0.19.

Top-ranked dates for people with no counted mention: March 14 and February 12 dominate under T1.
Top-ranked date of the 7B base model for the 350 people with nc = 0, under T1 (top) and T2 (bottom).

Post-training

OLMo-2 1B across base, SFT, DPO and Instruct: delta rises from 0.03 to 0.33; top-answer ECE from 0.015 to 0.027.
OLMo-2 1B under T1 at each post-training stage. Left: δs with 95% intervals. Right: top-answer ECE.

What it does not establish

The note is AI-assisted; its disclosure section says how. It has not been peer reviewed.