A Corrected Metric Does Not Validate a Vocabulary Scaling Law

A re-analysis of the compute-optimal vocabulary law of Tao et al. (2024) on their own 60 released training runs: what changes when the metric is fixed, and whether either version picks the right vocabulary at a model size it has not seen.

DOI 10.5281/zenodo.23205402 · PDF · code and results · alphaXiv · CC BY 4.0 (paper), MIT (code) · preprint, 7 Oct 2026 · ORCID 0009-0005-0419-4070
Re-expressing every checkpoint in bits per character moves the 70B prescription from 212K to 106K–119K or 550K–764K, depending on which of Tao et al.'s two fitting approaches you use. With a whole model size held out, every procedure misses the best vocabulary at 1.13B by 8–40% compute-equivalent, and the corrected metric does not reduce the miss.

The metric picks a winner even when there is none

The law is fitted to a unigram-normalised loss Lu. Its normaliser depends on the tokenizer, so Lu moves with vocabulary size even when every vocabulary predicts the text equally well. Give all ten released tokenizers the same bits per character, and Lu still has a minimum, which grows as the common loss falls. That is the shape of the published conclusion, produced with no real difference between vocabularies.

L_u against vocabulary size when every vocabulary has the same bits per character; the minimum moves from 4K to 10K to 24K as bits per character falls.
Lu at constant bits per character, shifted to minimum zero. The optimum (star) is 4K at 1.5 BPC and above, 10K at 1.1 and 24K at 0.85.

What the correction changes

An exact identity converts each checkpoint's Lu to bits per character from measured tokenizer statistics. Refitting with Tao et al.'s own code:

Lubits per character
Exponent of vocabulary parameters0.4160.302–0.316
Optimal vocabulary at 7B62K53–56K
At 70B, frontier approach212K106K–119K
At 70B, parametric approach222K550K–764K

On the original metric the two approaches give 212K and 222K at 70B. On the corrected one they disagree by a factor of five to six.

A held-out test

To ask whether a prescription is right, hide a whole model size, fit on the rest, and let each procedure pick a vocabulary at eight compute budgets for the hidden size. Score each pick against the best measured vocabulary, as the share of compute it wastes.

Regret at eight held-out budgets for five procedures at 1.13B in two designs; all procedures miss, and the corrected-metric procedures are not consistently better.
Regret at each held-out budget of the 1.13B model. The grey area is the regret of the worst vocabulary. Arm E was added after the pre-specified analysis.
ProcedureH1, 1.13BH2, 632MH2, 1.13B
L: parametric on Lu (Tao)14.6%20.8%12.9%
C: parametric on bits per character22.9%15.0%40.1%
D: frontier on bits per character12.5%5.0%22.9%
E: frontier on Lu (post hoc)8.1%13.5%9.0%
Worst vocabulary47.2%41.3%47.2%

Median regret over eight budgets. With the fitting method held fixed, Lu beats bits per character in four of six comparisons.

A simpler rule

On DataDecide and on FAIR's released granularity grids, keeping every option within noise of the best at the largest measured scale contains the eventual winner, or misses it by at most 3.2 × 10−4 bits per byte. On Tao et al.'s release the same rule covers seven of eight budgets in one design and none in the other.

What it does not establish

The paper is AI-assisted; its first-page footnote says how. Code, derived inputs, pre-written analysis rules and result tables are in the Zenodo supplement; third-party data are referenced by pinned source.