Long documents make your batch smaller than it looks

Chunks cut from one long document pull the gradient in the same direction. I measured how much that shrinks the effective batch, using 512-token chunks of real documents on open model checkpoints. Grouped 32 to a document, chunks of code repositories hold only 20–28% as many effective independent samples as the same number of tokens of packed web text.

9 Oct 2026 · Hanyu Yang · code and data · preregistration · doi:10.5281/zenodo.23270769

Two ways to fill a batch

Pretraining usually packs many short web documents into each sequence. Long-context training does something different: it fills a 16K or 64K sequence with one long document. That could be a book, a code repository glued together file by file, or a paper.

Two rows of four 16K sequences. Top: each sequence is split into many short grey documents. Bottom: each sequence is a single green document, cut into 32 chunks.
Same token count, very different contents. On top, each batch row mixes many unrelated web documents. Below, each row is one document.

Both batches have the same number of tokens, so on paper they are the same batch size. But the 32 chunks of one repository are not 32 independent samples. They share variable names, style, topic and imports. If their gradients point the same way, averaging them cancels less noise than averaging 32 unrelated chunks would.

The idea in one formula

Survey statisticians have a name for this: the design effect. If you sample people in clusters (whole households, whole schools) instead of one by one, your estimate is noisier than the sample size suggests. For a block of m chunks whose noise is correlated, the variance of their average is

Var(average of m chunks) = (σ² / m) × R,    with   R = 1 + (2/m) Σpairs ρ

Here σ² is the noise of one chunk's gradient and ρ is the correlation between two chunks' noise. If the chunks were independent, R would be 1. If R = 4, your batch carries as much noise as one with a quarter of the samples. In the language of McCandlish et al., R is how much the noise scale Bsimple grows when you swap shuffled chunks for whole-document blocks of the same corpus.

How I measured it

The plan was frozen before any gradient was computed, and it is public. One rule was added later, after part of the data had been seen; that is explained under What went wrong. The setup:

modelsPythia-410M (4 checkpoints), Pythia-1.4B (2), OLMo-2-1B (2)
corporacode repositories (The Stack, concatenated by repo), arXiv papers, PG-19 books, and packed FineWeb as the control
sample16 documents per corpus, the same 16 for every model; from each, the span from token 2,048 to 18,432, cut into 32 chunks of 512 tokens
per cell512 chunk gradients and their inner products (a Gram matrix)
coverage32 planned cells: 24 complete, 5 partial, 3 missing; every final checkpoint is complete
hardwareCPU only, 4-core cloud machines; no GPU
Pipeline: 16K-token document span, 32 chunks of 512 tokens, full gradient per chunk, count sketch, Gram matrix of 512 chunks, then rho, R and B_simple.
From documents to numbers. Each chunk runs as its own 512-token sequence, so any correlation comes from the content and not from shared attention context.

That last point cuts both ways. It keeps the measurement clean, but it also means this is not the gradient of a real 16K training sequence, where each chunk sees everything before it. The numbers here describe batches of 512-token chunks grouped by document. They are a proxy for long-context training, and the real multiplier could be larger or smaller.

A full gradient has hundreds of millions of numbers, and storing 512 of them per cell was not an option on CPUs. So each gradient is compressed with a count sketch: every parameter is hashed into one of 131,072 buckets with a random sign, separately for attention, MLP, embedding and other parameters. Inner products survive this almost exactly. Exact squared norms are kept on the side.

From the inner products, three things fall out. Pairs from different documents give the squared length of the mean gradient. The diagonal, minus that, gives the noise variance of one chunk. Pairs from the same document at distance Δ give the correlation ρ(Δ). For the two 1B models, documents were processed in groups of four, so only pairs inside a group exist: 24 of the 120 possible document pairs.

Result 1: the correlation reaches far

Gradient correlation against chunk distance from 0.5K to 16K tokens. Code starts near 0.27 (Pythia) and is still above 0.1 at 16K; arXiv starts near 0.13 and falls to about 0.04 to 0.06; books stay near 0.03; web drops to zero by about 4K.
Correlation between two chunks of the same document, against how far apart they are, at the final checkpoint. The longest distance is 15,872 tokens. The shaded band is the 2K–16K range the preregistration used.

Two chunks of the same repository, about 16K tokens apart, still have a noise correlation of about 0.1: 0.14 on Pythia-1.4B and 0.09 on OLMo-2. ArXiv falls from 0.13 at 512 tokens to 0.07–0.08 at 2K and 0.04–0.06 at the longest distance. Books are lower but clearly positive. The web control behaves as you would expect from packing: some correlation at short range, where one web document spans neighbouring chunks, and essentially zero beyond about 4K.

A correlation of 0.1 sounds small. It is not, because the design effect adds it up over every pair, and a 16K block has 496 pairs of chunks. These are averages over pairs and documents; individual pairs vary a lot.

Result 2: the noise grows with block length

R(W) against block length from 0.5K to 16K. At 16K, code reaches 6.9 (Pythia-1.4B) and 4.7 (OLMo-2), arXiv 3.3 and 3.0, books 2.2 and 1.75, web 1.4 and 1.35.
R(W): the noise of one contiguous W-token block, relative to the same tokens drawn as independent chunks. It is computed directly from the block sums, with no model of ρ assumed.

The longer the block, the more correlated pairs it contains, and the bigger the penalty. At 16K tokens a code block carries 4.7–6.9× the noise of independent chunks of the same code.

Result 3: against the control

The packed-web control is not at 1 either. Short web documents straddle chunk boundaries, which gives it R = 1.3–1.4. The fair comparison is therefore relative: each corpus divided by packed web at the same model and checkpoint, Rrel = Rcorpus / Rweb. Each R is measured against shuffled chunks of its own corpus, so Rrel compares how much each layout shrinks the sample count. It does not compare absolute noise levels between corpora; that would also need the ratio of their per-chunk noise variances, which is not part of this comparison.

Bars of R_rel at 16K with 95% intervals: code 4.91 and 3.51, arXiv 2.36 and 2.23, books 1.56 and 1.30, for Pythia-1.4B and OLMo-2-1B.
Design effect of whole-document 16K blocks divided by that of packed web, final checkpoints, with 95% intervals from resampling documents.
corpusPythia-1.4BOLMo-2-1Braw R(16K)
code repositories4.91 [2.61, 7.78]3.51 [2.37, 5.23]6.91 / 4.73
arXiv papers2.36 [1.91, 2.96]2.23 [1.78, 2.92]3.32 / 3.01
PG-19 books1.56 [1.26, 1.98]1.30 [1.07, 1.67]2.20 / 1.75
packed web111.41 / 1.35

The two models disagree on the exact size but agree on the order: code, then arXiv, then books. The intervals come from resampling the 16 documents with replacement, so they reflect only the documents I drew. Books sit near 1.5, the threshold the plan set in advance for raw R and the addendum later applied to Rrel.

The four corpora side by side

loss (nats)ρ at 512 tokensρ, 2K–16K averageR(16K)per-document share: median / max
Pythia-1.4B, final checkpoint
code repositories1.140.270.166.910.07 / 1.06
arXiv papers1.710.130.063.320.06 / 0.10
PG-19 books2.760.060.032.200.03 / 0.07
packed web2.930.070.001.410.00 / 0.05
OLMo-2-1B, final checkpoint
code repositories1.420.190.104.730.07 / 0.26
arXiv papers1.780.130.053.010.05 / 0.09
PG-19 books2.870.040.021.750.02 / 0.03
packed web2.880.050.001.35−0.01 / 0.06

Note the web row: its short-range correlation (0.05–0.07 at 512 tokens) is as high as the books', but it vanishes at long range. The books are the reverse: weak, but it never goes away. The last column is each document's long-range term, scaled by the variance of the whole cell rather than its own, so it is not a correlation and can exceed 1.

What it means for your batch size

Turn the ratio around and you get an effective batch size: how many tokens of packed web would hold the same number of effective independent samples.

Horizontal bars: effective batch as a percentage of the same tokens of packed web. Code 20% and 28%, arXiv 42% and 45%, books 64% and 77%.
Effective independent samples in whole-document 16K blocks, as a share of those in the same number of tokens of packed web text (1 / Rrel).
1M tokens per step of…as many effective samples as (Pythia-1.4B)(OLMo-2-1B)
whole code repositories204K web tokens [129K, 382K]285K web tokens [191K, 421K]
whole arXiv papers424K [337K, 523K]448K [342K, 562K]
whole books639K [504K, 793K]769K [597K, 932K]

The brackets are the 95% intervals carried over from Rrel; the ranges in the text are the two models' point estimates. This matters because long-context data mixes lean heavily on code repositories. ProLong's long-document data, for instance, is mostly repositories and books. If you size a whole-document batch by its token count, as you would packed data, it may hold several times fewer independent samples than that count suggests.

Three cautions. These numbers count samples; they do not say that 1M tokens of code and 204K tokens of web have the same gradient noise in absolute terms, because a single code chunk and a single web chunk have different noise variance and different mean gradients. They come from 512-token chunks without long context, so they are a proxy for 16K training. And they do not give a learning-rate rule: Bsimple ignores curvature, and a noisier gradient does not by itself mean a worse model.

It is not only an early-training effect

R at 16K across Pythia-410M checkpoints at 0.7%, 5%, 25% and 100% of training. Code falls from 7.9 to 4.9, arXiv from 3.2 to 2.5, books around 1.6 to 2, web around 1.3 to 1.6.
R(16K) across four Pythia-410M checkpoints, with 95% bands.

One worry was that this only shows up very early, when the model has not learned much and every gradient looks alike. On Pythia-410M it does not. The code effect shrinks over training, from about 8 to about 5, and arXiv falls by about a fifth, from 3.2 to 2.5, but the ordering holds at every checkpoint. For the two 1B models only the final checkpoints are complete, so this check rests on the smaller model.

It is lumpy

Per-document long-range share, one dot per document. Code is widely spread from 0.01 to 1; arXiv is tight around 0.06; books around 0.02 to 0.03; web scattered around zero.
Each document's long-range share (its average long-lag term, scaled by the variance of the whole cell), final checkpoint, with medians. The vertical axis is linear near zero and logarithmic above 0.05.

ArXiv papers are remarkably uniform: every one of the 16 sits near the median. Code repositories are not. A few repositories contribute a lot, and some barely anything. This is why the code intervals are so wide, and it suggests the effect depends a lot on which repositories you train on. A sixteen-repository sample is enough to see it, not enough to pin it down, and it says nothing about kinds of repositories it did not include.

It shows up in every parameter group

R at 16K by parameter group for Pythia-1.4B: attention, MLP, embeddings and other parameters all show the same corpus ordering, code highest.
R(16K) computed separately for each parameter group, Pythia-1.4B final checkpoint. No intervals were computed per group.

The penalty is not driven by one part of the model, say the embedding rows of rare tokens that keep recurring in one repository. Attention, MLP, embeddings and the remaining parameters all show it, in the same order. These are four large groups; I did not look at single layers or heads.

Checking the instrument

Two things could fake a result like this: a sketch that distorts inner products, or a position effect where chunks late in a sequence look different. The checks below argue against both, but they are narrower than a full validation.

Scatter of sketched against exact cosine for 8,176 chunk pairs; all points lie on the diagonal.
Sketched against exact cosine similarity for 8,176 chunk pairs on Pythia-410M: the first gradient of each of 16 shards, computed exactly, against the 511 others in its shard. The error has mean −0.00007 and standard deviation 0.002.
Mean chunk loss against position in the 16K span, flat for all four corpora.
Mean loss of each chunk against where it sits in the 16K span, Pythia-1.4B final checkpoint.

The sketch check was run on the smallest model only. Losses are flat across the span, which rules out the crudest position effect, though not every way a gradient could depend on position. They also show the familiar ordering: code is the easiest text for these models, at about 1.1 nats per token, and web the hardest, at about 2.9. Two more gates are worth naming. The web control's long-range correlation passes the plan's 0.01 limit on its point estimate, but its intervals reach 0.010–0.012. And one gate fails: for Pythia-1.4B books, the squared mean gradient is too uncertain (its interval includes zero), which makes Bsimple unreliable there. R does not divide by it.

Try it yourself

Pick a corpus, a model and the tokens per step you train with. Then flip through the curves behind the numbers above.

   

 

What went wrong, and what I changed

Not everything went to plan, and the details are in the repository.

What this does not show

Long-context training can cost some short-context quality, depending on the data mix; ProLong, for one, reports both the losses and mixes that avoid them. A noisier gradient is one possible reason. Others include attention spreading too thin over long windows, a mismatch between training and test lengths, and a plain shift in domain. This measurement shows the extra noise is there for 512-token chunks grouped by document, and how big it is. It does not measure the noise of real 16K-context gradients, and it does not show how much of any quality cost the noise explains.

Testing that takes training runs. One design changes only how a fixed set of documents is cut and batched:

armwindowdocuments per stepwhat it changes
A16K, whole documents64the usual long-context setup
B1K, the same 64 documents chopped, accumulated into one update64A vs B: the window, with batch content fixed
C1K, chunks shuffled across the same pool of documents~1,000B vs C: batch composition, with the window and the token pool fixed

All three arms need the same tokens, the same number of updates and the same exposure to each document over training. If each explanation acted alone, you would expect roughly this:

explanation, acting aloneA vs B (window)B vs C (composition)
design effect (this post)about equalC better
the long window itself hurtsA worseabout equal
train/test length mismatchA worse on short evals onlyabout equal

In practice these can act together, and A and B differ in more than noise, since longer context changes the gradient itself. A gap between B and C would show that batch composition matters, not by itself that this design effect is the reason. This would take a few thousand GPU-hours at the 400M–1B scale. I have not run it.

Related work

I could not find an earlier measurement of within-document gradient correlation as a function of distance, or of the design effect it implies, on real language-model gradients. If you know of one, I would like to hear about it.

Data and code

Everything is in doc-batch-gradients: the probe, the analysis, the Gram matrices for all 57 shards, and the full report. Every number on this page can be recomputed on a laptop in about 20 minutes. The plan, the deviations and the addendum are in docs/.

How to cite

This post, the code and the analysis report are archived on Zenodo as doi:10.5281/zenodo.23270769. The Gram matrices are too large for that deposit; they stay in the GitHub repository at tag v2.0, and the deposit lists their SHA-256 hashes.

@techreport{yang2026longdocbatch,
  author = {Yang, Hanyu},
  title  = {Long documents make your batch smaller than it looks: measuring within-document gradient correlation},
  year   = {2026},
  institution = {Zenodo},
  doi    = {10.5281/zenodo.23270769},
  url    = {https://doi.org/10.5281/zenodo.23270769}
}