Scaling Laws for
Neural Language Models
How model size, data, and compute shape what a language model can learn.
A companion to Kaplan et al. (2020), with the later Chinchilla results kept in view. Start with next-token prediction, build the mathematics, then learn to make sense of a training budget.
Contents
What does it mean to get better?
The paper studies an unusually specific question: how does a language model's average prediction loss change as we scale the resources used to train it? Its result concerns a measurable training objective. A broad claim about intelligence would need different evidence.
An autoregressive language model receives a sequence of tokens and assigns a probability to each possible next token. Tokens can be words, pieces of words, punctuation, or other text units. For a particular observed token, the model's loss is its surprise: the negative logarithm of the probability it assigned to that token.
If the correct continuation of "The sky is ..." is "blue", a probability of 0.5 gives a loss of about 0.693. A probability of 0.1 gives about 2.303. Confident, correct predictions cost less.
One correct token, different predictions
The paper measures average loss on held-out text
Here, is the number of evaluated tokens, and is the model's learned parameters. Held-out text helps distinguish generalization from memorizing the training set. A lower loss means better probabilities on that evaluation distribution.
Perplexity is another way to express the same quantity: . A loss of 3 gives a perplexity of about 20.1; a loss of 2.7 gives about 14.9. A 0.3-nat reduction therefore lowers perplexity by roughly 26%. It does not mean accuracy improved by 26%.
Keep these four quantities separate
| Symbol | Meaning in this guide | What to watch for |
|---|---|---|
| Model size | Kaplan counts non-embedding parameters. | |
| Size of the available training dataset, in tokens | Reading the same dataset twice does not double the amount of distinct data. | |
| Total tokens processed during training | Includes repeats. For one pass through a dataset, . | |
| Training computation, in floating-point operations (FLOPs) | Measures arithmetic, not dollars, elapsed time, or the number of GPUs. |
Equal multipliers. Predictable returns.
A power law has the form , where sets the scale and controls how quickly loss falls. In the paper, can represent model size, dataset size, or an appropriately allocated compute budget. The other resources must not be the bottleneck.
This ratio removes the unknown constant . Under the fitted law, every doubling gives the same percentage reduction in loss. Because the starting loss is getting smaller, each doubling gives a smaller absolute reduction. This distinction is the precise meaning of diminishing returns here.
What does a tenfold increase buy?
For the parameter-limited exponent :
Loss falls by about 16.1%. If the baseline loss were 3.0 nats, the fitted law would predict about 2.52 nats, assuming sufficient data and training. This is a calculation from a fitted relationship, not a guaranteed outcome for an arbitrary model.
Scaling by a factor of 10
Why the line becomes straight
Taking logarithms gives a linear equation:
On a log-log plot, moving one equally spaced interval to the right means multiplying the resource by the same factor. A straight descending line means equal resource multipliers produce equal loss multipliers. On linear axes, the same relationship appears curved.
What if there is a loss floor?
A more general relationship is . Then the power law applies to the reducible part, , and total loss bends toward the floor on log-log axes. Kaplan's headline fits do not include this additive constant; the paper discusses eventual breakdown and possible floors in section 6.3. Chinchilla's later parametric model includes one explicitly.
The floor depends on the prediction task and data distribution. A fitted constant is not a known universal lower bound for language.
Three laws, three sets of conditions.
Kaplan et al. trained Transformer language models across a wide range of sizes and training configurations, primarily on WebText2. The study varied model size, data availability, and training compute to identify regimes in which one resource limits performance.
The following exponents summarize the paper's headline fits. They are empirical measurements for the studied setup, not universal constants of machine learning. [1, §1.2]
Limited by model size
Enough data and training to reveal the capacity limit.
Limited by dataset size
A large enough model, with training stopped appropriately.
Limited by optimal compute
Best allocation of model size and training effort.
The normalization constants in the first two headline fits are parameters and tokens. They locate the fitted lines. They are not recommended model or dataset sizes.
Read these figures with a question in mind
Is the trend a power law? Look for straight segments on log-log axes. Ask what was held sufficiently large to expose each limiting factor.
How do the resources interact? The left panel models data and parameter limits together; the right panel follows training progress.
Where does the curve flatten? At fixed dataset size, more parameters eventually give little additional benefit.
What was measured, and what was adjusted? Compare the empirical compute trend with the batch-adjusted estimate.
Changes to width, depth, and related architecture choices often mattered less than non-embedding parameter count within the studied family. Extreme shapes and other model families need separate analysis. The result does not establish that architecture never matters.
A bigger model still needs evidence.
Single-axis scaling laws describe useful limits. Real training runs sit between those limits. Kaplan combines finite model size and finite data in this fitted expression for early-stopped test loss:
Read the two terms inside the brackets as two constraints. When one is much larger, reducing the smaller term changes the total only slightly.
The model becomes the limit
As , the data term disappears and the equation reduces to . More data offers little relief at this capacity limit.
The dataset becomes the limit
As , the parameter term disappears and the equation reduces to . More parameters offer little relief at this data limit.
The joint fit is calibrated separately: table 2 gives , , , and . These differ from the headline single-axis fits. Keep parameter sets together when calculating. [1, §4]
Dataset size and training duration are different limitations
A model can have access to a huge dataset but process only a small portion before the compute budget runs out. Conversely, it can process many tokens by repeating a small dataset. The first case is limited by training duration; the second may be limited by data diversity and generalization.
Kaplan also fits loss as a function of model size and effective training steps:
The first term is a model-size limit. The second shrinks with more effective optimization steps. adjusts for batch-size effects; it is not simply the raw update count. This learning-curve analysis, along with batch-size scaling, underpins the paper's compute allocation result.
How should a fixed budget be spent?
For the dense Transformer regime considered here, a useful estimate is:
is tokens per batch and is update steps, so . Roughly operations per token account for the forward pass, with about twice that for the backward pass. The approximation omits some embedding and attention costs; long contexts or different architectures can change the accounting.
The same budget can train different models
Take FLOPs. Then , allowing all three combinations below:
| Parameters | Tokens processed | Approximate compute |
|---|---|---|
| 0.5 billion | 200 billion | FLOPs |
| 1 billion | 100 billion | FLOPs |
| 5 billion | 20 billion | FLOPs |
The compute equation cannot tell you which wins. A fitted loss model or actual training experiments supply that missing information.
Kaplan's answer: favor larger models, stop earlier
Kaplan's analysis recommends increasing model size faster than the tokens consumed as the budget grows, and stopping well before the chosen model converges:
With fresh data and comparable batch-efficiency assumptions, 100× more compute means roughly 28.8× more parameters and 3.5× more training tokens. Their product is 100. These are scaling multipliers; they do not specify an absolute model size without fitted constants.
The surprising part is early stopping. A larger model can reach a desired loss in fewer processed tokens. At a fixed budget, the partially trained larger model can outperform a smaller model trained closer to its own limit. Converging every model is therefore a different objective from obtaining the lowest loss for a given compute budget.
Undertraining, overtraining, and overfitting
Undertrained relative to a compute optimum: the model is too large for the tokens allocated to it; a different size/token choice could achieve better loss with the same training budget.
Trained beyond a training-compute optimum: the model is trained on more tokens than the allocation rule would choose. Loss can still improve. Smaller models trained longer may be attractive when inference cost or deployment memory matters.
Overfitting: training loss improves while generalization to held-out data deteriorates. This is a different phenomenon from training beyond a compute-optimal token count.
A model can be short of convergence and still be the best choice for its budget. "Not converged" alone does not establish inefficient training.
The pattern held. The allocation changed.
Hoffmann et al. (2022), often called the Chinchilla paper, revisited compute-optimal training with more than 400 models and varied both model size and training duration. Its headline recommendation was to scale parameters and training tokens in roughly equal proportions. [2]
More of the growth goes to size
100× compute
28.8× parameters · 3.5× tokens
Grow size and tokens together
100× compute
~10× parameters · ~10× tokens
Parameter exponent Token exponent. Bar widths describe growth exponents, not percentages of FLOPs spent on separate activities. Each model-token operation contributes to the same product .
The authors trained Chinchilla with 70 billion parameters and 1.4 trillion tokens. It outperformed the 280-billion-parameter Gopher on a broad range of evaluations at approximately the same training-compute budget. This was an empirical test of the revised allocation, not just an extrapolated curve.
Why did the answer move?
Compute-optimal estimates depend on the experiments and fitting procedure. Chinchilla sampled model sizes and training lengths more directly and emphasized matching the learning-rate schedule to the intended duration. Estimating a short run from an intermediate checkpoint of a long run can give a different result from training with a schedule designed for that short run. Its three approaches also produced somewhat different exponents.
The often repeated "20 tokens per parameter" is a useful historical rule of thumb around some Chinchilla recommendations. It is not a universal constant, and the paper's parametric fit does not imply a perfectly fixed ratio at every scale.
Find the bottom of the curve
Chinchilla's parametric fit: . Parameters and tokens are raw counts; compute uses . [2, App. D.2]
At the fitted optimum, the marginal benefits of parameters and tokens balance.
Derive the optimum with one derivative
Start with the general form . Under a fixed budget, substitute :
The parameter term decreases with ; the token term increases with . At the minimum, the derivative with respect to is zero:
Solving gives:
Therefore and . When , both grow as . Using 0.34 and 0.28 gives approximately 0.45 for parameters and 0.55 for tokens. That explains why the lab's exact fit differs slightly from the headline 50/50 simplification.
Notice that the optimum equates weighted marginal improvements. The two loss contributions themselves need not be equal.
What the curves can tell you.
Scaling laws turn a large training run into a more informed bet. Small runs can help estimate a trend, reveal bottlenecks, and guide a budget. Their value comes from calibrated prediction within a regime, together with tests of how far the regime extends.
Loss is an average, not a capability inventory
A smooth improvement in next-token prediction does not imply that each benchmark or skill improves smoothly. Thresholded evaluations can look abrupt even when underlying performance changes gradually.
Quality and distribution matter
A billion duplicated, irrelevant, or low-quality tokens need not act like a billion useful new tokens. Changes to tokenizer, context length, and evaluation corpus can also change loss and fitted coefficients. Compare losses only when the measurement setup is compatible.
Extrapolation is an additional claim
A straight line over the measured range does not prove indefinite scaling. Kaplan explicitly discusses contradictions between extrapolated trends in section 6.3. A floor or a new bottleneck can bend the curve.
The optimum belongs to an objective
Minimum pretraining loss at fixed training FLOPs differs from minimum lifetime cost. Inference volume, latency, memory, data limits, and post-training can favor a different allocation.
For every scaling claim, ask: Which metric, on which data, with which architecture, under which budget?
Turn recognition into understanding.
Work through the calculation before opening its solution. Each one tests a different piece of the argument.
Double the model
Using , what is the loss multiplier after doubling ? What happens after doubling it again?
Solution
: about a 5.1% reduction each time. After two doublings, the multiplier is . The percentage improvement per doubling stays constant; the absolute improvement shrinks.
Make the target expensive
Under the same pure parameter law, how much larger must a model be to reduce loss by 10%?
Solution
Solve . Then . Approximately four times the parameters are required, with sufficient training and data. Small exponents make large improvements expensive.
Keep compute fixed
A model processes 100 billion tokens. Its replacement has four times as many parameters. How many tokens can the replacement process with the same approximate training FLOPs?
Solution
Since , the token budget falls by a factor of four: 25 billion tokens. The compute constraint alone cannot determine which model has lower loss.
Separate exposure from new evidence
A training dataset contains 20 billion tokens and is processed for five epochs. What are and ? Is this equivalent to 100 billion distinct training tokens?
Solution
billion and billion. Repetition uses compute and may improve learning, but it does not supply the same amount of new evidence as a larger, diverse dataset.
Connect loss and perplexity
Loss falls from 3.0 to 2.8 nats. What is the perplexity multiplier? Does that give a change in benchmark accuracy?
Solution
The multiplier is , about an 18.1% reduction. Benchmark accuracy cannot be inferred from this number alone.
Five quick decisions
"In the studied regime, held-out language-model loss follows approximate power laws in model size, data, and well-allocated compute. The best model for a budget depends on how size and training tokens are balanced. Later experiments changed that balance without discarding the idea of predictable scaling."
A more productive second reading.
- Abstract + section 1.2. Translate the three headline equations into words. State the conditions for each one.
- Figures 1, 4, and 9 + section 4. Identify a power-law segment, a saturation region, and the effect of a second bottleneck.
- Sections 5-6 + Figure 13. Track the distinction between training steps, critical batch size, raw compute, and adjusted compute.
- Section 6.3 + Appendix C. Read the authors' limits on extrapolation and experimental caveats.
- Chinchilla sections 3-5 + Appendix D. Compare allocation methods and the large-model validation. Keep the approximate 50/50 headline separate from the exact parametric fit.
A compact glossary
- Cross-entropy
- Average negative log probability assigned to the observed tokens.
- Perplexity
- for loss measured with natural logarithms.
- Epoch
- One pass through a dataset.
- Power-law exponent
- The negative slope of log loss versus log resource in a pure power law.
- Compute frontier
- The lowest loss attainable at each compute budget among the configurations considered.
- IsoFLOP comparison
- A comparison of training runs that use the same compute budget.
- Critical batch size
- A scale marking the tradeoff between parallel speed and efficient use of compute.
- Loss floor
- An asymptotic loss term in a model of a particular prediction setup.
Sources and provenance
arXiv:2001.08361. Headline exponents: equations 1.1-1.3. Joint data/size fit: equation 1.5 and table 2. Training-time law: equation 1.6. Compute allocation: section 6 and Appendix A. Extrapolation limits: section 6.3.
arXiv:2203.15556. Three estimation methods: section 3. Chinchilla validation: sections 4-5. Parametric fit used in the budget example: Appendix D.2. Learning-rate schedule: Appendix B.
This guide is an original explanation of the papers. All embedded plots are calculated illustrations, not copied figures or recovered experimental data. Numerical examples use the formulas identified next to them. The document and its interactive examples work offline; source links open online.
Typesetting and icon credits
Equations are typeset with KaTeX. Interface icons are from Lucide. Both are included under their respective licenses.
KATEX The MIT License (MIT) Copyright (c) 2013-2020 Khan Academy and other contributors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. LUCIDE ISC License Copyright (c) 2026 Lucide Icons and Contributors Permission to use, copy, modify, and/or distribute this software for any purpose with or without fee is hereby granted, provided that the above copyright notice and this permission notice appear in all copies. THE SOFTWARE IS PROVIDED "AS IS" AND THE AUTHOR DISCLAIMS ALL WARRANTIES WITH REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR ANY SPECIAL, DIRECT, INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY DAMAGES WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS, WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISING OUT OF OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THIS SOFTWARE. --- The following Lucide icons are derived from the Feather project: airplay, alert-circle, alert-octagon, alert-triangle, aperture, arrow-down-circle, arrow-down-left, arrow-down-right, arrow-down, arrow-left-circle, arrow-left, arrow-right-circle, arrow-right, arrow-up-circle, arrow-up-left, arrow-up-right, arrow-up, at-sign, calendar, cast, check, chevron-down, chevron-left, chevron-right, chevron-up, chevrons-down, chevrons-left, chevrons-right, chevrons-up, circle, clipboard, clock, code, columns, command, compass, corner-down-left, corner-down-right, corner-left-down, corner-left-up, corner-right-down, corner-right-up, corner-up-left, corner-up-right, crosshair, database, divide-circle, divide-square, dollar-sign, download, external-link, feather, frown, hash, headphones, help-circle, info, italic, key, layout, life-buoy, link-2, link, loader, lock, log-in, log-out, maximize, meh, minimize, minimize-2, minus-circle, minus-square, minus, monitor, moon, more-horizontal, more-vertical, move, music, navigation-2, navigation, octagon, pause-circle, percent, plus-circle, plus-square, plus, power, radio, rss, search, server, share, shopping-bag, sidebar, smartphone, smile, square, table-2, tablet, target, terminal, trash-2, trash, triangle, tv, type, upload, x-circle, x-octagon, x-square, x, zoom-in, zoom-out The MIT License (MIT) (for the icons listed above) Copyright (c) 2013-present Cole Bemis Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.