Liu Wei / THE LEARNING COLLECTION
PAPER COMPANION 01
A Guided Reading 2020 / 2022

Scaling Laws for
Neural Language Models

How model size, data, and compute shape what a language model can learn.

A companion to Kaplan et al. (2020), with the later Chinchilla results kept in view. Start with next-token prediction, build the mathematics, then learn to make sense of a training budget.

Conceptual + mathematicalPrerequisites: basic algebra and logarithms
MODEL SIZECapacityHow much can be represented?
TRAINING DATAEvidenceWhat can be learned from?
COMPUTEBudgetHow much learning is affordable?
Contents
01 First Principles

What does it mean to get better?

The paper studies an unusually specific question: how does a language model's average prediction loss change as we scale the resources used to train it? Its result concerns a measurable training objective. A broad claim about intelligence would need different evidence.

An autoregressive language model receives a sequence of tokens and assigns a probability to each possible next token. Tokens can be words, pieces of words, punctuation, or other text units. For a particular observed token, the model's loss is its surprise: the negative logarithm of the probability it assigned to that token.

ℓt=−ln⁡pθ(xt∣x<t)\ell_t=-\ln p_\theta(x_t\mid x_{<t})ONE TOKEN

If the correct continuation of "The sky is ..." is "blue", a probability of 0.5 gives a loss of about 0.693. A probability of 0.1 gives about 2.303. Confident, correct predictions cost less.

Probability & Surprise

One correct token, different predictions

EXAMPLE 01
50%
1% · unlikely99% · expected
0.693nats of surprise
Natural logarithms measure loss in nats. Halving the assigned probability adds ln⁡2≈0.693\ln 2\approx0.693 to the loss.

The paper measures average loss on held-out text

L=−1M∑t=1Mln⁡pθ(xt∣x<t)L=-\frac{1}{M}\sum_{t=1}^{M}\ln p_\theta(x_t\mid x_{<t})CROSS-ENTROPY

Here, MM is the number of evaluated tokens, and θ\theta is the model's learned parameters. Held-out text helps distinguish generalization from memorizing the training set. A lower loss means better probabilities on that evaluation distribution.

Perplexity is another way to express the same quantity: PPL=eL\mathrm{PPL}=e^L. A loss of 3 gives a perplexity of about 20.1; a loss of 2.7 gives about 14.9. A 0.3-nat reduction therefore lowers perplexity by roughly 26%. It does not mean accuracy improved by 26%.

Keep these four quantities separate

SymbolMeaning in this guideWhat to watch for
NNModel sizeKaplan counts non-embedding parameters.
DDSize of the available training dataset, in tokensReading the same dataset twice does not double the amount of distinct data.
TTTotal tokens processed during trainingIncludes repeats. For one pass through a dataset, T=DT=D.
CCTraining computation, in floating-point operations (FLOPs)Measures arithmetic, not dollars, elapsed time, or the number of GPUs.
02 The Power Law

Equal multipliers. Predictable returns.

A power law has the form L(x)=Ax−αL(x)=Ax^{-\alpha}, where AA sets the scale and α\alpha controls how quickly loss falls. In the paper, xx can represent model size, dataset size, or an appropriately allocated compute budget. The other resources must not be the bottleneck.

L(kx)L(x)=k−α\frac{L(kx)}{L(x)}=k^{-\alpha}THE USEFUL FORM

This ratio removes the unknown constant AA. Under the fitted law, every doubling gives the same percentage reduction in loss. Because the starting loss is getting smaller, each doubling gives a smaller absolute reduction. This distinction is the precise meaning of diminishing returns here.

Worked Example

What does a tenfold increase buy?

For the parameter-limited exponent αN=0.076\alpha_N=0.076:

L(10N)L(N)=10−0.076≈0.839\frac{L(10N)}{L(N)}=10^{-0.076}\approx0.839

Loss falls by about 16.1%. If the baseline loss were 3.0 nats, the fitted law would predict about 2.52 nats, assuming sufficient data and training. This is a calculation from a fitted relationship, not a guaranteed outcome for an arbitrary model.

Explore The Relationship

Scaling by a factor of 10

Chart axes
10×
1×10×100×1,000×
Loss / baseline0.839
Loss reduction16.1%
Fitted exponent0.076
Calculated curves from Kaplan et al., equations 1.1-1.3; normalized to L(1)=1L(1)=1. These are illustrations of the fits, not experimental datapoints. Each resource curve describes its own limiting regime.

Why the line becomes straight

Taking logarithms gives a linear equation:

ln⁡L=ln⁡A−αln⁡x\ln L=\ln A-\alpha\ln xSLOPE = −α-\alpha

On a log-log plot, moving one equally spaced interval to the right means multiplying the resource by the same factor. A straight descending line means equal resource multipliers produce equal loss multipliers. On linear axes, the same relationship appears curved.

What if there is a loss floor?

A more general relationship is L(x)=L∞+Ax−αL(x)=L_\infty+Ax^{-\alpha}. Then the power law applies to the reducible part, L−L∞L-L_\infty, and total loss bends toward the floor on log-log axes. Kaplan's headline fits do not include this additive constant; the paper discusses eventual breakdown and possible floors in section 6.3. Chinchilla's later parametric model includes one explicitly.

The floor depends on the prediction task and data distribution. A fitted constant is not a known universal lower bound for language.

03 Read The Evidence

Three laws, three sets of conditions.

Kaplan et al. trained Transformer language models across a wide range of sizes and training configurations, primarily on WebText2. The study varied model size, data availability, and training compute to identify regimes in which one resource limits performance.

The following exponents summarize the paper's headline fits. They are empirical measurements for the studied setup, not universal constants of machine learning. [1, §1.2]

N

Limited by model size

Enough data and training to reveal the capacity limit.

L(N)∝N−0.076L(N)\propto N^{-0.076}10× size → ~16.1% less loss
D

Limited by dataset size

A large enough model, with training stopped appropriately.

L(D)∝D−0.095L(D)\propto D^{-0.095}10× data → ~19.6% less loss
C

Limited by optimal compute

Best allocation of model size and training effort.

L(Cmin⁡)∝Cmin⁡−0.050L(C_{\min})\propto C_{\min}^{-0.050}10× compute → ~10.9% less loss

The normalization constants in the first two headline fits are Nc≈8.8×1013N_c\approx8.8\times10^{13} parameters and Dc≈5.4×1013D_c\approx5.4\times10^{13} tokens. They locate the fitted lines. They are not recommended model or dataset sizes.

Read these figures with a question in mind

Figure 1

Is the trend a power law? Look for straight segments on log-log axes. Ask what was held sufficiently large to expose each limiting factor.

Figure 4

How do the resources interact? The left panel models data and parameter limits together; the right panel follows training progress.

Figure 9

Where does the curve flatten? At fixed dataset size, more parameters eventually give little additional benefit.

Figure 13

What was measured, and what was adjusted? Compare the empirical compute trend with the batch-adjusted estimate.

Changes to width, depth, and related architecture choices often mattered less than non-embedding parameter count within the studied family. Extreme shapes and other model families need separate analysis. The result does not establish that architecture never matters.

04 Two Bottlenecks

A bigger model still needs evidence.

Single-axis scaling laws describe useful limits. Real training runs sit between those limits. Kaplan combines finite model size and finite data in this fitted expression for early-stopped test loss:

L(N,D)=[(NcN)αN/αD+DcD]αDL(N,D)=\left[\left(\frac{N_c}{N}\right)^{\alpha_N/\alpha_D}+\frac{D_c}{D}\right]^{\alpha_D}KAPLAN · EQ. 1.5

Read the two terms inside the brackets as two constraints. When one is much larger, reducing the smaller term changes the total only slightly.

Abundant Data

The model becomes the limit

As D→∞D\to\infty, the data term disappears and the equation reduces to L(N)=(Nc/N)αNL(N)=(N_c/N)^{\alpha_N}. More data offers little relief at this capacity limit.

Abundant Capacity

The dataset becomes the limit

As N→∞N\to\infty, the parameter term disappears and the equation reduces to L(D)=(Dc/D)αDL(D)=(D_c/D)^{\alpha_D}. More parameters offer little relief at this data limit.

The joint fit is calibrated separately: table 2 gives αN=0.076\alpha_N=0.076, αD=0.103\alpha_D=0.103, Nc=6.4×1013N_c=6.4\times10^{13}, and Dc=1.8×1013D_c=1.8\times10^{13}. These differ from the headline single-axis fits. Keep parameter sets together when calculating. [1, §4]

Dataset size and training duration are different limitations

A model can have access to a huge dataset but process only a small portion before the compute budget runs out. Conversely, it can process many tokens by repeating a small dataset. The first case is limited by training duration; the second may be limited by data diversity and generalization.

Kaplan also fits loss as a function of model size and effective training steps:

L(N,S)=(NcN)αN+(ScSmin⁡(S))αSL(N,S) = \left(\frac{N_c}{N}\right)^{\alpha_N}+\left(\frac{S_c}{S_{\min}(S)}\right)^{\alpha_S}KAPLAN · EQ. 1.6

The first term is a model-size limit. The second shrinks with more effective optimization steps. Smin⁡S_{\min} adjusts for batch-size effects; it is not simply the raw update count. This learning-curve analysis, along with batch-size scaling, underpins the paper's compute allocation result.

05 The Compute Budget

How should a fixed budget be spent?

For the dense Transformer regime considered here, a useful estimate is:

C≈6NT=6NBSC\approx6NT=6NBSTRAINING FLOPs

BB is tokens per batch and SS is update steps, so T=BST=BS. Roughly 2N2N operations per token account for the forward pass, with about twice that for the backward pass. The approximation omits some embedding and attention costs; long contexts or different architectures can change the accounting.

Worked Example

The same budget can train different models

Take C=6×1020C=6\times10^{20} FLOPs. Then NT=1020NT=10^{20}, allowing all three combinations below:

ParametersTokens processedApproximate compute
0.5 billion200 billion6×10206\times10^{20} FLOPs
1 billion100 billion6×10206\times10^{20} FLOPs
5 billion20 billion6×10206\times10^{20} FLOPs

The compute equation cannot tell you which wins. A fitted loss model or actual training experiments supply that missing information.

Kaplan's answer: favor larger models, stop earlier

Kaplan's analysis recommends increasing model size faster than the tokens consumed as the budget grows, and stopping well before the chosen model converges:

Nopt∝Cmin⁡0.73,Topt∝Cmin⁡0.27N_{\mathrm{opt}}\propto C_{\min}^{0.73},\qquad T_{\mathrm{opt}}\propto C_{\min}^{0.27}KAPLAN · ALLOCATION

With fresh data and comparable batch-efficiency assumptions, 100× more compute means roughly 28.8× more parameters and 3.5× more training tokens. Their product is 100. These are scaling multipliers; they do not specify an absolute model size without fitted constants.

The surprising part is early stopping. A larger model can reach a desired loss in fewer processed tokens. At a fixed budget, the partially trained larger model can outperform a smaller model trained closer to its own limit. Converging every model is therefore a different objective from obtaining the lowest loss for a given compute budget.

Undertraining, overtraining, and overfitting

Undertrained relative to a compute optimum: the model is too large for the tokens allocated to it; a different size/token choice could achieve better loss with the same training budget.

Trained beyond a training-compute optimum: the model is trained on more tokens than the allocation rule would choose. Loss can still improve. Smaller models trained longer may be attractive when inference cost or deployment memory matters.

Overfitting: training loss improves while generalization to held-out data deteriorates. This is a different phenomenon from training beyond a compute-optimal token count.

A model can be short of convergence and still be the best choice for its budget. "Not converged" alone does not establish inefficient training.

06 The Chinchilla Update

The pattern held. The allocation changed.

Hoffmann et al. (2022), often called the Chinchilla paper, revisited compute-optimal training with more than 400 models and varied both model size and training duration. Its headline recommendation was to scale parameters and training tokens in roughly equal proportions. [2]

Kaplan · 2020

More of the growth goes to size

0.730.27

100× compute
28.8× parameters · 3.5× tokens

Chinchilla · 2022

Grow size and tokens together

~0.50~0.50

100× compute
~10× parameters · ~10× tokens

Parameter exponent Token exponent. Bar widths describe growth exponents, not percentages of FLOPs spent on separate activities. Each model-token operation contributes to the same product NTNT.

The authors trained Chinchilla with 70 billion parameters and 1.4 trillion tokens. It outperformed the 280-billion-parameter Gopher on a broad range of evaluations at approximately the same training-compute budget. This was an empirical test of the revised allocation, not just an extrapolated curve.

Why did the answer move?

Compute-optimal estimates depend on the experiments and fitting procedure. Chinchilla sampled model sizes and training lengths more directly and emphasized matching the learning-rate schedule to the intended duration. Estimating a short run from an intermediate checkpoint of a long run can give a different result from training with a schedule designed for that short run. Its three approaches also produced somewhat different exponents.

The often repeated "20 tokens per parameter" is a useful historical rule of thumb around some Chinchilla recommendations. It is not a universal constant, and the paper's parametric fit does not imply a perfectly fixed ratio at every scale.

Explore A Fixed Budget

Find the bottom of the curve

Chinchilla's parametric fit: L=1.69+406.4N−0.34+410.7T−0.28L=1.69+406.4N^{-0.34}+410.7T^{-0.28}. Parameters and tokens are raw counts; compute uses C≈6NTC\approx6NT. [2, App. D.2]

10^21 FLOPs
10²⁰10²²10²⁴ FLOPs
1.00×
0.05×1×20×
Parameters—
Training tokens—
Predicted loss—
Above best fit—

At the fitted optimum, the marginal benefits of parameters and tokens balance.

An educational reproduction of one published fit, not a modern training recommendation. The lowest point depends on the chosen fit, data distribution, and compute approximation. It does not represent all three Chinchilla estimation methods.
Derive the optimum with one derivative

Start with the general form L=E+AN−α+BT−βL=E+AN^{-\alpha}+BT^{-\beta}. Under a fixed budget, substitute T=C/(6N)T=C/(6N):

L(N)=E+AN−α+B(6NC)βL(N)=E+AN^{-\alpha}+B\left(\frac{6N}{C}\right)^\beta

The parameter term decreases with NN; the token term increases with NN. At the minimum, the derivative with respect to ln⁡N\ln N is zero:

αAN−α=βB(6NC)β\alpha AN^{-\alpha}=\beta B\left(\frac{6N}{C}\right)^\beta

Solving gives:

N∗=(αAβB)1/(α+β)(C6)β/(α+β)N_* = \left(\frac{\alpha A}{\beta B}\right)^{1/(\alpha+\beta)}\left(\frac{C}{6}\right)^{\beta/(\alpha+\beta)}

Therefore N∗∝Cβ/(α+β)N_*\propto C^{\beta/(\alpha+\beta)} and T∗∝Cα/(α+β)T_*\propto C^{\alpha/(\alpha+\beta)}. When α=β\alpha=\beta, both grow as C1/2C^{1/2}. Using 0.34 and 0.28 gives approximately 0.45 for parameters and 0.55 for tokens. That explains why the lab's exact fit differs slightly from the headline 50/50 simplification.

Notice that the optimum equates weighted marginal improvements. The two loss contributions themselves need not be equal.

07 Read Critically

What the curves can tell you.

Scaling laws turn a large training run into a more informed bet. Small runs can help estimate a trend, reveal bottlenecks, and guide a budget. Their value comes from calibrated prediction within a regime, together with tests of how far the regime extends.

01

Loss is an average, not a capability inventory

A smooth improvement in next-token prediction does not imply that each benchmark or skill improves smoothly. Thresholded evaluations can look abrupt even when underlying performance changes gradually.

02

Quality and distribution matter

A billion duplicated, irrelevant, or low-quality tokens need not act like a billion useful new tokens. Changes to tokenizer, context length, and evaluation corpus can also change loss and fitted coefficients. Compare losses only when the measurement setup is compatible.

03

Extrapolation is an additional claim

A straight line over the measured range does not prove indefinite scaling. Kaplan explicitly discusses contradictions between extrapolated trends in section 6.3. A floor or a new bottleneck can bend the curve.

04

The optimum belongs to an objective

Minimum pretraining loss at fixed training FLOPs differs from minimum lifetime cost. Inference volume, latency, memory, data limits, and post-training can favor a different allocation.

For every scaling claim, ask: Which metric, on which data, with which architecture, under which budget?

08 Make It Stick

Turn recognition into understanding.

Work through the calculation before opening its solution. Each one tests a different piece of the argument.

01

Double the model

Using L∝N−0.076L\propto N^{-0.076}, what is the loss multiplier after doubling NN? What happens after doubling it again?

Solution

2−0.076≈0.9492^{-0.076}\approx0.949: about a 5.1% reduction each time. After two doublings, the multiplier is 4−0.076≈0.9004^{-0.076}\approx0.900. The percentage improvement per doubling stays constant; the absolute improvement shrinks.

02

Make the target expensive

Under the same pure parameter law, how much larger must a model be to reduce loss by 10%?

Solution

Solve k−0.076=0.9k^{-0.076}=0.9. Then k=0.9−1/0.076≈4.0k=0.9^{-1/0.076}\approx4.0. Approximately four times the parameters are required, with sufficient training and data. Small exponents make large improvements expensive.

03

Keep compute fixed

A model processes 100 billion tokens. Its replacement has four times as many parameters. How many tokens can the replacement process with the same approximate training FLOPs?

Solution

Since C≈6NTC\approx6NT, the token budget falls by a factor of four: 25 billion tokens. The compute constraint alone cannot determine which model has lower loss.

04

Separate exposure from new evidence

A training dataset contains 20 billion tokens and is processed for five epochs. What are DD and TT? Is this equivalent to 100 billion distinct training tokens?

Solution

D=20D=20 billion and T=100T=100 billion. Repetition uses compute and may improve learning, but it does not supply the same amount of new evidence as a larger, diverse dataset.

05

Connect loss and perplexity

Loss falls from 3.0 to 2.8 nats. What is the perplexity multiplier? Does that give a change in benchmark accuracy?

Solution

The multiplier is e2.8/e3.0=e−0.2≈0.819e^{2.8}/e^{3.0}=e^{-0.2}\approx0.819, about an 18.1% reduction. Benchmark accuracy cannot be inferred from this number alone.

Concept Check

Five quick decisions

0 / 5 answered
1. A straight descending line on log-log axes suggests...
2. At fixed compute, doubling the parameter count approximately...
3. A model stopped before convergence is...
4. Chinchilla revised which central recommendation?
5. With an additive loss floor, a pure power law describes...

Explain It In Your Own Words

"In the studied regime, held-out language-model loss follows approximate power laws in model size, data, and well-allocated compute. The best model for a budget depends on how size and training tokens are balanced. Later experiments changed that balance without discarding the idea of predictable scaling."

09 Return To The Paper

A more productive second reading.

  1. Abstract + section 1.2. Translate the three headline equations into words. State the conditions for each one.
  2. Figures 1, 4, and 9 + section 4. Identify a power-law segment, a saturation region, and the effect of a second bottleneck.
  3. Sections 5-6 + Figure 13. Track the distinction between training steps, critical batch size, raw compute, and adjusted compute.
  4. Section 6.3 + Appendix C. Read the authors' limits on extrapolation and experimental caveats.
  5. Chinchilla sections 3-5 + Appendix D. Compare allocation methods and the large-model validation. Keep the approximate 50/50 headline separate from the exact parametric fit.

A compact glossary

Cross-entropy
Average negative log probability assigned to the observed tokens.
Perplexity
eLe^L for loss measured with natural logarithms.
Epoch
One pass through a dataset.
Power-law exponent
The negative slope of log loss versus log resource in a pure power law.
Compute frontier
The lowest loss attainable at each compute budget among the configurations considered.
IsoFLOP comparison
A comparison of training runs that use the same compute budget.
Critical batch size
A scale marking the tradeoff between parallel speed and efficient use of compute.
Loss floor
An asymptotic loss term in a model of a particular prediction setup.

Sources and provenance

[1]
Jared Kaplan et al. (2020)Scaling Laws for Neural Language Models

arXiv:2001.08361. Headline exponents: equations 1.1-1.3. Joint data/size fit: equation 1.5 and table 2. Training-time law: equation 1.6. Compute allocation: section 6 and Appendix A. Extrapolation limits: section 6.3.

[2]
Jordan Hoffmann et al. (2022)Training Compute-Optimal Large Language Models

arXiv:2203.15556. Three estimation methods: section 3. Chinchilla validation: sections 4-5. Parametric fit used in the budget example: Appendix D.2. Learning-rate schedule: Appendix B.

This guide is an original explanation of the papers. All embedded plots are calculated illustrations, not copied figures or recovered experimental data. Numerical examples use the formulas identified next to them. The document and its interactive examples work offline; source links open online.

Typesetting and icon credits

Equations are typeset with KaTeX. Interface icons are from Lucide. Both are included under their respective licenses.

KATEX

The MIT License (MIT)

Copyright (c) 2013-2020 Khan Academy and other contributors

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.


LUCIDE

ISC License

Copyright (c) 2026 Lucide Icons and Contributors

Permission to use, copy, modify, and/or distribute this software for any
purpose with or without fee is hereby granted, provided that the above
copyright notice and this permission notice appear in all copies.

THE SOFTWARE IS PROVIDED "AS IS" AND THE AUTHOR DISCLAIMS ALL WARRANTIES
WITH REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES OF
MERCHANTABILITY AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR
ANY SPECIAL, DIRECT, INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY DAMAGES
WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS, WHETHER IN AN
ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISING OUT OF
OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THIS SOFTWARE.

---

The following Lucide icons are derived from the Feather project:

airplay, alert-circle, alert-octagon, alert-triangle, aperture, arrow-down-circle, arrow-down-left, arrow-down-right, arrow-down, arrow-left-circle, arrow-left, arrow-right-circle, arrow-right, arrow-up-circle, arrow-up-left, arrow-up-right, arrow-up, at-sign, calendar, cast, check, chevron-down, chevron-left, chevron-right, chevron-up, chevrons-down, chevrons-left, chevrons-right, chevrons-up, circle, clipboard, clock, code, columns, command, compass, corner-down-left, corner-down-right, corner-left-down, corner-left-up, corner-right-down, corner-right-up, corner-up-left, corner-up-right, crosshair, database, divide-circle, divide-square, dollar-sign, download, external-link, feather, frown, hash, headphones, help-circle, info, italic, key, layout, life-buoy, link-2, link, loader, lock, log-in, log-out, maximize, meh, minimize, minimize-2, minus-circle, minus-square, minus, monitor, moon, more-horizontal, more-vertical, move, music, navigation-2, navigation, octagon, pause-circle, percent, plus-circle, plus-square, plus, power, radio, rss, search, server, share, shopping-bag, sidebar, smartphone, smile, square, table-2, tablet, target, terminal, trash-2, trash, triangle, tv, type, upload, x-circle, x-octagon, x-square, x, zoom-in, zoom-out

The MIT License (MIT) (for the icons listed above)

Copyright (c) 2013-present Cole Bemis

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.