How much data do I need?
Scaling laws set a floor and a ceiling on useful dataset size. Pick a model and see where your
data lands.
Pre-training data requirements
Chinchilla scaling laws (Hoffmann et al. 2022) for a 7.00B-parameter model.
Each bar spans the token range for one tier, measured against
7.00B parameters.
In real examples
Token counts are hard to picture. Convert them into the thing you actually collect.
520 tokens per text example.
- Hard floor
- < 13.46M
700.00M–7.00B tokens
- Practical minimum
- 13.46M–134.62M
7.00B–70.00B tokens
- Compute-optimal
- 134.62M–403.85M
70.00B–210.00B tokens
- Inference-optimal
- 403.85M–2.69B
210.00B–1.40T tokens
Text examples needed to reach each tier.