floply

What a training run costs, before you commit the budget.

Cached · fetched just now

How much data do I need?

Scaling laws set a floor and a ceiling on useful dataset size. Pick a model and see where your data lands.

7,000,000,000

Pre-training data requirements

Chinchilla scaling laws (Hoffmann et al. 2022) for a 7.00B-parameter model.

1.00B10.00B100.00B1.00THard floorHard floor — < 1 tok/param. Fewer tokens than parameters — training will likely diverge or severely underfit.700.00M–7.00BPractical minimumPractical minimum — 1–10 tok/param. Usable but significantly undertrained — expect poor generalisation and high loss.7.00B–70.00BCompute-optimalCompute-optimal — 10–30 tok/param. Chinchilla-optimal zone (Hoffmann et al. 2022) — best loss per FLOP.70.00B–210.00BInference-optimalInference-optimal — 30–200 tok/param. Training smaller models longer reduces inference cost — the LLaMA / Mistral strategy.210.00B–1.40T
Each bar spans the token range for one tier, measured against 7.00B parameters.

In real examples

Token counts are hard to picture. Convert them into the thing you actually collect.

1 word ≈ 1.3 BPE tokens.

520 tokens per text example.

Hard floor
< 13.46M

700.00M–7.00B tokens

Practical minimum
13.46M–134.62M

7.00B–70.00B tokens

Compute-optimal
134.62M–403.85M

70.00B–210.00B tokens

Inference-optimal
403.85M–2.69B

210.00B–1.40T tokens

Text examples needed to reach each tier.

A Floating Point Labs project

Estimates only. FLOPs-based modelling assumes ideal scaling; real runs vary with data loading, checkpointing overhead, and cluster utilisation.