floply

What a training run costs, before you commit the budget.

Cached · fetched just now

What can this budget buy?

For a fixed spend, a bigger model means less data. Set your budget and see the frontier — and where the scaling laws say the sweet spot is.

10,000

GPU compute, tuning runs and storage.

The frontier

1.00B10.00B100.00B1.00T10.00T100.00M1.00B10.00B100.00BBudget-optimal: 88.67B tokens, 4.43BYour selection: 89.13B tokens, 4.41BDataset size (tokens)Max model size (params)
What the budget buys Chinchilla optimal (N = D / 20)
89.13B tokens

More data, smaller model for the same spend.

At this point on the curve

Max model size (params)
4.41B
Dataset
89.13B tokens
Tokens per parameter
20.2

against model params

Wall-clock
6.27 days
Compute
$9,946
Storage
$54
Total
$10,000
Budget used
100%

of $10,000

Chinchilla-optimal. 20.2 tok/param — within the 10–30× compute-optimal zone This is a well-balanced configuration.

Training schedule & storage

Derived from your setup. Change any of these and they hold until you edit the project setup above.

Share of a full run's cost per trial.

Auto-recommended: 1 for a ≤60 day run.

p5.48xlarge (8x H100 80GB) · bf16 · MFU 55% · $55.04/hr per instance · 8 GPUs total

A Floating Point Labs project

Estimates only. FLOPs-based modelling assumes ideal scaling; real runs vary with data loading, checkpointing overhead, and cluster utilisation.