What can this budget buy?
For a fixed spend, a bigger model means less data. Set your budget and see the frontier — and where the scaling laws say the sweet spot is.
10,000
GPU compute, tuning runs and storage.
The frontier
89.13B tokens
More data, smaller model for the same spend.
At this point on the curve
- Max model size (params)
- 4.41B
- Dataset
- 89.13B tokens
- Tokens per parameter
- 20.2
- Wall-clock
- 6.27 days
against model params
- Compute
- $9,946
- Storage
- $54
- Total
- $10,000
- Budget used
- 100%
of $10,000
Chinchilla-optimal. 20.2 tok/param — within the 10–30× compute-optimal zone This is a well-balanced configuration.
Training schedule & storage
Derived from your setup. Change any of these and they hold until you edit the project setup above.
Share of a full run's cost per trial.
Auto-recommended: 1 for a ≤60 day run.
p5.48xlarge (8x H100 80GB) · bf16 · MFU 55% · $55.04/hr per instance · 8 GPUs total