Training Compute Calculator

· LIVE

How many FLOPs, GPU-hours, and days a pretraining run needs, from the 6ND rule and the hardware you point at it.

Parameters (active for MoE)70B
Training tokens1.4T
GPUs1,024
MFU (BF16 reference)40%
GPU
Your GPU-hour rate$2.00
Estimate
Training FLOPs
5.88e23
GPU-hours
412.9K
Wall clock
16.8 days
Tokens per parameter20.0 · near the dense-model 20× reference
Peak BF16 (H100 SXM, dense)989 TFLOPS
Effective @ 40% MFU396 TFLOPS
Cluster (1,024 GPUs)405.1 PFLOPS
Cost @ $2.00/GPU-hour$825.7K
Approximate model compute using 6ND and dense BF16 peak throughput. MFU must reflect your workload and cluster size; the estimate does not check memory fit or predict scaling efficiency. Attention compute is omitted. Published GPU-hours are reference totals, not predictions at the selected MFU.

What it does

Pick a parameter count, a token budget, a GPU and a cluster size. The tool returns the total training compute, how many GPU-hours that is at a given utilization, how long the run takes on your cluster, and what it costs at your hourly rate.

The math

  • Training FLOPs ≈ 6 × N × D, where N is the parameter count and D is the training token count. The calculator converts billions of parameters and trillions of tokens to raw counts. This approximates the forward and backward parameter-matrix operations; it is not a complete operation count.
  • GPU-hours = FLOPs ÷ (peak FLOPS × MFU) ÷ 3600. TFLOPS are multiplied by 10¹². MFU is entered as a percentage and converted to a fraction.
  • Days = GPU-hours ÷ GPU count ÷ 24. This assumes the chosen MFU holds at the chosen cluster size. More GPUs do not automatically preserve utilization.
  • Cost = GPU-hours × hourly rate. This is GPU rental cost only.

The 20 tokens per parameter button is a rough dense-model reference inspired by Chinchilla, not a universal optimum. Data, architecture, and the intended inference budget change the best allocation. Applying that ratio to active MoE parameters does not establish compute optimality.

PaLM’s MFU accounting includes an additional attention term: approximately 12 × layers × query heads × head dimension × sequence length per token. This calculator omits it. The error depends on model size and context length; even at 8K context it can be substantial for smaller models. For example, 32 layers, 32 heads, head dimension 128, and 8,192 tokens add about 12.9 billion FLOPs per token, or 27% on top of 6N for an 8B model under that convention.

Published reference runs

RunParameters × tokensApproximate 6ND FLOPsReported GPU-hours
Llama 2 70B70B × 2T8.4 × 10²³1,720,320 A100
Llama 3.1 405B405B × 15.6T3.79 × 10²⁵30.84M H100
DeepSeek-V337B active × 14.8T3.29 × 10²⁴2.664M H800, pretraining only

Sources: Llama 2 model card, Llama 3.1 model card, Llama 3 report, and DeepSeek-V3 report.

Dividing approximate 6ND compute by reported GPU-hours and BF16 peak gives about 43.5% for Llama 2 70B and 34.5% for Llama 3.1 405B. These are inferred averages, not independent validation of the estimator or directly measured in-run MFU. Reported totals can cover different training stages and operating conditions.

For scale, 30.84M GPU-hours divided by a constant 16,384 GPUs is 78.4 days. Meta’s 54-day reliability observation is a snapshot, not the full training duration. DeepSeek’s 2.664M pretraining GPU-hours divided by 2,048 GPUs is 54.2 days. Selecting a preset keeps your chosen MFU and hourly rate, so the estimate need not equal the reported total.

Hardware and limitations

  • Peak values use dense BF16 throughput, without structured sparsity: A100 312 TFLOPS, H100 SXM and H200 SXM approximately 989, and B200 2,250. See NVIDIA A100, H100, H200, and NVIDIA’s non-sparse throughput table.
  • The DeepSeek preset uses an H100 compute proxy for an H800 FP8 mixed-precision run. A BF16 denominator produces a higher, not lower, utilization percentage than a larger FP8 denominator for the same work and time. Mixed-precision utilization needs an explicit accounting convention, so this preset shows reported hours without an implied MFU comparison.
  • For MoE, active parameters give only a rough compute estimate. Routing, attention, auxiliary objectives, and communication add work. Memory capacity still depends on total parameters and training state.
  • Memory fit, batch size, parallelism, interconnect, failures, checkpointing, and data loading are not modeled independently. Choose an end-to-end MFU appropriate to those conditions; the calculator cannot determine whether the configuration is feasible.
  • Rental cost excludes storage, networking, and separately billed failed or experimental runs.