The A100 40GB and 80GB share 312 TFLOP/s

The quote has two A100 lines. 40GB and 80GB. The Tensor cell is 312 TFLOP/s on both. You pick as if only capacity doubled.

Read the NVIDIA datasheet down the column and the split is memory. 40GB is HBM2 at 1,555 GB/s. 80GB SXM is HBM2e at 2,039 GB/s. 80GB PCIe is 1,935 GB/s. Core count is the same. The belt is not.

Wall time is max(memory time, math time). When FLOP per byte sits below the knee, sharpening the millstone does not add flour. The millstone waits on the belt.

It is one millstone and two belt widths. In LLM decode the millstone is the Tensor Core. The belt is HBM.

The numbers are an equation. Units are SI. 1 GB = 10^9 B. Tensor peak is the datasheet base 312 TFLOP/s. The sparsity 624 is not used here.

1. The knee depends on the SKU

The knee is peak math divided by bandwidth.

  • 80GB SXM: 312e12 / 2039e9 ≈ 153 FLOP/B
  • 80GB PCIe: 312e12 / 1935e9 ≈ 161 FLOP/B
  • 40GB: 312e12 / 1555e9 ≈ 201 FLOP/B

Same 312, narrower belt, higher knee. Each memory byte must buy more math before the millstone is busy.

Decode on weights alone is 1 FLOP/B, from the previous essay. That is 1/200 of the 40GB knee and 1/150 of the 80GB SXM knee. Util can be high while Tensor cores wait. Both SKUs stay on the belt in decode.

A100 roofline knee 153 vs 201 FLOP/B
The knee depends on the SKU. 153 FLOP/B on 80GB SXM, 201 on 40GB. Points are assumptions.

2. Batch has a ceiling

The handle that moves the point right is batch. Read 140 GB of weights once and stack questions on that read, and intensity from weights alone is about batch B FLOP/B. The 80GB SXM knee is B≈153. The 40GB knee is B≈201.

KV changes the formula. Intensity is 2P / (2P/B + KV). As B grows it stops at 2P/KV. For 70B GQA at 32K, KV is 0.328 MB per token × 32,768 ≈ 10.7 GB. 140 / 10.7 ≈ 13 FLOP/B. Below the 153 knee. Decode stays memory-side even if you raise batch.

3. A same-generation forecast

40GB versus 80GB is not extra cores. It is a forecast when belts differ. In batch-1 decode the 80GB SXM can be faster by the bandwidth ratio. 2039 / 1555 ≈ 1.31×. That is a datasheet division, not a measurement in this post.

MIG in seven slices cuts the belt per instance. 5 GB each on the 40GB, 10 GB on the 80GB. More slices do not widen the card.

“An A100 is an A100.” That is one datasheet cell. “Raise batch and decode becomes compute-bound.” KV fills the belt again before the knee.

Notes

  • Tensor 312 TFLOP/s and 1,555 / 1,935 / 2,039 GB/s: NVIDIA A100 datasheet (SXM4 and PCIe). Sparsity 624 omitted.
  • 0.328 MB/token for 70B GQA is the same formula as the previous essay. 32K KV 10.7 GB and the 13 FLOP/B ceiling are that division.

Leave a Comment