Chapter 3: The Real Bottleneck Is Memory, Not Math
Take our 70B reference model, stored at 8-bit (70 GB), on one H100. To produce a token, every one of the 70 billion weights must travel from memory into the compute units.
Moving the weights takes about 150 times longer than doing the math on them.
Time budget for one output token (70B model, 8-bit, one H100)
An RTX 4090 has ~1,008 GB/s of bandwidth. An 8B model at 8-bit is ~8 GB.
Predicted ceiling: 8 GB ÷ 1,008 GB/s ≈ 7.9 ms per token → ~126 tokens/s
The same model at 4-bit (~4.5 GB): ≈ 4.5 ms → ~220 tokens/s ceiling
Local-inference benchmarks typically land at 70–85% of these ceilings.
From Your World: The Pantry Door
Your weekend dinners: an absurdly fast chef, one narrow pantry door. Dinner comes out at the pace of the door.Your weekend dinners: an absurdly fast chef, one narrow pantry door. Dinner comes out at the pace of the door.
The Tensor Cores are busy 0.14 ms out of every 21 ms, under 1% of the time.
The GPU’s idle compute is where that difference comes from, and Chapter 4 shows how providers put it to work.
3. Arithmetic Intensity and the Roofline
There’s a single number that tells you which limit a workload hits.