Learn · Intermediate
Amdahl’s law: why a faster AI kernel barely changes the waiting time
Amdahl’s law says an optimization can speed up an application only in proportion to the time the application originally spent on the part being improved. Making one component twice as fast does not make the whole system twice as fast when other work remains unchanged. This principle explains why impressive AI kernel benchmarks can coexist with modest improvements in a user’s waiting time.
Start with elapsed time
Gene Amdahl set out the classic limitation in his 1967 paper on large-scale computing. The original setting concerned the limits of parallel computation, but the same accounting applies whenever an optimization affects only part of a fixed workload.
Imagine a restaurant meal that takes 20 minutes to arrive. Cooking takes 10 minutes, while preparation and delivery take another 10. A new oven halves the cooking time. The customer waits 15 minutes rather than 20: the meal arrives about 1.33 times as fast. The oven improved by a factor of two, but it affected only half the original waiting time.
That difference is not a defect in the oven’s benchmark. It is a difference between a component and the process containing it. AI systems contain many such components: reading data, moving it to a device, loading a model, executing operations, coordinating tools, and returning results. A correct optimization can still leave most of the critical path untouched.
The calculation
Let p be the fraction of the original elapsed time that an optimization improves. Let s be that component’s speedup. Normalize the original total time to one. The unchanged part still takes 1 minus p; the improved part now takes p divided by s.
The overall speedup is therefore:
overall speedup = 1 / ((1 - p) + p / s)
Suppose a particular operation accounts for 20% of an application’s time and becomes four times faster. The new time is 0.8 plus 0.2 divided by 4, or 0.85 of the original. Overall speedup is about 1.18 times. That is a 15% reduction in elapsed time, not a fourfold improvement in the whole application.
Even infinite speed has a ceiling. If the improved component could finish instantly, the other 80% would remain. The maximum possible speedup would be 1.25 times. This ceiling tells an engineer when further effort on the same component offers little user-visible value.
Why today’s browser result needs this distinction
Hugging Face’s WebGPU kernels announcement reports a 2.57-times geometric-mean speedup for selected individual operations against a reference runtime on one Apple M4 GPU. The authors Nico Martin and Joshua Lochner exclude setup and data-transfer overhead, and the experiment is not a complete-model benchmark.
Those facts do not give the p needed to predict an application’s speed. If the affected operations occupy 30% of original elapsed time, and if that portion uniformly improves by 2.57 times, the simplified formula gives roughly 1.22 times overall speedup. This is an illustrative calculation, not a measured result for their library. An average across different operations is not itself the measured speedup of one application’s affected portion.
The useful next experiment is to profile the actual application, identify how much time the eligible operations consume, and measure the full workflow after replacing them. Ground Truth’s lesson on memory-bound inference explains one possible underlying bottleneck; Amdahl’s law explains how a local improvement translates into the complete waiting time.
Keep the workload fixed
The standard formula assumes the same work before and after, with the unchanged portion staying unchanged. Adding setup costs, changing precision, producing different outputs, or moving work onto another device can violate those assumptions. An apparently faster implementation may also trade away correctness. Measure comparable outputs and include any new overhead.
Parallel systems add another complication: elapsed times do not always add neatly. If loading overlaps computation, halving the load stage might save nothing until it becomes the longest stage. Use the critical path rather than blindly summing durations from overlapping tasks. Continuous batching can increase throughput while changing latency, so completed requests per second and one person’s waiting time require separate measurements.
Likewise, an agent can generate tokens faster yet finish a task later if it makes extra tool calls or retries. Each retry changes the amount of work required for completion, as the lesson on inference economics explains. Performance measurements must include the surrounding process that defines a finished task, rather than holding only the model’s token-generation rate constant.
Use the law to choose the next change
Profile a representative workload before selecting an optimization. Estimate the affected time fraction, calculate the best plausible gain, and compare that gain with the implementation cost. Afterward, rerun the entire workflow and check output quality as well as speed.
Amdahl’s law does not say component improvements are unimportant. A widely reused kernel can save substantial resources across many workloads, and a bottleneck may move after the first improvement. It says the benefit must be measured at the level being claimed. Faster arithmetic, faster tokens, faster completed tasks, and shorter user waiting times are related outcomes with different denominators.
Key questions
How do I calculate the speedup of one optimized stage?
Why can a 2.57-times-faster kernel leave an AI app almost unchanged?
Does Amdahl’s law predict performance when the workload grows?
Cite this
APA
Ground Truth. (2026, October 1). Amdahl’s law: why a faster AI kernel barely changes the waiting time. Ground Truth. https://groundtruth.day/learn/amdahls-law-and-end-to-end-ai-speed.html
BibTeX
@misc{groundtruth:amdahls-law-and-end-to-end-ai-speed,
title = {Amdahl’s law: why a faster AI kernel barely changes the waiting time},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/learn/amdahls-law-and-end-to-end-ai-speed.html}
}