
The Launch Tax: How CUDA Graphs Keep Your GPU From Sitting Idle
Running one token through a language model isn't one big GPU operation. It's hundreds to thousands of tiny ones, and each of those has to be launched by the CPU first. During decode, those little operations finish so fas








