81 Kernels, 97.4% of a Megakernel: Reconstructing Persistence with PDL
How an 81-kernel CUDA Graph uses PDL to recover 97.4% of a persistent megakernel's end-to-end throughput on H100.
Research notes · Paper reading · Engineering
Longer notes on GPU systems, compilers, distributed systems, and the ideas encountered along the way.
How an 81-kernel CUDA Graph uses PDL to recover 97.4% of a persistent megakernel's end-to-end throughput on H100.