Daniel Kim · 1mo ago

Contributed CUDA kernel optimization reducing inference latency by 34%

Found a warp divergence issue in the attention mechanism. Restructured the memory access pattern to be coalesced. The perf jump was immediate and reproducible across hardware.