Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 5: GPUs, TPUs
Stanford Online
GPU systems prioritize massive throughput over serial execution, relying on a hierarchical memory structure where global memory access remains the primary bottleneck. Achieving peak performance requires minimizing data movement through operator fusion, which combines multiple operations into single kernels, and recomputation, which trades extra compute for a reduced memory footprint. Tiling plays a critical role by loading sub-matrices into fast shared memory, allowing for repeated data reuse and improved arithmetic intensity. These principles underpin FlashAttention, which optimizes attention mechanisms by computing softmax online and tiling matrix multiplications to avoid storing large intermediate activations. Understanding hardware-level constraints—including burst memory access and wave quantization—is essential for effectively scaling modern machine learning models and ensuring efficient resource utilization across both GPU and TPU architectures.
Sign in to continue reading, translating and more.
Open full episode in Podwise
