All posts

PMPP Chapter 3 Notes

PMPP Chapter 3 Notes

Note: I've learned some things previous in CPU microarchitecture land (in college) so reading this chapters raised some questions there, which I've asked chatgpt about (shoutout to 5.6 Sol + custom instructions). So these notes are not purely what's covered in PMPP.

Architecture

  • blockIdx and threadIdx have syntax candy for saying .x .y .z, still same linear array underneath.
  • CUDA streams are thread safe
  • Typical topology in terms of latency/scale considerations
    • Network fabric
      • NVLink domain
        • GPU
          • SM
            • registers

Definitions

  • BLAS
    • Basic Linear Algebra Subprograms
    • Three levels of linear algebra functions
      • Level 1: vector operations
      • Level 2: Matrix vector multiplications
      • Level 3: Matrix Matrix multiplications

Syntax

  • struct dim3(unsigned x = 1, unsigned y = 1, unsigned z = 1)
  • (z, y, x) when drawing typically and note that CUDA lays things out row major
  • Implicitly provided variables
    • threadIdx
    • blockIdx
    • blockDim (shape of threads in a block)
    • gridDim (shape of blocks in a grid)

Numbers

  • allowed values of gridDim.x is [1, 2^31 - 1]
  • Allowed values of gridDim.y, gridDim.z are [1, 2^16-1]
  • Hopper family empirical benchmarks
    • L2 access in hundreds of cycles
    • HBM misses ~ 600-700

Additional notes

  • Parallelism in multi-GPU work
    • Data parallelism
      • GPUs process different batch shards
    • Tensor parallelism
      • individual layers and/or matrix operations are sharded
    • Pipeline parallelism
      • Different layer ranges live on different GPUs
    • Context parallelism
      • sequence positions are divided
    • Expert parallelism
      • MoE experts are distributed
    • Fully sharded data parallelism
      • parameters, gradients, optimizer state are partitioned
  • CUDA supports atomics
    • i.e. atomicAdd
    • memory ordering is relaxed (at least for legacy functions afaik)
    • newer cuda::atomic supports C++ memory model w/ release, acquire syntax
  • atomic_ref types of up to 8 bytes are always lockfree
  • lock free vs wait free vs contention free
    • Lock free - system as whole makes progress
    • Wait free - each operation guaranteed to complete w/in a number of its own steps
    • Contention free - threads do not interfere w/ each other
  • Natural aligned load/store performed indivisibly != race condition free
    • It does not make an operation atomic
  • volatile does not exist in CUDA
  • NaNs and special floating point can complicate CAS implementations
  • Relaxed does not mean the atomic result can drift around for an arbitrary time, it merely means that the atomic operation does not order unrelated memory operations for communication w/ other threads
  • acquire/release and relaxed describe which executions and observations are legal, not how quickly cache lines propagate
  • Typically, device scoped atomic shared across SMs is coordinated at an appropriate coherence point in GPU memory hierarchy
    • generally, this should be the L2
    • uncontended operation may become observable quite quickly
    • hot address w/ thousands of contenders develops a serialized queue
    • thread scheduler can also delay when a particular producer or consumer runs
  • What's hard about kernel optimization?
    Problem Typical starting point Difficulty
    Standard dense GEMM on one GPU cuBLASLt Mostly already industrialized
    Custom fused or irregular single-GPU operator Triton/CUDA/CUTLASS Tiling, layouts, registers, pipelines, numerics
    Standard multi-GPU collective NCCL ?
    Custom distributed attention or MoE operator Custom kernels plus communication Local kernels, topology, overlap, balance, synchronization