Note: I've learned some things previous in CPU microarchitecture land (in college) so reading this chapters raised some questions there, which I've asked chatgpt about (shoutout to 5.6 Sol + custom instructions). So these notes are not purely what's covered in PMPP.
Architecture
- blockIdx and threadIdx have syntax candy for saying
.x .y .z, still same linear array underneath. - CUDA streams are thread safe
- Typical topology in terms of latency/scale considerations
- Network fabric
- NVLink domain
- GPU
- SM
- registers
- SM
- GPU
- NVLink domain
- Network fabric
Definitions
- BLAS
- Basic Linear Algebra Subprograms
- Three levels of linear algebra functions
- Level 1: vector operations
- Level 2: Matrix vector multiplications
- Level 3: Matrix Matrix multiplications
Syntax
struct dim3(unsigned x = 1, unsigned y = 1, unsigned z = 1)(z, y, x)when drawing typically and note that CUDA lays things out row major- Implicitly provided variables
threadIdxblockIdxblockDim(shape of threads in a block)gridDim(shape of blocks in a grid)
Numbers
- allowed values of
gridDim.xis[1, 2^31 - 1] - Allowed values of
gridDim.y,gridDim.zare[1, 2^16-1] - Hopper family empirical benchmarks
- L2 access in hundreds of cycles
- HBM misses ~ 600-700
Additional notes
- Parallelism in multi-GPU work
- Data parallelism
- GPUs process different batch shards
- Tensor parallelism
- individual layers and/or matrix operations are sharded
- Pipeline parallelism
- Different layer ranges live on different GPUs
- Context parallelism
- sequence positions are divided
- Expert parallelism
- MoE experts are distributed
- Fully sharded data parallelism
- parameters, gradients, optimizer state are partitioned
- Data parallelism
- CUDA supports atomics
- i.e. atomicAdd
- memory ordering is relaxed (at least for legacy functions afaik)
- newer cuda::atomic supports C++ memory model w/ release, acquire syntax
atomic_reftypes of up to 8 bytes are always lockfree- lock free vs wait free vs contention free
- Lock free - system as whole makes progress
- Wait free - each operation guaranteed to complete w/in a number of its own steps
- Contention free - threads do not interfere w/ each other
- Natural aligned load/store performed indivisibly != race condition free
- It does not make an operation atomic
volatiledoes not exist in CUDA- NaNs and special floating point can complicate CAS implementations
- Relaxed does not mean the atomic result can drift around for an arbitrary time, it merely means that the atomic operation does not order unrelated memory operations for communication w/ other threads
- acquire/release and relaxed describe which executions and observations are legal, not how quickly cache lines propagate
- Typically, device scoped atomic shared across SMs is coordinated at an appropriate coherence point in GPU memory hierarchy
- generally, this should be the L2
- uncontended operation may become observable quite quickly
- hot address w/ thousands of contenders develops a serialized queue
- thread scheduler can also delay when a particular producer or consumer runs
- What's hard about kernel optimization?
Problem Typical starting point Difficulty Standard dense GEMM on one GPU cuBLASLt Mostly already industrialized Custom fused or irregular single-GPU operator Triton/CUDA/CUTLASS Tiling, layouts, registers, pipelines, numerics Standard multi-GPU collective NCCL ? Custom distributed attention or MoE operator Custom kernels plus communication Local kernels, topology, overlap, balance, synchronization
Links and references
- https://docs.nvidia.com/cuda/cuda-programming-guide/05-appendices/cpp-language-extensions.html
- ^ these are not complete I will add as I remember to