All posts

JAX Book Chapter 0 Notes

JAX Book Chapter 0 Notes

Symbology (in the transformer)

  • B
    • batch?
  • L
    • Number of layers
  • T
    • Sequence length (query)
  • S
    • Sequence length (key value)
  • V
    • vocab
  • D
    • embedding dimension of model
  • F
    • MLP hidden dimension
  • H
    • attention head width
  • N
    • query heads
  • K
    • Key Value heads

Architecture

  • High level data flow of a transformer
    • token ids
    • vectorize each
    • inter-position effects (Attention)
    • intra-position effect (MLP)
    • repeat
    • output next token probabilities

Principles

  • How well models score on benchmarks can be a bit magical - but performance doesn't have to be
    • a grounded understnding of how hardware works will deterministically help guide you towards ways of improving performance

Definitions

  • Strong scaling - throughput increases approximately linearly to chips used in training or inference

Notes

  • Expect to learn in section 1 roofline analysis & hardware limitations that determine what is/is not possible on today's class of hardware
  • TPUs, interconnected system, interchip-links
  • The exact transformer model
    • i..e memory used
    • time spent on compute or comms
    • when attention will become important relative to feed forward blocks