Skip to contentAll posts
- B
- L
- T
- S
- Sequence length (key value)
- V
- D
- embedding dimension of model
- F
- H
- N
- K
Architecture
- High level data flow of a transformer
- token ids
- vectorize each
- inter-position effects (Attention)
- intra-position effect (MLP)
- repeat
- output next token probabilities
Principles
- How well models score on benchmarks can be a bit magical - but performance doesn't have to be
- a grounded understnding of how hardware works will deterministically help guide you towards ways of improving performance
Definitions
- Strong scaling - throughput increases approximately linearly to chips used in training or inference
Notes
- Expect to learn in section 1 roofline analysis & hardware limitations that determine what is/is not possible on today's class of hardware
- TPUs, interconnected system, interchip-links
- The exact transformer model
- i..e memory used
- time spent on compute or comms
- when attention will become important relative to feed forward blocks