Architecture
CPU chips are typically decomposed into multiple subsections. Each vendor has their own term for it
- AMD: CCD
- Intel: Die
The general hierarchy looks something like
- the chip itself
- die
- physical core
- 2 HW threads (typically)
- private l1 instruction cache
- private l1 data cache
- (typically) private l2
- coherence connection
- L3 either for itself or shared w/ other dies
- physical core
- memory controller
- PCIe/CXL I/O
- socket level coherence stuff
- die
- the chip itself
L3 structure
- AMD: Each CCD has one shared L3
- Intel: Cores connected via mesh to logically shared, but physically separate L3s.
Intel exposes relevant uncore performance counters
- LLC lookups
- Directory lookups
- Snoops sent/responses
A normal store can enter the store buffer and retire before the coherence transaction that makes it globally visible has finished
These are different timings
- Latency until store instruction can retire
- Latency until store has obtained exclusive ownership of cache line
- Latency until another core can observe the store
Loads that are executed speculatively will need to be replayed when coherence order conflicts with it
Definitions
CAT
- Cache Allocation Technology (Intel terminology)
- Can limit destructive interference - i.e. limits how many ways each core can occupy
CHA
- Caching home agent
TOR
- Transaction tracking/order data structure in CHA
Socket level fabric
- AMD: Infinity fabric
- Intel: Mesh
SNC
- Sub Numa Clustering (Intel terminology)
NUMA subdivision terminology
- AMD: NPS1, NPS2, NPS4, L3 as NUMA
- Intel: SNC2, SNC4
UPI
- Ultra Path Interconnect (Intel terminology)
- Inter-CPU interconnect
xGMI
- External/Socket Global Memory Interconnect (AMD terminology)
Numbers
AMD
- Older CCDs contained multiple CCX's
- Newer CCDs have exactly one CCX
- 8-16 cores per CCX
- L3 cache on entire chip on order of ~512mb-1gb ish size
- CCX L3 ~32MiB size
Intel
- LLC cache hit ~50 cycles
- Line in M/E in another core, same SNC ~ 120-132 cycles
- Line in M/E in another SNC domain, same socket ~ 134-145 cycles
- Difference between nearby vs distance cores in one SNC domain ~ 5 cycles
Tools
- Intel's memory latency checker has a dedicated
mlc --c2c_latencyfor cache-to-cache transfer latency
Conceptual distinctions
- Coherence vs Atomic vs Locks
Coherence - hardware level, who has data & can write
Atomic - can other threads interleave
Memory ordering - in what order may others observe operations to different memory addresses
Lock - purely software level protocol
Coherence required for atomics
Atomic + ordering guarantees gives locks
Some Resources
- https://chipsandcheese.com/p/a-look-into-intel-xeon-6s-memory
- These are not exhaustive - I will just put ones here as I remember to.