All posts

Notes on CPU Cache Hierarchies

CPU Cache Hierarchies

Architecture

  • CPU chips are typically decomposed into multiple subsections. Each vendor has their own term for it

    • AMD: CCD
    • Intel: Die
  • The general hierarchy looks something like

    • the chip itself
      • die
        • physical core
          • 2 HW threads (typically)
          • private l1 instruction cache
          • private l1 data cache
          • (typically) private l2
        • coherence connection
        • L3 either for itself or shared w/ other dies
      • memory controller
      • PCIe/CXL I/O
      • socket level coherence stuff
  • L3 structure

    • AMD: Each CCD has one shared L3
    • Intel: Cores connected via mesh to logically shared, but physically separate L3s.
  • Intel exposes relevant uncore performance counters

    • LLC lookups
    • Directory lookups
    • Snoops sent/responses
  • A normal store can enter the store buffer and retire before the coherence transaction that makes it globally visible has finished

  • These are different timings

    • Latency until store instruction can retire
    • Latency until store has obtained exclusive ownership of cache line
    • Latency until another core can observe the store
  • Loads that are executed speculatively will need to be replayed when coherence order conflicts with it

Definitions

  • CAT

    • Cache Allocation Technology (Intel terminology)
    • Can limit destructive interference - i.e. limits how many ways each core can occupy
  • CHA

    • Caching home agent
  • TOR

    • Transaction tracking/order data structure in CHA
  • Socket level fabric

    • AMD: Infinity fabric
    • Intel: Mesh
  • SNC

    • Sub Numa Clustering (Intel terminology)
  • NUMA subdivision terminology

    • AMD: NPS1, NPS2, NPS4, L3 as NUMA
    • Intel: SNC2, SNC4
  • UPI

    • Ultra Path Interconnect (Intel terminology)
    • Inter-CPU interconnect
  • xGMI

    • External/Socket Global Memory Interconnect (AMD terminology)

Numbers

AMD

  • Older CCDs contained multiple CCX's
  • Newer CCDs have exactly one CCX
  • 8-16 cores per CCX
  • L3 cache on entire chip on order of ~512mb-1gb ish size
  • CCX L3 ~32MiB size

Intel

  • LLC cache hit ~50 cycles
  • Line in M/E in another core, same SNC ~ 120-132 cycles
  • Line in M/E in another SNC domain, same socket ~ 134-145 cycles
  • Difference between nearby vs distance cores in one SNC domain ~ 5 cycles

Tools

  • Intel's memory latency checker has a dedicated
    • mlc --c2c_latency for cache-to-cache transfer latency

Conceptual distinctions

  • Coherence vs Atomic vs Locks
    • Coherence - hardware level, who has data & can write

    • Atomic - can other threads interleave

    • Memory ordering - in what order may others observe operations to different memory addresses

    • Lock - purely software level protocol

    • Coherence required for atomics

    • Atomic + ordering guarantees gives locks

Some Resources