Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Section H: Advanced Computer Architecture

This section covers the deep, graduate-level topics in computer architecture that separate candidates who truly understand modern processors from those who only know textbook basics. These are the topics that come up in senior/staff-level interviews at Intel, AMD, Apple, NVIDIA, Google, and Amazon.

Topic Map

mindmap
  root((Advanced Architecture))
    OoO Execution
      Tomasulo Algorithm
      Register Renaming
      Reorder Buffers
      Reservation Stations
      Instruction Windows
    Branch Prediction
      Perceptron Predictors
      TAGE
      Indirect Branches
      Return Prediction
      Speculative Execution
    Side Channels
      Spectre Variants
      Meltdown Variants
      Transient Execution
      Cache Timing Attacks
    Cache Coherence
      MESI/MOESI Deep Dive
      Directory Protocols
      Memory Consistency
      TSO / ARM / RISC-V Models
      Store Buffers & Load Queues
    Memory Systems
      Hardware Prefetchers
      Cache Replacement
      DRAM Scheduling
      RowHammer
      Refresh Mechanisms
    Modern Interconnects
      CXL / CXL.mem
      Chiplets & UCIe
      2.5D / 3D Packaging
      HBM / DDR5 / LPDDR
      Persistent Memory
    Accelerators
      DPUs / Smart NICs
      FPGAs / CGRAs
      TPU / Tensor Cores
      GPU Architecture Deep
      PIM & Computational Storage

Reading Order

OrderFilePrerequisitesCore Focus
1Out-of-Order ExecutionBasic pipelining, data hazardsTomasulo, renaming, ROB internals
2Branch Prediction AdvancedBasic branch predictionNeural/TAGE predictors, indirect branches
3Side ChannelsOoO execution, branch prediction, cache coherenceSpectre, Meltdown, transient execution
4Cache Coherence AdvancedMESI/MOESI basicsDirectory protocols, memory models, TSO
5Memory System AdvancedCache basics, DRAM basicsPrefetching, replacement, RowHammer
6Modern InterconnectsPCIe basics, memory hierarchyCXL, chiplets, packaging, NVRAM
7AcceleratorsGPU basics, SIMDDPUs, TPUs, FPGAs, PIM

How This Differs from Earlier Sections

Earlier SectionsThis Section (Advanced)
What is OoO?How Tomasulo’s algorithm works in hardware
2-bit saturating countersPerceptron and TAGE neural predictors
MESI state diagramDirectory coherence, memory consistency models
Basic cache prefetchingStream prefetchers, Markov prefetchers, stride detection
PCIe bandwidth numbersCXL coherent interconnect, memory pooling
GPU programming basicsWarp scheduling, tensor cores, GPU cache hierarchy

Cross-References

  • Pipelining — Foundation for OoO and branch prediction
  • Cache Coherence Basics — Prerequisite for coherence deep dive
  • DRAM — Prerequisite for DRAM scheduling and RowHammer
  • GPU Basics — Prerequisite for accelerator deep dive
  • x86-64 — Real processor implementations
  • AMD Zen — AMD-specific architecture details
  • Apple Silicon — ARM-based high-performance design