This section covers the deep, graduate-level topics in computer architecture that separate candidates who truly understand modern processors from those who only know textbook basics. These are the topics that come up in senior/staff-level interviews at Intel, AMD, Apple, NVIDIA, Google, and Amazon.
mindmap
root((Advanced Architecture))
OoO Execution
Tomasulo Algorithm
Register Renaming
Reorder Buffers
Reservation Stations
Instruction Windows
Branch Prediction
Perceptron Predictors
TAGE
Indirect Branches
Return Prediction
Speculative Execution
Side Channels
Spectre Variants
Meltdown Variants
Transient Execution
Cache Timing Attacks
Cache Coherence
MESI/MOESI Deep Dive
Directory Protocols
Memory Consistency
TSO / ARM / RISC-V Models
Store Buffers & Load Queues
Memory Systems
Hardware Prefetchers
Cache Replacement
DRAM Scheduling
RowHammer
Refresh Mechanisms
Modern Interconnects
CXL / CXL.mem
Chiplets & UCIe
2.5D / 3D Packaging
HBM / DDR5 / LPDDR
Persistent Memory
Accelerators
DPUs / Smart NICs
FPGAs / CGRAs
TPU / Tensor Cores
GPU Architecture Deep
PIM & Computational Storage
Order File Prerequisites Core Focus
1 Out-of-Order Execution Basic pipelining, data hazards Tomasulo, renaming, ROB internals
2 Branch Prediction Advanced Basic branch prediction Neural/TAGE predictors, indirect branches
3 Side Channels OoO execution, branch prediction, cache coherence Spectre, Meltdown, transient execution
4 Cache Coherence Advanced MESI/MOESI basics Directory protocols, memory models, TSO
5 Memory System Advanced Cache basics, DRAM basics Prefetching, replacement, RowHammer
6 Modern Interconnects PCIe basics, memory hierarchy CXL, chiplets, packaging, NVRAM
7 Accelerators GPU basics, SIMD DPUs, TPUs, FPGAs, PIM
Earlier Sections This Section (Advanced)
What is OoO? How Tomasulo’s algorithm works in hardware
2-bit saturating counters Perceptron and TAGE neural predictors
MESI state diagram Directory coherence, memory consistency models
Basic cache prefetching Stream prefetchers, Markov prefetchers, stride detection
PCIe bandwidth numbers CXL coherent interconnect, memory pooling
GPU programming basics Warp scheduling, tensor cores, GPU cache hierarchy