Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Hardware Interrupts

Introduction

Hardware interrupts are electrical signals generated by peripheral devices to request CPU attention. The mechanisms for delivering these signals — from the ancient 8259 PIC to the modern APIC architecture and PCI Message Signaled Interrupts — represent one of the most important evolution paths in PC architecture. This chapter covers the interrupt controller hardware, the Linux kernel’s IRQ domain abstraction, and the modern MSI/MSI-X interrupt delivery mechanism.

The 8259 PIC (Legacy)

The Intel 8259 Programmable Interrupt Controller was the original interrupt controller in the IBM PC. It manages 8 interrupt lines and can be cascaded (two chips) to support 15 usable IRQ lines (IRQ 2 is consumed by the cascade connection).

graph LR
    subgraph "Slave 8259"
        S[IRQ 8-15]
    end
    subgraph "Master 8259"
        M[IRQ 0-7]
    end
    S -->|IRQ 2 cascade| M
    M -->|INTR| CPU[CPU]

How it works:

  1. A device asserts a voltage on one of the IRQ lines.
  2. The 8259 sets the corresponding bit in its Interrupt Request Register (IRR).
  3. The 8259 checks if the interrupt priority is higher than the current In-Service Register (ISR) level.
  4. If so, it asserts the INTR line to the CPU.
  5. The CPU responds with INTA (Interrupt Acknowledge) pulses.
  6. The 8259 places the interrupt vector number on the data bus.
  7. The CPU reads the vector and dispatches through the IDT.

Limitations:

  • Only 15 usable IRQ lines — far too few for modern systems
  • Fixed priority scheme (lower IRQ = higher priority)
  • No support for multi-processor interrupt delivery
  • Level-triggered and edge-triggered mixing is problematic
  • No per-CPU interrupt masking (only global enable/disable via CLI)

In modern Linux kernels, the 8259 is typically emulated by the IOAPIC or handled through ACPI firmware for backward compatibility.

8259 Programming Interface

The 8259 is programmed through two I/O ports:

/* Master 8259 */
#define PIC1_CMD    0x20    /* Command port */
#define PIC1_DATA   0x21    /* Data port */

/* Slave 8259 */
#define PIC2_CMD    0xA0    /* Command port */
#define PIC2_DATA   0xA1    /* Data port */

/* Initialization Command Words (ICW) sequence: */
/* ICW1: Initialization */
outb(0x11, PIC1_CMD);  /* ICW4 needed, cascade mode, edge triggered */

/* ICW2: Vector offset */
outb(0x20, PIC1_DATA); /* Master vectors start at 0x20 (32) */

/* ICW3: Cascade identity */
outb(0x04, PIC1_DATA); /* Slave on IRQ 2 */

/* ICW4: Mode */
outb(0x01, PIC1_DATA); /* 8086 mode, non-buffered, normal EOI */

/* Masking: disable all except keyboard */
outb(0xFD, PIC1_DATA); /* IRQ 1 enabled */

End of Interrupt (EOI):

/* Send EOI to master */
outb(0x20, PIC1_CMD);

/* If interrupt came from slave, also send to slave */
outb(0x20, PIC2_CMD);

The Linux kernel disables the 8259 early in boot when it switches to APIC mode:

/* arch/x86/kernel/i8259.c */
void __init make_8259A_irq(unsigned int irq)
{
    /* ... setup 8259A for legacy interrupt ... */
}

static void mask_and_ack_8259A(struct irq_data *data)
{
    /* Mask the IRQ and send EOI */
}

The APIC Architecture

The Advanced Programmable Interrupt Controller (APIC) replaced the 8259 and is the standard interrupt delivery mechanism on x86 systems since the mid-1990s. It consists of two components:

Local APIC (LAPIC)

Each CPU core has a built-in Local APIC. It is responsible for:

  • Receiving interrupts from the IOAPIC (inter-processor interrupts, timers, error signals)
  • Accepting inter-processor interrupts (IPIs) from other CPUs’ LAPICs
  • Managing the CPU’s interrupt priority level
  • Sending End-of-Interrupt (EOI) acknowledgments

The LAPIC is memory-mapped to a configurable base address (default 0xFEE00000 on x86). Key registers include:

RegisterOffsetDescription
ID0x020LAPIC identifier
TPR0x080Task Priority Register (filters interrupts)
APR0x090Arbitration Priority Register
PPR0x0A0Processor Priority Register
EOI0x0B0End of Interrupt register (write to acknowledge)
SVR0x0F0Spurious Interrupt Vector Register
ICR0x300Interrupt Command Register (for IPIs)
LVT Timer0x320Local Vector Table — Timer
LVT LINT00x350Local Vector Table — Local Interrupt 0
LVT LINT10x360Local Vector Table — Local Interrupt 1
LVT Error0x370Local Vector Table — Error
ICR high0x310Interrupt Command Register — high 32 bits
# Read LAPIC base from the MSR
$ sudo rdmsr 0x1B
fee000900

LAPIC Timer

The LAPIC includes a built-in per-CPU timer, which is the source of the kernel’s timer tick:

/* LAPIC timer modes */
#define APIC_LVT_TIMER_PERIODIC   0x00020000  /* Periodic mode */
#define APIC_LVT_TIMER_TSCDEADLINE 0x00040000  /* TSC-deadline mode */

/* The kernel calibrates the LAPIC timer against the TSC */
/* View LAPIC timer frequency */
$ dmesg | grep -i "calibrating"
[    0.000000] Calibrating delay loop (skipped), value calculated using timer frequency.. 4800.00 BogoMIPS (lpj=2400000)

# Check which timer mode is in use
$ dmesg | grep -i "tsc-deadline"
[    0.000000] TSC deadline timer enabled

IOAPIC (I/O APIC)

The IOAPIC sits on the I/O bus (typically PCI) and routes hardware interrupt signals to the LAPIC of a target CPU. Each IOAPIC typically has 24 input lines (called “pins”), and systems can have multiple IOAPICs.

The IOAPIC contains a Redirection Table (IRT) with one entry per pin. Each entry specifies:

  • Vector: The interrupt vector number delivered to the CPU
  • Delivery Mode: Fixed, lowest priority, SMI, NMI, INIT, ExtINT
  • Destination Mode: Physical (specific APIC ID) or Logical (CPU set)
  • Polarity: Active high or active low
  • Trigger Mode: Edge-triggered or level-triggered
  • Destination: Target CPU(s) by APIC ID or logical cluster
# View IOAPIC configuration
$ cat /proc/ioapic
# Or via ACPI MADT table
$ sudo acpidump -t | grep -A 5 IOAPIC

# Detailed IOAPIC register dump
$ sudo cat /sys/kernel/debug/irq/irqs/16
handler:  handle_fasteoi_irq
status:  0x00000040 (IRQD_IRQ_STARTED)
depth:   0
irq_count: 892341
chip:    ioapic
domain:  IO-APIC

IOAPIC Redirection Table Entries

Each Redirection Table Entry (RTE) is 64 bits wide:

struct ioapic_rte {
    union {
        struct {
            u32 vector       : 8;   /* Vector 0x10-0xFE */
            u32 delivery     : 3;   /* 0=fixed, 1=lowest, 2=SMI, 4=NMI, 5=INIT */
            u32 dest_mode    : 1;   /* 0=physical, 1=logical */
            u32 delivery_stat: 1;   /* Read-only: delivery status */
            u32 polarity     : 1;   /* 0=high, 1=low */
            u32 irr          : 1;   /* Read-only: remote IRR */
            u32 trigger      : 1;   /* 0=edge, 1=level */
            u32 mask         : 1;   /* 0=enabled, 1=masked */
            u32 reserved     : 15;
        };
        u32 low;
        u32 high;    /* Destination field (APIC ID or logical cluster) */
    };
};

Interrupt Delivery Flow (APIC)

sequenceDiagram
    participant DEV as PCI Device
    participant IOAPIC as IOAPIC
    participant LAPIC as Local APIC
    participant CPU as CPU Core

    DEV->>IOAPIC: Assert IRQ line (edge/level)
    IOAPIC->>IOAPIC: Look up Redirection Table entry
    IOAPIC->>LAPIC: APIC message (vector, delivery mode)
    LAPIC->>LAPIC: Check priority vs TPR/PPR
    LAPIC->>CPU: Deliver interrupt
    CPU->>CPU: Save context, vector through IDT
    CPU->>LAPIC: Write EOI
    LAPIC->>IOAPIC: EOI broadcast (level-triggered)

APIC Priority Mechanism

The LAPIC uses a priority scheme to determine whether an interrupt should be delivered:

/* Priority comparison */
/* Interrupt priority = vector / 16 */
/* Current priority = PPR (Processor Priority Register) */
/* Interrupt is delivered only if its priority > PPR */

/* TPR (Task Priority Register) can be used to block low-priority interrupts */
/* Setting TPR = 0x20 blocks vectors 0x20-0x2F */

The Linux kernel uses TPR sparingly — it’s primarily relevant for virtualization (where the hypervisor manages guest interrupt priorities).

Inter-Processor Interrupts (IPIs)

IPIs are interrupts sent from one CPU to another (or to all CPUs) via the LAPIC’s Interrupt Command Register (ICR):

/* ICR fields for sending IPIs */
#define APIC_ICR_DM_FIXED    0x00000000  /* Fixed delivery mode */
#define APIC_ICR_DM_LOWEST   0x00000100  /* Lowest priority */
#define APIC_ICR_DM_SMI      0x00000200  /* SMI */
#define APIC_ICR_DM_NMI      0x00000400  /* NMI */
#define APIC_ICR_DM_INIT     0x00000500  /* INIT */
#define APIC_ICR_DM_SIPI     0x00000600  /* Start-up IPI */
#define APIC_ICR_DEST_SELF   0x00040000  /* Send to self */
#define APIC_ICR_DEST_ALL    0x00080000  /* Send to all */
#define APIC_ICR_DEST_NOTSELF 0x000C0000 /* Send to all except self */

Common IPI types in Linux:

IPI VectorPurposeTriggered By
RESCHEDULE_VECTORForce reschedulesmp_send_reschedule()
CALL_FUNCTION_VECTORExecute function on remote CPUsmp_call_function()
CALL_FUNCTION_SINGLE_VECTORExecute on specific CPUsmp_call_function_single()
REBOOT_VECTORForce rebootsmp_send_stop()
IRQ_WORK_VECTORProcess irq_work itemsirq_work_queue()
X86_PLATFORM_IPI_VECTORPlatform-specific (e.g., thermal)Various
UINTR_VECTORUser interruptsuintr subsystem
# View IPI statistics in /proc/interrupts
$ grep -E 'RES|CAL|TLB|TRM' /proc/interrupts
RES:   1234  2345  3456  4567   Rescheduling interrupts
CAL:   5678  6789  7890  8901   Function call interrupts
TLB:   1234  2345  3456  4567   TLB shootdowns

MSI and MSI-X

Message Signaled Interrupts (MSI) were introduced with PCI 2.2 and represent a fundamental departure from the pin-based interrupt model. Instead of asserting a physical IRQ line, the device writes a special message to a memory-mapped address — the LAPIC’s interrupt command register.

MSI

  • Supports 1, 2, 4, 8, 16, or 32 interrupt vectors per device
  • Each vector is a separate interrupt with its own handler
  • The message address encodes the destination CPU APIC ID
  • The message data encodes the vector number and delivery mode

Advantages over pin-based interrupts:

  • No shared IRQ lines — eliminates the “interrupt storm” problem
  • No need for IRQ routing through IOAPIC — lower latency
  • Each MSI vector can target a specific CPU — enables perfect multi-queue scaling
  • No lost interrupts (race-free acknowledgment via the write transaction)

MSI-X

MSI-X (PCI 3.0) extends MSI with:

  • Up to 2048 vectors per device (vs MSI’s 32 maximum)
  • Each vector can be independently routed to a different CPU
  • Table entries can be in either BAR-mapped memory or I/O space
  • More flexible — each vector has its own message address and data

A typical high-performance NVMe drive or network card uses MSI-X with one vector per hardware queue, pinned to the CPU that processes that queue:

$ cat /proc/interrupts | grep nvme
120:  452108  0  0  0  PCI-MSI  524289-edge  nvme0q1
121:  0  387654  0  0  PCI-MSI  524290-edge  nvme0q2
122:  0  0  298123  0  PCI-MSI  524291-edge  nvme0q3
123:  0  0  0  401567  PCI-MSI  524292-edge  nvme0q4

Each NVMe queue pair is handled by a dedicated MSI-X vector on the CPU closest to the NUMA node of the queue.

MSI-X Configuration Space

Each MSI-X table entry consists of 12 bytes:

Offset 0x00: Message Address (32 bits)
Offset 0x04: Message Upper Address (32 bits, for 64-bit addressing)
Offset 0x08: Message Data (32 bits)
Offset 0x0A: Vector Control (32 bits, bit 0 = mask bit)
# Inspect MSI-X capability of a device
$ sudo lspci -vvv -s 00:04.0
Capabilities: [b0] MSI-X: Enable+ Count=32 Masked-
        Vector table: BAR=0 offset=00000000
        PBA: BAR=0 offset=00001000

MSI Capability Structure (PCI Config Space)

The MSI capability structure in PCI configuration space:

/* MSI Capability Structure */
struct msi_cap {
    u8  cap_id;         /* 0x05 = MSI */
    u8  next;           /* Next capability pointer */
    u16 msg_control;    /* Message control */
    /* Bits 0-2: Multiple Message Enable (log2 vectors) */
    /* Bit 7: 64-bit address capable */
    u32 msg_addr_lo;    /* Message address low 32 bits */
    u32 msg_addr_hi;    /* Message address high 32 bits (if 64-bit) */
    u16 msg_data;       /* Message data */
    /* ... */
};

/* MSI-X Capability Structure */
struct msix_cap {
    u8  cap_id;         /* 0x11 = MSI-X */
    u8  next;           /* Next capability pointer */
    u16 msg_control;    /* Message control */
    /* Bits 0-10: Table size (N-1, where N = number of vectors) */
    /* Bit 14: Function mask */
    u32 table_offset;   /* BAR and offset for vector table */
    u32 pba_offset;     /* BAR and offset for pending bit array */
};

MSI Interrupt Flow

sequenceDiagram
    participant NIC as Network Card (MSI-X)
    participant MEM as Memory Write
    participant LAPIC as Local APIC
    participant CPU as CPU Core

    NIC->>MEM: Write vector to LAPIC address
    Note over MEM: 0xFEE000xx address
    MEM->>LAPIC: Memory-mapped write triggers interrupt
    LAPIC->>LAPIC: Decode vector from data
    LAPIC->>CPU: Deliver interrupt
    CPU->>CPU: IDT lookup, handler execution
    CPU->>LAPIC: EOI (automatic for MSI)

Key advantage: MSI delivery is a simple memory write — no acknowledgment protocol, no shared lines, no race conditions.

Kernel MSI Allocation

/* Allocate MSI-X vectors */
int pci_alloc_irq_vectors(struct pci_dev *dev,
                          unsigned int min_vectors,
                          unsigned int max_vectors,
                          unsigned int flags);

/* Example: allocate 4 MSI-X vectors for a NIC */
ret = pci_alloc_irq_vectors(pdev, 4, 4, PCI_IRQ_MSIX);
if (ret < 0)
    return ret;

/* Get IRQ number for vector 0 */
irq = pci_irq_vector(pdev, 0);

/* Request IRQ for each vector */
for (i = 0; i < num_queues; i++) {
    irq = pci_irq_vector(pdev, i);
    request_irq(irq, my_handler, 0, "nic", &queues[i]);
}

IRQ Domains

The IRQ domain abstraction (introduced in Linux 3.3) provides a clean mapping between hardware interrupt numbers and Linux virtual IRQ numbers (virq). This is essential because:

  1. Different interrupt controllers (IOAPIC, GIC on ARM, GPIO controllers) use different numbering schemes.
  2. A system may have multiple interrupt controllers.
  3. Hardware numbers can collide across controllers.

IRQ Domain Architecture

graph TD
    subgraph "Hardware Layer"
        IOAPIC[IOAPIC: hwirq 0-23]
        GPIO[GPIO Controller: hwirq 0-31]
        MSI[MSI: hwirq per vector]
    end
    subgraph "IRQ Domain Layer"
        D1[IOAPIC Domain]
        D2[GPIO Domain]
        D3[MSI Domain]
    end
    subgraph "Linux IRQ Numbers"
        V1[virq 16-39]
        V2[virq 40-71]
        V3[virq 72-103]
    end
    IOAPIC --> D1
    GPIO --> D2
    MSI --> D3
    D1 --> V1
    D2 --> V2
    D3 --> V3

Key Data Structures

/* Representation of an IRQ domain */
struct irq_domain {
    struct list_head link;
    const char *name;
    const struct irq_domain_ops *ops;
    void *host_data;
    unsigned int flags;
    unsigned int mapcount;
    struct fwnode_handle *fwnode;
    enum irq_domain_bus_token bus_token;
    struct irq_domain_chip_generic *gc;
    /* Radix tree mapping hwirq -> irq_desc */
    struct radix_tree_root revmap_tree;
    unsigned int revmap_size;
    struct irq_desc **revmap;
};

/* Maps a hardware IRQ to a Linux IRQ */
unsigned int irq_create_mapping(struct irq_domain *domain,
                                irq_hw_number_t hwirq);

IRQ Domain Operations

Each interrupt controller provides an irq_domain_ops structure:

struct irq_domain_ops {
    int (*map)(struct irq_domain *d, unsigned int virq,
               irq_hw_number_t hw);       /* Map hwirq to virq */
    void (*unmap)(struct irq_domain *d, unsigned int virq);
    int (*xlate)(struct irq_domain *d, struct device_node *node,
                 const u32 *intspec, unsigned int intsize,
                 unsigned long *out_hwirq,
                 unsigned int *out_type);  /* DT/ACPI translation */
    /* ... */
};

Domain Types

TypeDescriptionUsed By
irq_domain_linearArray-based lookup, O(1)IOAPIC, GIC
irq_domain_treeRadix tree, sparseGPIO controllers
irq_domain_nomapNo virq mapping, direct hwirqSome legacy controllers
irq_domain_hierarchyStacked domains (chained controllers)Modern complex topologies

Stacked IRQ Domains

Modern systems often have hierarchical interrupt controllers. The kernel models this with stacked domains:

graph TD
    subgraph "Hardware"
        GPIO[GPIO Controller]
        GIC[GICv3 Distributor]
    end
    subgraph "IRQ Domain Stack"
        D_GPIO[GPIO Domain]
        D_GIC[GIC Domain]
        D_PLATFORM[Platform Domain]
    end
    subgraph "Linux IRQ"
        VIRQ[virq 50]
    end
    GPIO --> D_GPIO
    D_GPIO --> D_GIC
    D_GIC --> D_PLATFORM
    D_PLATFORM --> VIRQ
/* Stacked domain allocation */
struct irq_domain *gpio_domain = irq_domain_create_hierarchy(
    parent_domain,      /* GIC domain */
    IRQ_DOMAIN_FLAG_HIERARCHY,
    nr_irqs,
    ops,
    host_data);

Trigger Modes: Edge vs Level

Hardware interrupts use two electrical signaling modes:

Edge-Triggered

The interrupt is signaled by a transition (low-to-high or high-to-low). The controller latches the edge and the line can return to its idle state. The CPU must detect the transition even if it’s brief.

  • Used by: legacy ISA interrupts, MSI/MSI-X (always edge)
  • Advantage: no need for explicit acknowledgment of the line state
  • Disadvantage: can miss interrupts if the CPU is not ready

Level-Triggered

The interrupt is signaled by a voltage level (high or low). The line remains asserted until the device is explicitly told to de-assert it.

  • Used by: most PCI interrupts (legacy INTx), IOAPIC pins
  • Advantage: cannot miss interrupts — the line stays asserted
  • Disadvantage: requires explicit EOI and device de-assert; shared lines cause “interrupt storms” if a handler fails to acknowledge

In the IOAPIC redirection table, the trigger mode is encoded per entry:

#define IRQ_TYPE_NONE           0x00000000
#define IRQ_TYPE_EDGE_RISING    0x00000001
#define IRQ_TYPE_EDGE_FALLING   0x00000002
#define IRQ_TYPE_EDGE_BOTH      (IRQ_TYPE_EDGE_FALLING | IRQ_TYPE_EDGE_RISING)
#define IRQ_TYPE_LEVEL_HIGH     0x00000004
#define IRQ_TYPE_LEVEL_LOW      0x00000008

Edge vs Level: Handling Differences

sequenceDiagram
    participant DEV as Device
    participant IRQ as IRQ Line
    participant CTRL as Controller
    participant CPU as CPU

    Note over DEV,CPU: Edge-Triggered
    DEV->>IRQ: Assert (pulse)
    IRQ->>CTRL: Rising edge detected
    DEV->>IRQ: De-assert (line returns to idle)
    CTRL->>CPU: Deliver interrupt
    CPU->>CPU: Handle interrupt
    Note over DEV,CPU: No EOI needed for line

    Note over DEV,CPU: Level-Triggered
    DEV->>IRQ: Assert (held high)
    IRQ->>CTRL: Level detected
    CTRL->>CPU: Deliver interrupt
    CPU->>CPU: Handle interrupt
    CPU->>DEV: Acknowledge → device de-asserts
    DEV->>IRQ: De-assert
    CPU->>CTRL: EOI

Level-Triggered Shared IRQ Problem

sequenceDiagram
    participant DEV_A as Device A
    participant DEV_B as Device B
    participant IRQ as Shared IRQ Line
    participant HANDLER_A as Handler A
    participant HANDLER_B as Handler B

    DEV_A->>IRQ: Assert (level high)
    IRQ->>HANDLER_A: IRQ fires
    HANDLER_A->>HANDLER_A: Check device A status
    alt Device A caused interrupt
        HANDLER_A->>DEV_A: Acknowledge
        DEV_A->>IRQ: De-assert
    else Device A did NOT cause interrupt
        HANDLER_A->>HANDLER_A: Return IRQ_NONE
    end
    IRQ->>HANDLER_B: IRQ fires (shared line)
    HANDLER_B->>HANDLER_B: Check device B status
    Note over IRQ: If neither handler acknowledges...
    Note over IRQ: Interrupt storm! Line stays asserted.

Interrupt Controllers Beyond x86

ARM GIC (Generic Interrupt Controller)

ARM systems use the GIC architecture (GICv2, GICv3, GICv4):

  • Distributor (GICD): Routes interrupts to CPU interfaces
  • Redistributor (GICR): Per-PE (Processing Element) component in GICv3+
  • CPU Interface (GICC): Per-CPU interrupt acknowledgment and priority management
  • ITS (Interrupt Translation Service): MSI-like support for GICv3+

GICv3 supports:

  • Up to 1020 SPIs (Shared Peripheral Interrupts)
  • 32 SGIs (Software Generated Interrupts) for IPI
  • 16 PPIs (Private Peripheral Interrupts) per CPU
  • Direct injection of virtual interrupts for VMs (GICv4)
/* GIC interrupt types */
#define GIC_SGI    0   /* Software Generated Interrupt (IPI) */
#define GIC_PPI    1   /* Private Peripheral Interrupt (per-CPU timer, etc.) */
#define GIC_SPI    2   /* Shared Peripheral Interrupt (device interrupts) */
#define GIC_ESPI   3   /* Extended SPI (GICv4, up to 960 more) */

GICv3 Architecture

graph TD
    subgraph "CPU 0"
        RD0[Redistributor 0]
        ICC0[CPU Interface 0]
    end
    subgraph "CPU 1"
        RD1[Redistributor 1]
        ICC1[CPU Interface 1]
    end
    subgraph "Distributor"
        GICD[GICD]
    end
    subgraph "ITS"
        ITS[ITS - MSI translation]
    end
    GICD --> RD0
    GICD --> RD1
    RD0 --> ICC0
    RD1 --> ICC1
    ITS --> GICD

GICv3 System Register Interface

GICv3 uses system registers (not memory-mapped) for the CPU interface:

/* Reading ICC (Interrupt Controller CPU interface) registers */
static inline u64 gic_read_iar(void)
{
    u64 irq;
    asm volatile("mrs %0, ICC_IAR1_EL1" : "=r" (irq));
    return irq;
}

static inline void gic_write_eoir(u64 irq)
{
    asm volatile("msr ICC_EOIR1_EL1, %0" :: "r" (irq));
}

RISC-V PLIC/PLIC

RISC-V uses the Platform-Level Interrupt Controller (PLIC) for external interrupts and the Core-Local Interrupt Controller (CLINT) for timer and IPI interrupts.

/* RISC-V interrupt types */
#define RISCV_SMODE_SOFT_IRQ    1   /* Supervisor software interrupt */
#define RISCV_SMODE_TIMER_IRQ   5   /* Supervisor timer interrupt */
#define RISCV_SMODE_EXT_IRQ     9   /* Supervisor external interrupt */
graph TD
    subgraph "RISC-V System"
        CLINT["CLINT<br>Timer + IPI"]
        PLIC["PLIC<br>External interrupts"]
        CORE0[Core 0]
        CORE1[Core 1]
    end
    CLINT -->|Timer/IPI| CORE0
    CLINT -->|Timer/IPI| CORE1
    PLIC -->|External IRQs| CORE0
    PLIC -->|External IRQs| CORE1

ACPI and Interrupt Routing

On x86, the ACPI firmware provides interrupt routing information through several tables:

  • MADT (Multiple APIC Description Table): Lists all interrupt controllers (LAPIC, IOAPIC, etc.)
  • DSDT/SSDT: Device-specific interrupt routing via _CRS (Current Resource Settings) and _PRS (Possible Resource Settings) methods
  • IRQT/PCI Routing Table: Maps PCI interrupt pins (INTA-INTD) to IOAPIC input pins
# Dump ACPI MADT table
$ sudo acpidump -b -t MADT | xxd | head -20

# Or use the decoded tables
$ sudo dmidecode | grep -i interrupt

# View ACPI interrupt source overrides
$ dmesg | grep -i "ACPI:.*IRQ"
[    0.000000] ACPI: INT_SRC_OVR (bus 0 bus_irq 9 global_irq 9 high level)
[    0.000000] ACPI: INT_SRC_OVR (bus 0 bus_irq 0 global_irq 2 high edge)

ACPI Interrupt Source Override

The MADT table can contain Interrupt Source Override entries that remap ISA IRQs to different IOAPIC pins:

struct acpi_madt_interrupt_override {
    struct acpi_subtable_header header;
    u8 bus;           /* 0 = ISA */
    u8 source_irq;    /* ISA IRQ number */
    u32 global_irq;   /* IOAPIC pin number */
    u16 flags;        /* MPS INTI flags */
};

Common override: ISA IRQ 0 (timer) is remapped to IOAPIC pin 2 on most systems:

$ dmesg | grep "ACPI.*timer"
[    0.000000] ACPI: INT_SRC_OVR (bus 0 bus_irq 0 global_irq 2 high edge)

Kernel Internals: irq_chip and irq_domain

The Linux kernel models each interrupt controller as an irq_chip:

struct irq_chip {
    struct device   *parent_device;
    const char      *name;
    void            (*irq_enable)(struct irq_data *data);
    void            (*irq_disable)(struct irq_data *data);
    void            (*irq_ack)(struct irq_data *data);
    void            (*irq_mask)(struct irq_data *data);
    void            (*irq_unmask)(struct irq_data *data);
    void            (*irq_eoi)(struct irq_data *data);
    int             (*irq_set_affinity)(struct irq_data *data,
                                        const struct cpumask *dest,
                                        bool force);
    int             (*irq_set_type)(struct irq_data *data,
                                    unsigned int flow_type);
    /* ... */
};

The irq_data structure ties it all together:

struct irq_data {
    unsigned int        irq;        /* Linux IRQ number */
    unsigned long       hwirq;      /* Hardware IRQ number */
    struct irq_common_data *common;
    struct irq_chip     *chip;      /* Interrupt controller */
    struct irq_domain   *domain;    /* IRQ domain */
    void                *chip_data; /* Controller-private data */
};

irq_chip Examples

/* IOAPIC irq_chip */
static struct irq_chip ioapic_chip = {
    .name           = "IO-APIC",
    .irq_enable     = ioapic_enable,
    .irq_disable    = ioapic_disable,
    .irq_ack        = ioapic_ack_irq,
    .irq_mask       = ioapic_mask_irq,
    .irq_unmask     = ioapic_unmask_irq,
    .irq_eoi        = ioapic_eoi,
    .irq_set_affinity = ioapic_set_affinity,
    .irq_set_type   = ioapic_set_type,
};

/* MSI irq_chip */
static struct irq_chip msi_chip = {
    .name           = "PCI-MSI",
    .irq_enable     = unmask_msi_irq,
    .irq_disable    = mask_msi_irq,
    .irq_ack        = ack_apic_edge,
    .irq_mask       = mask_msi_irq,
    .irq_unmask     = unmask_msi_irq,
    .irq_set_affinity = msi_set_affinity,
};

Practical Example: Tracing IRQ Routing

# Show all IRQ chip information
$ sudo cat /proc/irq/*/chip_name 2>/dev/null
IO-APIC
IO-APIC
PCI-MSI
PCI-MSI

# Detailed IRQ info via debugfs
$ sudo ls /sys/kernel/debug/irq/irqs/
0  1  8  9  16  23  120  121  122  123 ...

$ sudo cat /sys/kernel/debug/irq/irqs/120
handler:  handle_edge_irq
status:  0x00000040 (IRQD_IRQ_STARTED)
depth:   0
irq_count: 452108
chip:    pci-msi
domain:  PCI-MSI

Viewing IRQ Domain Hierarchy

# List all IRQ domains
$ sudo cat /sys/kernel/debug/irq_domain/domains
Domain  Name                 Parent
------  ----                 ------
0       IO-APIC              (root)
1       IO-APIC-2            IO-APIC
2       PCI-MSI              (root)
3       GPIO-0               (root)

# View domain mappings
$ sudo cat /sys/kernel/debug/irq_domain/mappings
hwirq    virq   chip          domain
------   ----   ----          ------
0        0      timer         IO-APIC
1        1      i8042         IO-APIC
16       16     ehci_hcd      IO-APIC
524289   120    nvme          PCI-MSI

Interrupt Storm Detection and Handling

An interrupt storm occurs when an interrupt fires repeatedly without being properly handled, consuming all CPU time:

# Symptoms: 100% CPU usage in top, system unresponsive
# /proc/interrupts shows rapidly incrementing counter for one IRQ

# Monitor interrupt rate
$ watch -n1 'cat /proc/interrupts | grep -E "^[0-9]+:" | awk "{print \$1, \$2}"'

# Disable a problematic IRQ
$ echo 1 > /proc/irq/16/spurious_count  # Not real API
# Real approach: unbind the driver
$ echo 0 > /proc/irq/16/smp_affinity

# Kernel spurious interrupt handling
$ dmesg | grep -i spurious
[  123.456789] spurious 8259A interrupt: IRQ7.

Kernel Spurious IRQ Detection

The kernel tracks unhandled interrupts and automatically disables IRQs that appear spurious:

/* kernel/irq/spurious.c */
#define MAX_INTERRUPTS  100000

static void note_interrupt(struct irq_desc *desc, irqreturn_t retval)
{
    if (retval == IRQ_WAKE_THREAD) {
        /* Threaded handler — handled */
        desc->threads_handled++;
        return;
    }

    if (retval == IRQ_NONE) {
        desc->irqs_unhandled++;
        /* If too many unhandled, mark as spurious */
        if (desc->irqs_unhandled > 99900 &&
            time_after(jiffies, desc->last_unhandled + HZ/10)) {
            /* Disable the IRQ */
            desc->irq_count++;
            if (desc->irq_count > 10)
                __report_bad_irq(desc);
        }
    }
}

References