How Flush+Reload Attacks Exploit Shared CPU Cache Lines
Flush+Reload attacks measure L3 cache hit latencies with clflush and rdtscp. Learn how shared memory pages and timing thresholds expose cryptographic keys.

Modern operating systems conserve physical DRAM by mapping shared read-only files into the address spaces of multiple processes. When two independent user applications link against the same dynamic library, such as libc.so or libcrypto.so, the kernel assigns distinct virtual addresses to the same underlying physical memory frames. The processor manages this shared memory transparently through hardware-managed page tables, enforcing process isolation by marking user pages as unalterable without triggering a copy-on-write fault.
This shared mapping creates an exploitable physical overlap in the CPU cache hierarchy. Shared physical frames reside within the same set of cache lines regardless of which process requests them. Flush+Reload attacks exploit this shared physical mapping to monitor victim memory access patterns with single-cache-line granularity. By measuring the elapsed execution cycles of cache management and serialized read instructions, an unprivileged spy process determines whether an isolated victim process touched a specific 64-byte segment of code or static data.
Unlike contention-based attacks that monitor entire cache sets by flooding them with dummy lines, this attack operates on precise virtual addresses mapped to shared pages. An attacker needs no administrative capabilities, kernel extensions, or inter-process communication permissions. The timing discrepancy between an on-chip Last-Level Cache hit and an off-chip main memory access provides a high-signal channel that leaks cryptographic operations, keystrokes, and control flow decisions across sandbox boundaries.
How Modern CPU Cache Hierarchies Handle Shared Physical Pages
Modern multi-core x86 processors split memory access into a hierarchy: private Level 1 (L1) and Level 2 (L2) caches dedicated to each physical core, backed by a unified, shared Last-Level Cache (LLC, or L3). L1 caches are split into separate 32 KiB or 48 KiB instruction (L1I) and data (L1D) caches with single-cycle or four-cycle latencies. L2 caches typically hold 512 KiB to 2 MiB per core with roughly 12 to 14 cycles of latency. L3 caches are shared across all cores on a silicon die, scaling from 16 MiB to over 96 MiB with latencies ranging from 38 to 65 cycles.
Cache memory is organized into 64-byte blocks called cache lines. For a physical address on an x86-64 CPU, the lowest 6 bits (bits 0 through 5) define the byte offset within the 64-byte line. The next set of bits defines the cache set index, selecting which set within the N-way set-associative cache array holds the data. The remaining upper bits form the address tag. When a core issues a memory read, the memory execution unit checks the tag across all ways in the selected set.
+------------------------------------+------------------+---------------+
| Tag Bits | Set Index Bits | Line Offset |
| (Bits 63 down to 18) | (Bits 17 down 6) | (Bits 5 to 0) |
+------------------------------------+------------------+---------------+
|<------------------------ Physical Address --------------------------->|
When an application loads a shared library via mmap(), the Linux page cache allocates physical pages backed by the filesystem. If Process A and Process B both map /usr/lib/x86_64-linux-gnu/libcrypto.so.3, their distinct Page Table Entries (PTEs) store identical Physical Frame Numbers (PFNs). Attackers abusing kernel structure layouts often exploit shared file pointers in similar ways, as seen when How Linux Dirty Cred Exploits Bypass Kernel Mitigations demonstrates swapping unprivileged credentials for privileged ones within common kernel structs.
Because L3 caches on Intel processors are physically indexed and physically tagged (PIPT), identical physical addresses map to the exact same L3 cache set and slice, regardless of which core or virtual address generated the request. Most Intel desktop and server processors implement inclusive L3 caches: any line resident in a core's private L1 or L2 cache must also reside in the shared L3 cache. If an instruction evicts a line from the L3 cache, the cache coherence controller broadcasts back-invalidation requests across the internal ring bus or mesh interconnect, clearing that line from every core's private L1 and L2 caches. This inclusive property provides the mechanical basis for eviction-driven side channels.
How Flush+Reload Attacks Exploit Inclusive Cache Hierarchies
A Flush+Reload attack cycle consists of three sequential phases: Flush, Wait, and Reload. The spy process identifies a target memory address inside a shared library that the victim process accesses during a sensitive computation.
In the Flush phase, the spy process executes the clflush (Cache Line Flush) or clflushopt (Optimized Cache Line Flush) instruction on the target virtual address. The processor converts the virtual address to its physical frame, locates the corresponding line in the cache hierarchy, and invalidates it. The line is evicted from L1D, L1I, L2, and L3 caches across all cores on the socket. If the line contains modified data, it is written back to DRAM before invalidation, though shared library pages are read-only and require no writeback.
[Attacker Core] [Victim Core]
| |
1. clflush(addr) -----------------------------+ |
Evicts addr from L1, L2, L3 | |
| v |
2. Wait interval ----------------------> Executes code?
| | |
| | (If accessed:
| | addr fetched to L3/L1)
| v |
3. rdtscp + read(addr) + rdtscp <-------------+-------+
Timing < Threshold -> Victim accessed addr
Timing > Threshold -> Victim idle
During the Wait phase, the spy process yields the processor or executes a calibrated delay loop. This interval provides a time window for the victim process to execute. If the victim triggers a code branch or accesses a data lookup table located within the flushed 64-byte memory line, the victim's core issues a read request to memory. The hardware fetches the missing 64-byte block from main DRAM, installing the line into the shared L3 cache and the victim core's private L1 cache. If the victim does not execute that path, the target line remains absent from all cache levels.
In the Reload phase, the spy re-accesses the target virtual address while measuring the exact number of clock cycles consumed by the read operation. If the victim accessed the target line during the Wait window, the line already resides in the shared L3 cache; the reload completes within 40 to 65 cycles. If the victim did not touch the line, the reload must traverse the memory controller to fetch the data from off-chip DRAM, consuming 180 to 320 cycles. By evaluating this timing differential against a predetermined threshold, the spy detects victim activity with near-zero error rates.
Calibrating Cycle Thresholds and the Measurement Primitive
Precise timing requires instruction serialization. Modern out-of-order processors aggressively execute memory instructions ahead of preceding instructions if no register dependencies exist. An unconstrained cycle count read using the raw rdtsc instruction will yield corrupted timings because the CPU can execute the memory load before the first timestamp read or after the second timestamp read.
The rdtscp instruction resolves this issue by reading the 64-bit Time Stamp Counter into the EDX:EAX registers while simultaneously reading the IA32_TSC_AUX register into ECX. Critically, rdtscp is partially serializing: it waits until all previous instructions in program order have executed and retired before sampling the counter. To prevent subsequent instructions from executing prematurely, a full memory barrier (mfence) or execution barrier (cpuid) locks pipeline retirement.
The following C implementation implements the core measurement primitive:
#include
#include
static inline void flush_line(const void *addr) {
asm volatile("clflush 0(%0)" : : "r"(addr) : "memory");
}
static inline uint64_t measure_read_latency(const void *addr) {
uint32_t aux;
uint64_t start_cycles, end_cycles;
volatile uint8_t *target = (const volatile uint8_t *)addr;
// Serialize pipeline and read timestamp
start_cycles = __rdtscp(&aux);
asm volatile("mfence" ::: "memory");
// Force memory load
(void)*target;
// Serialize pipeline and capture final timestamp
end_cycles = __rdtscp(&aux);
asm volatile("mfence" ::: "memory");
return end_cycles - start_cycles;
}
Establishing the discrimination threshold requires empirical calibration on the target host before launching an extraction loop. The calibration routine allocates a buffer, repeatedly flushes it, and records the cycle counts of uncached accesses. Next, it accesses the buffer to prime the cache, immediately measures the access latency, and records the cycle counts of cached accesses.
#define SAMPLES 10000
uint64_t calibrate_threshold(const void *sample_addr) {
uint64_t hit_cycles = 0;
uint64_t miss_cycles = 0;
for (int i = 0; i < SAMPLES; i++) {
// Measure Cache Miss
flush_line(sample_addr);
miss_cycles += measure_read_latency(sample_addr);
// Measure Cache Hit
(void)*(volatile uint8_t *)sample_addr;
hit_cycles += measure_read_latency(sample_addr);
}
uint64_t avg_hit = hit_cycles / SAMPLES;
uint64_t avg_miss = miss_cycles / SAMPLES;
// Set threshold between average hit and average miss
return avg_hit + ((avg_miss - avg_hit) / 3);
}
Latency (Cycles)
0 40 80 120 160 200 240 280 320
[ L1/L2 ]
[ L3 Hit ]
|
THRESHOLD (105)
|
[ DRAM Miss ]
On an Intel Core i7-11700K running at a fixed 3.6 GHz clock speed, typical distributions show L3 cache hits clustering tightly between 42 and 58 cycles. DRAM accesses cluster between 190 and 260 cycles. A cycle threshold of 100 to 110 cycles cleanly bifurcates the states.
Memory page configurations influence these latencies. When virtualization platforms configure nested page tables with transparent huge pages, Translation Lookaside Buffer (TLB) misses add variable page walk latencies that broaden the distribution curves. Memory layout patterns and paging structures impact these walks directly, as analyzed in Why Transparent Huge Pages on a VPS Degrade Memory Latency, shifting cache miss latencies from 200 cycles to over 350 cycles when TLB misses cascade into page table walks.
Extracting Secret Keys From Cryptographic Lookup Tables
The primary historical targets for Flush+Reload attacks are table-based implementations of the Advanced Encryption Standard (AES) and modular exponentiation routines in RSA. In legacy OpenSSL releases (such as OpenSSL 0.9.8 through 1.0.1e), AES encryption was optimized for 32-bit processors using precomputed lookup tables termed "T-tables."
Standard AES-128 operates on a 4x4 state array of 16 bytes over 10 algorithmic rounds. The round function combines byte substitution (SubBytes), row shifting (ShiftRows), column mixing (MixColumns), and round key addition (AddRoundKey). To accelerate this sequence, developers combined these four transformations into four static 32-bit tables: Te0, Te1, Te2, and Te3. Each table contains 256 4-byte entries, occupying exactly 1,024 bytes (1 KiB) of memory:
$$\text{Table Size} = 256 \times 4\text{ bytes} = 1024\text{ bytes}$$
Because a single cache line spans 64 bytes, each T-table occupies:
$$\frac{1024\text{ bytes}}{64\text{ bytes/line}} = 16\text{ cache lines}$$
Each cache line holds 16 consecutive table entries:
$$\frac{64\text{ bytes}}{4\text{ bytes/entry}} = 16\text{ entries/line}$$
In Round 1 of AES, the state byte input to the table lookup is the direct XOR combination of the known plaintext byte $P_i$ and the unknown secret encryption key byte $K_i$. The lookup computation for the first column is:
$$S_0 = Te0[P_0 \oplus K_0] \oplus Te1[P_5 \oplus K_5] \oplus Te2[P_{10} \oplus K_{10}] \oplus Te3[P_{15} \oplus K_{15}]$$
The index into table Te0 is index $j = P_0 \oplus K_0$. The line number $L$ in memory containing entry $j$ is calculated by:
$$L = \lfloor \frac{j}{16} \rfloor = \lfloor \frac{P_0 \oplus K_0}{16} \rfloor = (P_0 \oplus K_0) \gg 4$$
The memory line accessed depends directly on the upper 4 bits of the intermediate state $P_0 \oplus K_0$.
To execute the key recovery attack, the spy process maps libcrypto.so and calculates the virtual address of the 16 cache lines comprising Te0:
Line 0: Te0[0] through Te0[15] (Offsets 0x00 to 0x3F)
Line 1: Te0[16] through Te0[31] (Offsets 0x40 to 0x7F)
Line 2: Te0[32] through Te0[47] (Offsets 0x80 to 0xBF)
...
Line 15: Te0[240] through Te0[255] (Offsets 0x3C0 to 0x3FF)
The attack proceeds across thousands of encryption operations where the attacker knows the plaintexts $P$:
- The spy flushes line $L_m$ of table
Te0usingclflush. - The spy triggers the victim process to encrypt a known plaintext $P$ (or waits for an unprivileged network request to prompt encryption).
- The spy executes the reload measurement on line $L_m$.
- If the measured latency falls below the cycle threshold, a cache hit occurred: the victim must have accessed an entry inside line $L_m$.
A recorded hit on line $L_m$ establishes the mathematical identity:
$$(P_0 \oplus K_0) \gg 4 = m$$
$$m \times 16 \le (P_0 \oplus K_0) \le (m \times 16) + 15$$
Since $P_0$ is known, this single hit restricts the upper 4 bits of secret key byte $K_0$ to:
$$K_0 \gg 4 = m \oplus (P_0 \gg 4)$$
A single observation eliminates 15 of every 16 candidate values for the upper nibble of $K_0$. Repeating this measurement over several thousand plaintexts aggregates a histogram of hits across all 16 cache lines. Because the lower 4 bits of the index create noise across the other three T-table accesses in subsequent rounds, the attacker performs a correlation attack over the collected samples. The correlation peak identifies the true upper 4 bits of all 16 key bytes with mathematical certainty within 10,000 to 50,000 encryptions. The remaining lower 4 bits per byte ($2^{64}$ total space) can be solved via standard key schedule relationships and brute force in seconds.
Kernel and Hardware Mitigations Against Cache Line Probing
Mitigating Flush+Reload attacks requires removing shared memory, eliminating timing variances, or restricting hardware flushing primitives.
The most effective software defense is constant-time cryptographic engineering. Modern cryptographic libraries, including OpenSSL 3.x, BoringSSL, and libsodium, eliminate data-dependent memory lookups entirely. Software AES implementations replace T-tables with bit-sliced implementations that evaluate the Galois field substitutions ($S\text{-box}$) using logical operations (AND, XOR, SHIFT) over CPU vector registers. Because every instruction executes in fixed cycle counts and loads no table data, memory access patterns convey zero entropy regarding secret keys.
At the instruction set level, hardware-accelerated cryptographic extensions remove software table implementations entirely. Intel and AMD CPUs provide the AES-NI instruction set (aesenc, aesenclast), which computes substitution and mixing phases within dedicated execution units in silicon. The execution latency is invariant to the data values, and no intermediate states leak to cache lines.
Operating systems can mitigate cross-process line sharing by disabling Kernel Samepage Merging (KSM) and breaking shared file mappings across isolated execution boundaries. When containers or virtual machines require side-channel isolation, the host kernel can mark shared executable text with private page allocations using distinct physical memory frames:
# Disable Kernel Samepage Merging globally
echo 0 > /sys/kernel/mm/ksm/run
# Check unprivileged access to cycle counters
sysctl -w kernel.perf_event_paranoid=3
Hardware designers have deployed changes to cache eviction behavior. AMD processors built on the Zen architecture utilize largely non-inclusive L3 cache designs. In a non-inclusive hierarchy, a line present in L1 or L2 is not required to exist in L3. Invalidating a line in L3 using clflush does not automatically purge it from the private L1 or L2 caches of neighboring cores. This invalidates the core assumption of cross-core eviction tracking, requiring attackers to resort to more complex, error-prone eviction mechanisms.
Intel processors feature Cache Allocation Technology (CAT), part of Intel Resource Director Technology (RDT). CAT allows system administrators to partition L3 cache ways into isolated masks assigned to specific hardware class-of-service (CLOS) tags. System daemons can isolate cryptographic service processes into a dedicated cache way, preventing an attacker running on another core from flushing or observing lines assigned to that secure slice.
System monitoring daemons can detect ongoing Flush+Reload attacks by monitoring hardware performance counters via perf subsystems. An active attack executes millions of clflush instructions while generating abnormal L3 hit-to-miss ratios within tight loops. Attackers seeking to evade detection must throttle their sampling rates, which lowers data transmission bandwidth from several kilobytes per second down to single bits per second. Detecting anomalous behavior at the kernel level mirrors the techniques used when monitoring eBPF execution channels, as covered in How Linux eBPF Rootkits Evade Detection in Production.
Conclusion
Flush+Reload attacks turn standard memory management behaviors—shared libraries, page deduplication, and inclusive caches—into a high-resolution surveillance vector. Because the attack works on exact 64-byte lines, it isolates single functions and lookup table segments with sub-microsecond precision. Software-only mitigations that fail to address constant-time execution cannot guarantee protection against an attacker measuring shared physical cache frames. Defending against these side channels requires hardware-accelerated cryptographic primitives like AES-NI, the elimination of secret-dependent memory indexing, and the careful separation of physical page frames across security boundaries.
Measured on our own hardware: What 64 Bytes of Padding Are Worth
How much throughput does false sharing cost when several threads increment counters that share one cache line, compared with the same counters padded onto separate lines?
We ran it. The numbers below come from a program executed on the server hosting this site on 2026-09-27 — an AMD EPYC 9354P 32-Core Processor with 8 cores visible, 31.3 GB of memory, Linux 6.8.0-139-generic.
| Metric | Value |
|---|---|
| slowdown factor | 8.49 |
| iterations per thread | 10000000 |
| padded best seconds | 0.0405 |
| padded mean seconds | 0.0592 |
| padded million ops per sec | 1976.88 |
| repeats | 5 |
| shared line best seconds | 0.3435 |
| shared line mean seconds | 0.4076 |
| shared line million ops per sec | 232.9 |
| threads | 8 |
This is a shared virtual server, not an isolated test rig, so treat the absolute figures as indicative and the ratio between the two cases as the finding. The full method, the machine specification, and the complete source code are on the What 64 Bytes of Padding Are Worth benchmark page, so you can check the method or run it yourself.
Measured on our own hardware: Do Transparent Huge Pages Help on a VPS? Random Access Over 1 GiB
On this virtual server, does asking the kernel for 2 MiB transparent huge pages make a dependent random read over a 1 GiB buffer faster than the same buffer on 4 KiB pages, and how much of the buffer actually gets huge pages?
We ran it. The numbers below come from a program executed on the server hosting this site on 2026-09-27 — an AMD EPYC 9354P 32-Core Processor with 8 cores visible, 31.3 GB of memory, Linux 6.8.0-139-generic.
| Metric | Value |
|---|---|
| huge page speedup factor | 0.88 |
| accesses | 50000000 |
| buffer mib | 1024 |
| checksum | 61078 |
| huge case anon huge pages kb | 120832 |
| huge case fraction backed by 2m pages | 0.115 |
| huge pages ns per access best | 308.26 |
| huge pages ns per access mean | 313.3 |
| mode | measure |
| pages 2m in buffer | 512 |
| pages 4k in buffer | 262144 |
| repeats | 5 |
| small case anon huge pages kb | 0 |
| small pages ns per access best | 271.99 |
| small pages ns per access mean | 288.2 |
This is a shared virtual server, not an isolated test rig, so treat the absolute figures as indicative and the ratio between the two cases as the finding. The full method, the machine specification, and the complete source code are on the Do Transparent Huge Pages Help on a VPS? Random Access Over 1 GiB benchmark page, so you can check the method or run it yourself.
References
- Bernstein, D. J. "Cache-timing attacks on AES", 2005. https://cr.yp.to/antiforgery/cachetiming-20050414.pdf
- Yarom, Y., & Falkner, K. "FLUSH+RELOAD: A High Resolution, Low Noise, L3 Cache Side-Channel Attack", USENIX Security Symposium, 2014. https://www.usenix.org/conference/usenixsecurity14/technical-sessions/presentation/yarom
- Intel Corporation. "Intel 64 and IA-32 Architectures Software Developer Manuals", Volume 2A: Instruction Set Reference, 2024. https://www.intel.com/content/www/us/en/developer/articles/technical/intel-sdm.html