Meltdown漏洞PoC代码技术问询:缓存计时与结果统计等疑问
Great question—this snippet is a core part of how side-channel attacks like Meltdown leverage cache behavior to leak secret data. Let’s break down every piece, including your specific questions.
What This Code Does Overall
This loop is a cache timing side-channel probe: it measures how long it takes to access each entry in array2 to figure out which ones are already in the CPU cache. Cache hits are much faster (tens of clock cycles) than misses (hundreds), so we can use timing to infer which memory addresses were secretly loaded into cache by the exploit.
Why Use Primes 167 and 13?
The line mix_i = ((i * 167) + 13) & 255 is designed to defeat CPU stride prediction:
- CPUs try to predict regular memory access patterns (like sequential loops) and prefetch data into cache before it’s needed. If we accessed
array2in order (0,1,2,...255), the CPU would prefetch most entries, making our timing results useless—every access would look like a cache hit. - Using primes in this linear congruential generator ensures the sequence of
mix_ivalues is as unordered as possible. 167 and 13 are small primes that, when combined with& 255(mod 256), create a full permutation: everyifrom 0-255 maps to a uniquemix_iwith no repeats or predictable steps. This guarantees we test every possible index without triggering prefetching.
Why Track Maximum and Second-Maximum Results?
In Meltdown/Spectre, we’re trying to leak a single byte of secret data (0-255), which maps directly to one of the array2 indices. Here’s why we need max/second-max:
- Ideally, the index corresponding to the secret value will have the highest
resultscount—because it was loaded into cache by the exploit, so every access to it is a hit. - But real-world systems have noise: random cache hits from other processes, tiny timing errors, or CPU interference.
- Comparing the top two values lets us validate the result’s reliability: if the max is significantly higher than the second-max (e.g., 2-3x more counts), we can be confident it’s the leaked secret. If they’re close, the probe was noisy and we need to re-run the loop.
Other Key Code Details
Let’s walk through the rest of the snippet:
addr = &array2[mix_i * 512]: Each index maps to a separate cache line. Cache lines are typically 64 bytes, so multiplying by 512 ensures eachmix_itargets a unique cache line (no overlap). This guarantees each access’s cache state is independent.__rdtscp(&junk): This x86 instruction reads the CPU’s Time Stamp Counter (TSC) and serializes execution—meaning it waits for all prior instructions to finish before taking the timestamp. Unlike plainrdtsc, this prevents out-of-order execution from skewing timing results.junk = *addr: Triggers the memory access we want to time. The value is stored injunkto stop the compiler from optimizing away the access entirely.time2 = __rdtscp(&junk) - time1: Calculates the total clock cycles taken for the memory access, giving us a precise measure of cache hit/miss status.if (time2 <= CACHE_HIT_THRESHOLD && mix_i != array1[tries % array1_size]):CACHE_HIT_THRESHOLDis a pre-defined value (usually ~100 cycles) that separates fast cache hits from slow misses.- The second condition filters out the index we intentionally loaded into cache during a prior setup step—this avoids counting our own "legitimate" cache hit and focuses on the secret one loaded by the exploit.
内容的提问来源于stack exchange,提问作者Antonin GAVREL

