You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Intel Kaby Lake架构下紧邻指令的RIP相对加载CPU行为问询

RIP-Relative Addressing with +1 Offset on Intel Kaby Lake: Cache Behavior Breakdown

Let’s break down exactly what happens when you do a RIP-relative load with a +1 offset (pointing to data immediately adjacent to your current instruction) on a Kaby Lake CPU, especially when that data is in L1 Instruction Cache (L1I) but not L1 Data Cache (L1D).

First, Clarify the Address Calculation

When executing an instruction like mov rax, [rip+1], remember that RIP always points to the next instruction to execute, not the current one. So if your current instruction starts at address 0x1000 and is 5 bytes long, RIP during execution is 0x1005. Adding a +1 offset gives a target address of 0x1006—which is immediately after the current instruction, as you described.

Step-by-Step Cache Interaction on Kaby Lake

  1. Initial Cache State Check

    • The target data is in the same 64-byte cache line as your instruction (since it’s adjacent), so it was already pulled into L1I when the CPU fetched the instruction (L1I loads entire cache lines).
    • But Kaby Lake uses a Harvard architecture for L1: L1I and L1D are separate, physically distinct caches. The CPU cannot directly read data from L1I for a data load operation—data loads only query L1D, L2, L3, then main memory.
  2. L1D Miss Triggered

    • Since the data hasn’t been loaded into L1D yet, the load instruction will cause an L1D cache miss. The CPU sends a request for the target physical address to the next level of cache.
  3. L2 Cache Hit (Almost Guaranteed)

    • L2 cache on Kaby Lake is a unified cache (stores both instructions and data), and the cache line containing your instruction and adjacent data was already loaded into L2 to feed L1I. So the L2 cache will immediately recognize the address and respond with the cache line.
  4. L1D Cache Line Allocation

    • The CPU will allocate the cache line to L1D, marking it as valid for data accesses. The requested byte(s) are then read from the newly allocated L1D line and loaded into the destination register.

Does This Count as a Cache Hit?

It depends on which cache level you’re referring to:

  • L1D: No, the first load will be a miss (since the data wasn’t in L1D initially). Subsequent loads to the same address will hit L1D, though.
  • L2: Yes, this will be a hit because the cache line was already present from the instruction prefetch.
  • Overall cache hierarchy: You could argue it’s a "partial hit" since the data was already in the CPU’s cache subsystem (L1I/L2), just not in the dedicated data cache.

Additional Notes

  • Performance Impact: An L2 hit is much faster than a main memory access (around 4-5 cycles vs. ~100+ cycles), so even though it’s not an L1D hit, the penalty is minimal compared to a full cache miss.
  • Code/Data Permissions: As long as the memory region is marked readable (which code segments typically are, with RX permissions), this load operation will work without issues. NX (No-eXecute) bits only block execution, not reads.
  • Subsequent Accesses: Once the cache line is in L1D, any future RIP-relative loads to the same address will hit L1D directly, with the full low-latency benefit (~1 cycle).

内容的提问来源于stack exchange,提问作者bumpbump

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 10:42:29