You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Skylake SP干净缓存行回写原因咨询:L3非包容性的影响

Clean Cache Line Writeback on Skylake SP: Why Non-Inclusive L3 Matters

Great question—this is a nuanced behavior tied to how non-inclusive cache hierarchies interact with multi-socket coherence. Let's break this down clearly:

First: Inclusive vs. Non-Inclusive L3 Recap

To get to the root cause, let's contrast the two cache designs:

  • Inclusive L3 (older Intel server chips): Every line in L1/L2 has a copy in L3. When an L2 line gets evicted, the L3 still holds a valid copy. So clean lines (data matches DRAM) can just be discarded—no need to write back, since L3 has you covered.
  • Non-Inclusive L3 (Skylake SP+): L3 doesn't guarantee to hold copies of all L1/L2 lines. This lets Intel ship larger L2 caches (up to 1MB per core) without making L3 impractically big, but it changes eviction logic.

Why Clean L2 Lines Get Written Back Here

The confusion comes from thinking "clean lines match DRAM, so why write them?" But in dual-socket Skylake SP setups, two edge cases trigger this:

  1. Coherence State Management: If a clean L2 line is in the Shared MESI state (meaning the other socket might have a copy), evicting it requires ensuring the global coherence state stays consistent. On inclusive L3, the L3 would act as a "coherence anchor"—but without that, the hardware might write back the clean line to DRAM to signal that the line's valid copy now lives only in memory (instead of any cache). This is a niche case, not a common occurrence.
  2. Simplified Eviction Logic: Some cache controller designs simplify their code paths by treating clean and dirty lines the same in certain eviction scenarios, especially when the line isn't present in L3. Even though the data matches DRAM, the hardware might perform a writeback anyway to avoid adding extra conditional logic for clean lines.

Your Perf Data Makes Perfect Sense

Looking at your numbers:

  • The tiny number of clean writebacks (20 out of ~3.7B reads) confirms this is an edge case, not a regular part of L2 evictions.
  • The fact you only see this on dual-socket platforms ties directly to the coherence trigger—single-socket systems don't have to handle cross-socket state, so clean L2 lines can be safely discarded without writing back.

For clarity on your perf events:

  • uncore_imc_1/event=0x4,umask=0x3/: Tracks total DRAM read CAS operations, which lines up with your massive read workload.
  • uncore_imc_1/event=0xb8,umask=0x11/: Counts write CAS operations to a specific DRAM rank/bank group—your 20 counts here are exactly those rare clean writebacks.

Wrapping Up

Skylake SP's non-inclusive L3 is a tradeoff for better per-core performance (via larger L2), but it introduces rare edge cases where clean lines need to be written back—almost exclusively in multi-socket setups due to coherence rules. Your measurement method is solid, and the low count aligns with how infrequent this behavior is.


内容的提问来源于stack exchange,提问作者alexghiti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 19:27:48