You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

跨cache line对循环解码的影响及相关性能变化原因技术问询

摘要

观测发现大量小型循环在跨cache line时会出现严重性能下降,该现象仅与跨cache line相关,与循环是否跨越其他取指块边界无关。
该问题看起来与解码过程存在关联。

核心问题

本次问询的核心问题如下:

  1. 跨cache line会如何影响循环的解码过程?
  2. 该影响会如何作用于循环的运行性能?

已完成的测试研究

采用如下测试程序开展验证,配套Makefile与运行脚本已公开:

#define LOOP_CNT (1000 * 1000 * 1000)

    /* Just #define for NOP{N}. */
#include "nops.h"
#define XOR 0
#define MOV 1
#define DSB0    2
#define DSB1    3
#define DSB2    4
#define DSB3    5

#ifdef RUNALL
# include "padding.h"
#else
# define PAYLOAD    DSB0

    /* Just to see if certain cacheline offsets within a page may
       change things. So far nothing.  */
# define CACHE_PADDING  NOP0

    /* Offset within cache line. Using NOP12, NOP28, NOP44, and
       NOP60.  */
# define LOOP_PADDING   NOP60
#endif
    .global _start
    .text
_start:
    movl    $LOOP_CNT, %ecx
    leaq    (2048 + buf_start)(%rip), %rsp
    movq    %rsp, %rsi
    leaq    64(%rsi), %rdi

    /* Page align.  */
    .p2align 12

    /* Padding config.  */
    LOOP_PADDING

    /* Loop is 8 bytes.  */
loop:
#if PAYLOAD == XOR
    /* Payload where loop is really just the decl/jnz.  */
    xorl    %eax, %eax
    xorl    %edx, %edx
#elif PAYLOAD == MOV
    /* Payload testing if it has to do with memory bandwidth.  */
    movl    (%rsi), %eax
    movl    (%rdi), %edx
#elif PAYLOAD == DSB0
    /* Payload testing how being able to run out of the LSD vs DSB
       changes things. LSD disabling behavior in first fetch block.  */
    popq    %rax
    movq    %rsi, %rsp
#elif PAYLOAD == DSB1
    /* Payload testing how being able to run out of the LSD vs DSB
       changes things. LSD disabling behavior in second fetch block.
     */
    movq    %rsi, %rsp
    popq    %rax
#elif PAYLOAD == DSB2
    /* Payload testing how being able to run out of the LSD vs DSB
       changes things. LSD disabling behavior in first fetch block and
       non-eliminatable mov cache line.  */
    popq    %rax
    leaq    (%rsi), %rsp
#elif PAYLOAD == DSB3
    /* Payload testing how being able to run out of the LSD vs DSB
       changes things. LSD disabling behavior in second fetch block and
       non-eliminatable first cache line.  */
    leaq    (%rsi), %rsp
    popq    %rax
#else
# error NO PAYLOAD
#endif
    decl    %ecx
    jnz loop

    movl    $60, %eax
    xorl    %edi, %edi
    syscall

    .section .data
    .balign 4096
buf_start:  .space 4096
buf_end:

选取NOP12、NOP28、NOP44、NOP60四个测试组,所有组的循环均会跨越取指块,仅NOP60组的循环会跨cache line。
测试观测到跨cache line的NOP60组性能出现明显下降,其余组性能均稳定维持在“优秀”水平(单迭代1周期)。
PAYLOAD设置为XOR时,cycles、lsd_uops、dsb_uops、mite_uops的测试结果如下。
注:所有数据为5次运行的中位数,测试平台为Tigerlake:

PAYLOADLOOP_PADDINGCYCLESLSD_UOPSDSB_UOPSMITE_UOPS
XOR121.000e+093.000e+094.629e+034.468e+03
XOR281.001e+093.000e+091.651e+048.511e+03
XOR441.000e+093.000e+094.869e+034.697e+03
XOR601.102e+092.693e+094.328e+033.069e+08

跨cache line场景下性能下降约10%,原因看起来是循环的部分解码工作由MITE完成,而非全部由LSD解码。
通过不匹配push/pop操作强制循环不使用LSD时,跨cache line的测试结果恶化更为明显。

PAYLOADLOOP_PADDINGCYCLESLSD_UOPSDSB_UOPSMITE_UOPS
DSB0121.000e+090.000e+003.000e+091.050e+04
DSB0281.001e+090.000e+003.000e+092.267e+04
DSB0441.003e+090.000e+003.000e+098.710e+04
DSB0602.001e+091.822e+052.522e+094.773e+08

该场景下跨cache line组的运行耗时是其他跨取指块组的2倍。
同时可观测到跨cache line的循环仍然从LSD获取了部分uops,这暗示循环可能沿cache line边界被“拆分”处理。
当前DSB0版本的循环拆分情况如下:

000000000000003c <loop>:
    003c:   58                      pop    %rax
    003d:   48 89 f4                mov    %rsi,%rsp
    0040:   ff c9                   dec    %ecx
    0042:   75 f8                   jne    103c <loop>

第二个cache line不存在LSD禁用行为,看起来至少部分场景下dec; jne指令由LSD解码。
如果调换pop与mov的顺序(DSB1版本),并将LOOP_PADDING加1,使用NOP61时循环结构如下:

000000000000003d <loop>:
    003d:   48 89 f4                mov    %rsi,%rsp
    0040:   58                      pop    %rax
    0041:   ff c9                   dec    %ecx
    0043:   75 f8                   jne    103c <loop>

得到如下测试结果:

PAYLOADLOOP_PADDINGCYCLESLSD_UOPSDSB_UOPSMITE_UOPS
DSB1131.008e+090.000e+003.000e+091.837e+05
DSB1291.008e+090.000e+003.000e+091.269e+05
DSB1451.005e+090.000e+003.000e+091.904e+05
DSB1612.007e+090.000e+002.699e+093.013e+08

该场景下循环完全不再从LSD取指,但仍存在2倍性能下降,看起来是由于循环部分从MITE而非更快的解码器(该场景下为DSB)取指。这也说明2倍性能下降是DSB与MITE切换带来的,而LSD与MITE切换仅带来约10%的性能损失。
在无LSD的Skylake平台上测试PAYLOAD == XOR的场景,结果如下:

PAYLOADLOOP_PADDINGCYCLESLSD_UOPSDSB_UOPSMITE_UOPS
XOR121.021e+090.000e+003.002e+091.190e+06
XOR281.019e+090.000e+003.003e+091.469e+06
XOR441.020e+090.000e+003.003e+091.526e+06
XOR602.044e+090.000e+002.755e+092.490e+08

该结果支持2倍性能下降是DSB/MITE切换导致的结论。
基于测试得到如下观测结论:

  1. 跨cache line会中断两类针对循环优化的解码器工作,该特性为跨cache line独有,与任意取指块边界交叉无关。
  2. 解码器会将跨cache line的循环识别为多个独立实体,该现象较为特殊。
  3. 循环从DSB取指时跨cache line的性能损失比从LSD取指时更严重。

目前暂无法解释上述观测现象的底层原因,请求专业人士解答背后的运行机制。


内容的提问来源于stack exchange,提问作者Noah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 21:18:02