跨cache line对循环解码的影响及相关性能变化原因技术问询
摘要
观测发现大量小型循环在跨cache line时会出现严重性能下降,该现象仅与跨cache line相关,与循环是否跨越其他取指块边界无关。
该问题看起来与解码过程存在关联。
核心问题
本次问询的核心问题如下:
- 跨cache line会如何影响循环的解码过程?
- 该影响会如何作用于循环的运行性能?
已完成的测试研究
采用如下测试程序开展验证,配套Makefile与运行脚本已公开:
#define LOOP_CNT (1000 * 1000 * 1000) /* Just #define for NOP{N}. */ #include "nops.h" #define XOR 0 #define MOV 1 #define DSB0 2 #define DSB1 3 #define DSB2 4 #define DSB3 5 #ifdef RUNALL # include "padding.h" #else # define PAYLOAD DSB0 /* Just to see if certain cacheline offsets within a page may change things. So far nothing. */ # define CACHE_PADDING NOP0 /* Offset within cache line. Using NOP12, NOP28, NOP44, and NOP60. */ # define LOOP_PADDING NOP60 #endif .global _start .text _start: movl $LOOP_CNT, %ecx leaq (2048 + buf_start)(%rip), %rsp movq %rsp, %rsi leaq 64(%rsi), %rdi /* Page align. */ .p2align 12 /* Padding config. */ LOOP_PADDING /* Loop is 8 bytes. */ loop: #if PAYLOAD == XOR /* Payload where loop is really just the decl/jnz. */ xorl %eax, %eax xorl %edx, %edx #elif PAYLOAD == MOV /* Payload testing if it has to do with memory bandwidth. */ movl (%rsi), %eax movl (%rdi), %edx #elif PAYLOAD == DSB0 /* Payload testing how being able to run out of the LSD vs DSB changes things. LSD disabling behavior in first fetch block. */ popq %rax movq %rsi, %rsp #elif PAYLOAD == DSB1 /* Payload testing how being able to run out of the LSD vs DSB changes things. LSD disabling behavior in second fetch block. */ movq %rsi, %rsp popq %rax #elif PAYLOAD == DSB2 /* Payload testing how being able to run out of the LSD vs DSB changes things. LSD disabling behavior in first fetch block and non-eliminatable mov cache line. */ popq %rax leaq (%rsi), %rsp #elif PAYLOAD == DSB3 /* Payload testing how being able to run out of the LSD vs DSB changes things. LSD disabling behavior in second fetch block and non-eliminatable first cache line. */ leaq (%rsi), %rsp popq %rax #else # error NO PAYLOAD #endif decl %ecx jnz loop movl $60, %eax xorl %edi, %edi syscall .section .data .balign 4096 buf_start: .space 4096 buf_end:
选取NOP12、NOP28、NOP44、NOP60四个测试组,所有组的循环均会跨越取指块,仅NOP60组的循环会跨cache line。
测试观测到跨cache line的NOP60组性能出现明显下降,其余组性能均稳定维持在“优秀”水平(单迭代1周期)。PAYLOAD设置为XOR时,cycles、lsd_uops、dsb_uops、mite_uops的测试结果如下。
注:所有数据为5次运行的中位数,测试平台为Tigerlake:
| PAYLOAD | LOOP_PADDING | CYCLES | LSD_UOPS | DSB_UOPS | MITE_UOPS |
|---|---|---|---|---|---|
| XOR | 12 | 1.000e+09 | 3.000e+09 | 4.629e+03 | 4.468e+03 |
| XOR | 28 | 1.001e+09 | 3.000e+09 | 1.651e+04 | 8.511e+03 |
| XOR | 44 | 1.000e+09 | 3.000e+09 | 4.869e+03 | 4.697e+03 |
| XOR | 60 | 1.102e+09 | 2.693e+09 | 4.328e+03 | 3.069e+08 |
跨cache line场景下性能下降约10%,原因看起来是循环的部分解码工作由MITE完成,而非全部由LSD解码。
通过不匹配push/pop操作强制循环不使用LSD时,跨cache line的测试结果恶化更为明显。
| PAYLOAD | LOOP_PADDING | CYCLES | LSD_UOPS | DSB_UOPS | MITE_UOPS |
|---|---|---|---|---|---|
| DSB0 | 12 | 1.000e+09 | 0.000e+00 | 3.000e+09 | 1.050e+04 |
| DSB0 | 28 | 1.001e+09 | 0.000e+00 | 3.000e+09 | 2.267e+04 |
| DSB0 | 44 | 1.003e+09 | 0.000e+00 | 3.000e+09 | 8.710e+04 |
| DSB0 | 60 | 2.001e+09 | 1.822e+05 | 2.522e+09 | 4.773e+08 |
该场景下跨cache line组的运行耗时是其他跨取指块组的2倍。
同时可观测到跨cache line的循环仍然从LSD获取了部分uops,这暗示循环可能沿cache line边界被“拆分”处理。
当前DSB0版本的循环拆分情况如下:
000000000000003c <loop>: 003c: 58 pop %rax 003d: 48 89 f4 mov %rsi,%rsp 0040: ff c9 dec %ecx 0042: 75 f8 jne 103c <loop>
第二个cache line不存在LSD禁用行为,看起来至少部分场景下dec; jne指令由LSD解码。
如果调换pop与mov的顺序(DSB1版本),并将LOOP_PADDING加1,使用NOP61时循环结构如下:
000000000000003d <loop>: 003d: 48 89 f4 mov %rsi,%rsp 0040: 58 pop %rax 0041: ff c9 dec %ecx 0043: 75 f8 jne 103c <loop>
得到如下测试结果:
| PAYLOAD | LOOP_PADDING | CYCLES | LSD_UOPS | DSB_UOPS | MITE_UOPS |
|---|---|---|---|---|---|
| DSB1 | 13 | 1.008e+09 | 0.000e+00 | 3.000e+09 | 1.837e+05 |
| DSB1 | 29 | 1.008e+09 | 0.000e+00 | 3.000e+09 | 1.269e+05 |
| DSB1 | 45 | 1.005e+09 | 0.000e+00 | 3.000e+09 | 1.904e+05 |
| DSB1 | 61 | 2.007e+09 | 0.000e+00 | 2.699e+09 | 3.013e+08 |
该场景下循环完全不再从LSD取指,但仍存在2倍性能下降,看起来是由于循环部分从MITE而非更快的解码器(该场景下为DSB)取指。这也说明2倍性能下降是DSB与MITE切换带来的,而LSD与MITE切换仅带来约10%的性能损失。
在无LSD的Skylake平台上测试PAYLOAD == XOR的场景,结果如下:
| PAYLOAD | LOOP_PADDING | CYCLES | LSD_UOPS | DSB_UOPS | MITE_UOPS |
|---|---|---|---|---|---|
| XOR | 12 | 1.021e+09 | 0.000e+00 | 3.002e+09 | 1.190e+06 |
| XOR | 28 | 1.019e+09 | 0.000e+00 | 3.003e+09 | 1.469e+06 |
| XOR | 44 | 1.020e+09 | 0.000e+00 | 3.003e+09 | 1.526e+06 |
| XOR | 60 | 2.044e+09 | 0.000e+00 | 2.755e+09 | 2.490e+08 |
该结果支持2倍性能下降是DSB/MITE切换导致的结论。
基于测试得到如下观测结论:
- 跨cache line会中断两类针对循环优化的解码器工作,该特性为跨cache line独有,与任意取指块边界交叉无关。
- 解码器会将跨cache line的循环识别为多个独立实体,该现象较为特殊。
- 循环从
DSB取指时跨cache line的性能损失比从LSD取指时更严重。
目前暂无法解释上述观测现象的底层原因,请求专业人士解答背后的运行机制。
内容的提问来源于stack exchange,提问作者Noah

