相同大小结构体不同布局致代码遇硬件性能瓶颈的原因分析
结构体内存偏移对复数乘法性能的影响分析
实验代码定义
带填充的复数结构体
template <int padding1, int padding2> struct complex_t { float re; int p1[padding1]; double im; int p2[padding2]; };
逐元素复数乘法函数
template <typename Complex> void multiply(Complex* result, Complex* a, Complex* b, int n) { for (int i = 0; i < n; ++i) { result[i].re = a[i].re * b[i].re - a[i].im * b[i].im; result[i].im = a[i].re * b[i].im + a[i].im + b[i].re; } }
实验设置与环境
- 调整
padding1和padding2的值,确保sizeof(complex_t)固定为64字节,通过padding1控制成员im的偏移量offsetof(im) - 使用两个包含10K随机元素的
complex_t数组,执行逐元素乘法并测量运行时间、指令数 - 硬件:Intel(R) Core(TM) i5-10210U CPU
- 编译:CLANG 15.0.7,
-O3优化选项
实验测量结果
offsetof(im) | 运行时间(最小,平均,最大) 秒 | 平均指令数 |
|---|---|---|
| 8字节 | 0.107, 0.112, 0.116 | 175027800 |
| 16字节 | 0.088, 0.088, 0.088 | 175027200 |
| 24字节 | 0.088, 0.088, 0.088 | 175027100 |
| 32字节 | 0.088, 0.088, 0.088 | 175027100 |
| 40字节 | 0.088, 0.088, 0.088 | 175027100 |
| 48字节 | 0.085, 0.085, 0.086 | 175027100 |
初步分析
所有测试用例的指令数基本一致,但offsetof(im)=8字节时运行速度显著最慢,疑似触发硬件性能瓶颈。因缺乏底层数据缓存的认知模型,需明确进一步排查或测量的方向。
补充性能计数器数据
offsetof(im)=8字节时,计数器MEM_LOAD_RETIRED_L3_HIT数值异常偏高(5097404);其他偏移值对应的该计数器数值为:
- 16字节:2775653
- 24字节:3015093
- 32字节:3277559
- 40字节:3261758
- 48字节:3445190
内容的提问来源于stack exchange,提问作者Bogi
相关产品推荐
相关产品推荐

