C语言for循环数组赋值的异常性能问题排查
C语言并行代码性能异常原因分析
问题背景
原串行代码运行耗时约1.3秒,添加#pragma omp parallel for指令实现并行后,耗时降至约0.2秒。测试中发现以下异常现象:
- 移除循环最后一行赋值语句
theta_rec[i] = (1.0 / 4.0) * atan2(sum_i, sum_r) - M_PI / 4;时,代码耗时变为0秒; - 仅计算表达式并赋值给临时变量
double tmp = (1.0 / 4.0) * atan2(sum_i, sum_r) - M_PI / 4;,耗时仍为0秒; - 将临时变量
tmp赋值给theta_rec[i]时,耗时回到0.2秒; - 直接给
theta_rec[i]赋值常量1.2,耗时同样为0秒。
测试代码
核心计算函数
void test(double* sigs_in_r, double* sigs_in_i, int Nsig, int nCPE, double* theta_rec) { #pragma omp parallel for for (int i = 0; i < Nsig; i++) { int start = i - nCPE > 0 ? i - nCPE : 0; int stop = i + nCPE < Nsig - 2 ? i + nCPE : Nsig - 2; double sum_r = 0.0; double sum_i = 0.0; for (int j = start; j <= stop; j++) { double sigs_in_r_j = sigs_in_r[j]; double sigs_in_i_j = sigs_in_i[j]; double abs_val = sqrt(sigs_in_r_j * sigs_in_r_j + sigs_in_i_j * sigs_in_i_j); double arg_val = atan2(sigs_in_i_j, sigs_in_r_j); double abs_val_pow4 = abs_val * abs_val * abs_val * abs_val; sum_r += abs_val_pow4 * cos(4 * arg_val); sum_i += abs_val_pow4 * sin(4 * arg_val); } theta_rec[i] = (1.0 / 4.0) * atan2(sum_i, sum_r) - M_PI / 4; //double tmp = (1.0 / 4.0) * atan2(sum_i, sum_r) - M_PI / 4; //theta_rec[i] = tmp; //theta_rec[i] = 1.2; } }
主测试函数
#define SIZE 100000 int main() { clock_t start, end; double cpu_time_used; double *arr1 = calloc(SIZE, sizeof(double)); double *arr2 = calloc(SIZE, sizeof(double)); double *theta = calloc(SIZE, sizeof(double)); for (int i = 0; i < SIZE; i++) { arr1[i] = 1.0; arr2[i] = 2.0; } start = clock(); // 启动计时器 test(arr1, arr2, SIZE, 100, theta); end = clock(); // 停止计时器 cpu_time_used = ((double) (end - start)) / CLOCKS_PER_SEC; // 计算耗时(秒) printf("Computation time: %f seconds\n", cpu_time_used); return 0; }
头文件依赖
#include <stdlib.h> #include <stdio.h> #include <omp.h> #define _USE_MATH_DEFINES #include <math.h> #include <time.h>
编译命令
gcc -Wall -Wextra -O3 -Ofast -fopenmp -std=c99 test.c -o test
原因分析
这是编译器高优化级别下的死代码消除导致的现象,具体逻辑如下:
- 无输出时的全量消除:当代码没有将计算结果写入
theta_rec数组、仅赋值给未使用的临时变量时,编译器发现所有计算逻辑的结果没有被后续代码读取(主函数未使用theta数组内容),在-O3/-Ofast优化级别下会直接删除所有无关计算,自然耗时为0秒。 - 赋值常量时的逻辑简化:直接给
theta_rec[i]赋值常量时,编译器能识别出内层循环的计算与最终赋值无关,会删除所有内层循环的计算逻辑,仅保留批量内存赋值操作——该操作耗时极短,因此显示为0秒。 - 赋值表达式结果时的计算保留:当把表达式结果写入
theta_rec[i]时,编译器必须保留内层循环的计算逻辑(需要sum_r和sum_i来推导最终值),同时执行内存写入操作,所有计算逻辑被完整保留,因此耗时回到0.2秒。
验证方式:在主函数计时结束后,添加一行读取theta数组元素的代码(如printf("%lf\n", theta[0]);),此时即使是仅计算临时变量的情况,耗时也会变回0.2秒——因为编译器无法再消除计算逻辑,结果需要被使用。
内容的提问来源于stack exchange,提问作者phw
相关产品推荐
相关产品推荐

