Python调用含complex.h的C DLL运行速度极慢的优化问询
问题:C标准库complex.h函数在DLL调用时性能暴跌的优化方案
我正在把Python函数转成C语言通过.dll供Python调用,实验阶段碰到核心问题:用complex.h里的函数(比如carg、cpow、cabs)时,代码运行速度慢得离谱。比如去掉theta_rec[i] = carg(sumv1);这句,代码执行只需要0.3秒,保留的话耗时直接超10秒。单独跑C代码没这问题,手动实现这些complex.h函数后性能有提升,但还想找更好的优化办法。
更新:发现complex.h所有相关函数都会导致速度骤降,手动实现后性能改善,但仍需更优方案。
Python端代码
import ctypes import numpy as np import os import _ctypes testnew = ctypes.CDLL("test.dll", winmode=0x8) # Define the function argument types testnew.testfunc.argtypes = [ np.ctypeslib.ndpointer(dtype=np.complex128, ndim=2, flags="C"), # sigs_in ctypes.c_int, # Nsymb ctypes.c_int, # Nsig ctypes.c_int, # nCPE np.ctypeslib.ndpointer(dtype=np.complex128, ndim=2, flags="C"), # sigs_out np.ctypeslib.ndpointer(dtype=np.double, ndim=2, flags="C") # theta_rec ] # Define the function return type testnew.testfunc.restype = None Nsig, Nsymb = data.shape nCPE = 100 # Create input and output arrays sigs_in = data.values sigs_in = np.ascontiguousarray(sigs_in, dtype=np.complex128) sigs_out = np.zeros_like(sigs_in) theta_rec = np.zeros_like( sigs_in, dtype=float ) sigs_out = np.ascontiguousarray(sigs_out, dtype=np.complex128) theta_rec = np.ascontiguousarray(theta_rec, dtype=np.double) sigs_in = np.asarray(sigs_in, order='C') sigs_out = np.asarray(sigs_out, order='C') theta_rec = np.asarray(theta_rec, order='C') # Call the function testnew.testfunc(sigs_in, Nsymb, Nsig, nCPE, sigs_out, theta_rec) # Print the results print("sigs_out:") print(sigs_out) print("theta_rec:") print(theta_rec) _ctypes.FreeLibrary(testnew._handle)
C端代码
#include <stdlib.h> #define _USE_MATH_DEFINES #include <math.h> #include <stdio.h> #include <complex.h> #if defined(_WIN32) # define DLL00_EXPORT_API __declspec(dllexport) #else # define DLL00_EXPORT_API #endif DLL00_EXPORT_API void testfunc(complex double *sigs_in, int Nsymb, int Nsig, int nCPE, complex double *sigs_out, double *theta_rec) { for (int i = 0; i < Nsig*Nsymb; i += Nsymb) { int start = (i - Nsymb*nCPE) > 0 ? (i - Nsymb*nCPE) : 0; int stop = (i + Nsymb*nCPE) < (Nsymb*Nsig - 2) ? (i + Nsymb*nCPE) : (Nsymb*Nsig - 2); double complex sumv1 = 0; double complex sumv2 = 0; for (int j = start; j < stop; j++) { sumv1 += cpow((sigs_in[j + 0] / cabs(sigs_in[j + 0] )), 4); sumv2 += cpow((sigs_in[j + 1] / cabs(sigs_in[j + 1] )), 4); } theta_rec[i] = carg(sumv1); // the function carg() slows down my code a lot } }
优化方案
- 拉满编译器优化
编译DLL时开启最高级别优化:MSVC用/O2 /Oi /Ot,GCC用-O3 -march=native。标准库complex函数在优化全开时会被内联或替换为硬件指令,消除函数调用开销。 - 替换complex.h函数为轻量实现
直接用底层数学函数替代封装层:cabs(z)→hypot(creal(z), cimag(z))或sqrt(creal(z)*creal(z) + cimag(z)*cimag(z))carg(z)→atan2(cimag(z), creal(z))- 针对代码中单位复数的4次方场景,完全避开
cpow:单位复数z = x+yi的4次方等价于角度乘4后的复数,直接用三角函数计算:
连除法和模运算都省了,性能提升显著。double x = creal(sigs_in[j]); double y = cimag(sigs_in[j]); double angle4 = atan2(y, x) * 4; sumv1 += cos(angle4) + sin(angle4)*I;
- 缓存重复计算结果
避免循环中重复计算cabs(sigs_in[j]),缓存后复用:double abs_val = cabs(sigs_in[j]); double complex unit = sigs_in[j] / abs_val; sumv1 += cpow(unit,4); - 优化内存访问模式
当前外层循环i += Nsymb会导致内存跳跃访问,尽量调整循环顺序,让内层循环连续访问内存,减少缓存失效。比如先按Nsymb维度循环,再处理Nsig维度。 - SIMD指令加速
若编译器自动优化不足,可手动使用SSE/AVX等SIMD指令批量处理复数运算,比如用_mm_add_pd、_mm_mul_pd等指令并行计算多个复数的累加和幂运算。
内容的提问来源于stack exchange,提问作者phw
相关产品推荐
相关产品推荐

