如何准确测量Cython运行时间并与Python代码对比?
问题描述
在Cython的.pyx文件中使用time.time()测量代码运行时间时,结果存在明显误差:相同代码每次运行耗时波动大,执行相同任务的循环耗时差一倍(如0.010和0.020),有时某段代码耗时显示为0,其他同类代码段却有0.010或0.020的耗时。查阅Cython文档未找到合适的解决方法,希望掌握正确的性能测量方式。
问题代码片段
t4 = time.time() # print('T3 =', t4 - t3) for j in range(np.shape(im1)[1]): # sum_c1[j] = np.shape(im1)[0] - (np.sum(im1[:, j])) sum_c1[j] = np.shape(im1)[0] - (np.count_nonzero(im1[:, j])) tt3 = time.time() print('TT3 =', tt3 - t4) cdef int amc1 = np.argmax(sum_c1) # argmax sum_c1 tt4 = time.time() # print('TT4 =', tt4 - tt3) for j in range(np.shape(im2)[1]): # sum_c2[j] = np.shape(im2)[0] - (np.sum(im2[:, j])) sum_c2[j] = np.shape(im2)[0] - (np.count_nonzero(im2[:, j])) t5 = time.time() print('TT5 =', t5 - tt4) # print('T4 =', t5 - t4) ## find of max zeros in row for j in range(np.shape(im1)[0]): # sum_r1[j] = np.shape(im1)[1] - (np.sum(im1[j, :])) sum_r1[j] = np.shape(im1)[1] - (np.count_nonzero(im1[j, :])) tt1 = time.time() print('TT1 =', tt1 - t5) cdef int amr1 = np.argmax(sum_r1) # argmax sum_r1 tt2 = time.time() # print('TT2 =', tt2 - tt1) for j in range(np.shape(im2)[0]): # sum_r2[j] = np.shape(im2)[1] - (np.sum(im2[j, :])) sum_r2[j] = np.shape(im2)[1] - (np.count_nonzero(im2[j, :])) t6 = time.time() print('T5 =', t6 - t5)
两次运行的输出结果
('TT3 =', 0.020589590072631836) ('TT5 =', 0.011527061462402344) ('TT1 =', 0.0) ('T5 =', 0.009999990463256836) ----------- ('TT3 =', 0.0100250244140625) ('TT5 =', 0.00996851921081543) ('TT1 =', 0.01003265380859375) ('T5 =', 0.020001888275146484)
正确的性能测量方法
1. 替换为高精度计时函数
time.time()的精度依赖操作系统(Windows上仅约15ms精度),改用time.perf_counter()——它是Python专门为性能测量设计的最高精度计时器,能捕获更细微的时间差,避免出现耗时为0的情况。
2. 多次运行取平均值
单次运行的耗时受系统负载(后台进程、CPU调度)影响极大,必须多次重复执行目标代码块,取平均耗时才能反映真实性能。建议循环执行100~1000次,总耗时除以运行次数得到平均值。
3. 移除测量中的干扰操作
- 将打印语句移出计时块:IO操作会严重拖慢代码执行,干扰计时结果,应先收集所有耗时数据,最后统一打印。
- 提前初始化变量:确保numpy数组等对象在测量前已完成内存分配,避免首次运行的初始化开销影响结果。
4. 优化Cython代码减少耗时波动
你的代码在循环中反复调用np.shape()、np.count_nonzero()等Python层面的函数,会增加开销和波动。可以做以下优化:
- 提前将数组shape存入Cython原生int变量,避免循环中重复调用:
cdef int im1_rows = im1.shape[0] cdef int im1_cols = im1.shape[1] - 为numpy数组声明Cython类型,加速元素访问:
cdef np.ndarray[np.uint8_t, ndim=2] im1, im2 cdef np.ndarray[np.int32_t, ndim=1] sum_c1, sum_c2, sum_r1, sum_r2
5. 修改后的示例代码
import time import numpy as np def measure_optimized(): # 假设im1、im2、sum_c1等已提前初始化并声明Cython类型 cdef int runs = 500 cdef double total_tt3 = 0.0 cdef double total_tt5 = 0.0 cdef double total_tt1 = 0.0 cdef double total_t5 = 0.0 cdef int im1_rows = im1.shape[0] cdef int im1_cols = im1.shape[1] cdef int im2_rows = im2.shape[0] cdef int im2_cols = im2.shape[1] for _ in range(runs): t4 = time.perf_counter() for j in range(im1_cols): sum_c1[j] = im1_rows - np.count_nonzero(im1[:, j]) tt3 = time.perf_counter() total_tt3 += tt3 - t4 amc1 = np.argmax(sum_c1) tt4 = time.perf_counter() for j in range(im2_cols): sum_c2[j] = im2_rows - np.count_nonzero(im2[:, j]) t5 = time.perf_counter() total_tt5 += t5 - tt4 for j in range(im1_rows): sum_r1[j] = im1_cols - np.count_nonzero(im1[j, :]) tt1 = time.perf_counter() total_tt1 += tt1 - t5 amr1 = np.argmax(sum_r1) tt2 = time.perf_counter() for j in range(im2_rows): sum_r2[j] = im2_cols - np.count_nonzero(im2[j, :]) t6 = time.perf_counter() total_t5 += t6 - t5 # 打印平均耗时 print(f'TT3 平均耗时: {total_tt3 / runs:.6f}') print(f'TT5 平均耗时: {total_tt5 / runs:.6f}') print(f'TT1 平均耗时: {total_tt1 / runs:.6f}') print(f'T5 平均耗时: {total_t5 / runs:.6f}')
内容的提问来源于stack exchange,提问作者sinamcr7
相关产品推荐
相关产品推荐

