You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何准确测量Cython运行时间并与Python代码对比?

问题描述

在Cython的.pyx文件中使用time.time()测量代码运行时间时,结果存在明显误差:相同代码每次运行耗时波动大,执行相同任务的循环耗时差一倍(如0.010和0.020),有时某段代码耗时显示为0,其他同类代码段却有0.010或0.020的耗时。查阅Cython文档未找到合适的解决方法,希望掌握正确的性能测量方式。

问题代码片段

t4 = time.time()
# print('T3 =', t4 - t3)
for j in range(np.shape(im1)[1]):
    # sum_c1[j] = np.shape(im1)[0] - (np.sum(im1[:, j]))
    sum_c1[j] = np.shape(im1)[0] - (np.count_nonzero(im1[:, j]))
tt3 = time.time()
print('TT3 =', tt3 - t4)
cdef int amc1 = np.argmax(sum_c1)  # argmax sum_c1
tt4 = time.time()
# print('TT4 =', tt4 - tt3)
for j in range(np.shape(im2)[1]):
    # sum_c2[j] = np.shape(im2)[0] - (np.sum(im2[:, j]))
    sum_c2[j] = np.shape(im2)[0] - (np.count_nonzero(im2[:, j]))
t5 = time.time()
print('TT5 =', t5 - tt4)
# print('T4 =', t5 - t4)
## find of max zeros in row
for j in range(np.shape(im1)[0]):
    # sum_r1[j] = np.shape(im1)[1] - (np.sum(im1[j, :]))
    sum_r1[j] = np.shape(im1)[1] - (np.count_nonzero(im1[j, :]))
tt1 = time.time()
print('TT1 =', tt1 - t5)
cdef int amr1 = np.argmax(sum_r1)  # argmax sum_r1
tt2 = time.time()
# print('TT2 =', tt2 - tt1)
for j in range(np.shape(im2)[0]):
    # sum_r2[j] = np.shape(im2)[1] - (np.sum(im2[j, :]))
    sum_r2[j] = np.shape(im2)[1] - (np.count_nonzero(im2[j, :]))
t6 = time.time()
print('T5 =', t6 - t5)

两次运行的输出结果

('TT3 =', 0.020589590072631836)
('TT5 =', 0.011527061462402344)
('TT1 =', 0.0)
('T5 =', 0.009999990463256836)

-----------
('TT3 =', 0.0100250244140625)
('TT5 =', 0.00996851921081543)
('TT1 =', 0.01003265380859375)
('T5 =', 0.020001888275146484)
正确的性能测量方法

1. 替换为高精度计时函数

time.time()的精度依赖操作系统(Windows上仅约15ms精度),改用time.perf_counter()——它是Python专门为性能测量设计的最高精度计时器,能捕获更细微的时间差,避免出现耗时为0的情况。

2. 多次运行取平均值

单次运行的耗时受系统负载(后台进程、CPU调度)影响极大,必须多次重复执行目标代码块,取平均耗时才能反映真实性能。建议循环执行100~1000次,总耗时除以运行次数得到平均值。

3. 移除测量中的干扰操作

  • 将打印语句移出计时块:IO操作会严重拖慢代码执行,干扰计时结果,应先收集所有耗时数据,最后统一打印。
  • 提前初始化变量:确保numpy数组等对象在测量前已完成内存分配,避免首次运行的初始化开销影响结果。

4. 优化Cython代码减少耗时波动

你的代码在循环中反复调用np.shape()、np.count_nonzero()等Python层面的函数,会增加开销和波动。可以做以下优化:

  • 提前将数组shape存入Cython原生int变量,避免循环中重复调用:
    cdef int im1_rows = im1.shape[0]
    cdef int im1_cols = im1.shape[1]
    
  • 为numpy数组声明Cython类型,加速元素访问:
    cdef np.ndarray[np.uint8_t, ndim=2] im1, im2
    cdef np.ndarray[np.int32_t, ndim=1] sum_c1, sum_c2, sum_r1, sum_r2
    

5. 修改后的示例代码

import time
import numpy as np

def measure_optimized():
    # 假设im1、im2、sum_c1等已提前初始化并声明Cython类型
    cdef int runs = 500
    cdef double total_tt3 = 0.0
    cdef double total_tt5 = 0.0
    cdef double total_tt1 = 0.0
    cdef double total_t5 = 0.0
    cdef int im1_rows = im1.shape[0]
    cdef int im1_cols = im1.shape[1]
    cdef int im2_rows = im2.shape[0]
    cdef int im2_cols = im2.shape[1]

    for _ in range(runs):
        t4 = time.perf_counter()
        for j in range(im1_cols):
            sum_c1[j] = im1_rows - np.count_nonzero(im1[:, j])
        tt3 = time.perf_counter()
        total_tt3 += tt3 - t4

        amc1 = np.argmax(sum_c1)
        tt4 = time.perf_counter()

        for j in range(im2_cols):
            sum_c2[j] = im2_rows - np.count_nonzero(im2[:, j])
        t5 = time.perf_counter()
        total_tt5 += t5 - tt4

        for j in range(im1_rows):
            sum_r1[j] = im1_cols - np.count_nonzero(im1[j, :])
        tt1 = time.perf_counter()
        total_tt1 += tt1 - t5

        amr1 = np.argmax(sum_r1)
        tt2 = time.perf_counter()

        for j in range(im2_rows):
            sum_r2[j] = im2_cols - np.count_nonzero(im2[j, :])
        t6 = time.perf_counter()
        total_t5 += t6 - t5

    # 打印平均耗时
    print(f'TT3 平均耗时: {total_tt3 / runs:.6f}')
    print(f'TT5 平均耗时: {total_tt5 / runs:.6f}')
    print(f'TT1 平均耗时: {total_tt1 / runs:.6f}')
    print(f'T5 平均耗时: {total_t5 / runs:.6f}')

内容的提问来源于stack exchange,提问作者sinamcr7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 21:35:16