为何将向量存为类属性时NumPy向量化运算性能下降?
问题描述
我编写了一个辅助类,用于在网格上评估参数化函数。由于网格不会随参数变化,我选择将其一次性创建为类属性。但我发现,将网格作为类属性时,相比作为全局变量会出现显著的性能下降。更有趣的是,2D网格存在的性能差异在1D网格中似乎消失了。
以下是展示该问题的最小可复现代码:
import numpy as np import time grid_1d_size = 25000000 grid_2d_size = 5000 x_min, x_max = 0, np.pi / 2 y_min, y_max = -np.pi / 4, np.pi / 4 # Grid evaluation (2D) with grid as class attribute class GridEvaluation2DWithAttribute: def __init__(self): self.x_2d_values = np.linspace(x_min, x_max, grid_2d_size, dtype=np.float128) self.y_2d_values = np.linspace(y_min, y_max, grid_2d_size, dtype=np.float128) def grid_evaluate(self): cost_values = np.cos(self.x_2d_values[:, None] * self.y_2d_values[None, :]) return cost_values grid_eval_2d_attribute = GridEvaluation2DWithAttribute() initial_time = time.process_time() grid_eval_2d_attribute.grid_evaluate() final_time = time.process_time() print(f"2d grid, with grid as class attribute: {final_time - initial_time} seconds") # Grid evaluation (1D) with grid as global variable x_2d_values = np.linspace(x_min, x_max, grid_2d_size) y_2d_values = np.linspace(y_min, y_max, grid_2d_size) class GridEvaluation2DWithGlobal: def __init__(self): pass def grid_evaluate(self): cost_values = np.cos(x_2d_values[:, None] * y_2d_values[None, :]) return cost_values grid_eval_2d_global = GridEvaluation2DWithGlobal() initial_time = time.process_time() grid_eval_2d_global.grid_evaluate() final_time = time.process_time() print(f"2d grid, with grid as global variable: {final_time - initial_time} seconds") # Grid evaluation (1D) with grid as class attribute class GridEvaluation1DWithAttribute: def __init__(self): self.x_1d_values = np.linspace(x_min, x_max, grid_1d_size, dtype=np.float128) def grid_evaluate(self): cost_values = np.cos(self.x_1d_values) return cost_values grid_eval_1d_attribute = GridEvaluation1DWithAttribute() initial_time = time.process_time() grid_eval_1d_attribute.grid_evaluate() final_time = time.process_time() print(f"1d grid, with grid as class attribute: {final_time - initial_time} seconds") # Grid evaluation (1D) with grid as global variable x_1d_values = np.linspace(x_min, x_max, grid_1d_size, dtype=np.float128) class GridEvaluation1DWithGlobal: def __init__(self): pass def grid_evaluate(self): cost_values = np.cos(x_1d_values) return cost_values grid_eval_1d_global = GridEvaluation1DWithGlobal() initial_time = time.process_time() grid_eval_1d_global.grid_evaluate() final_time = time.process_time() print(f"1d grid, with grid as global variable: {final_time - initial_time} seconds")
运行输出如下:
2d grid, with grid as class attribute: 0.8012442529999999 seconds 2d grid, with grid as global variable: 0.20206781899999982 seconds 1d grid, with grid as class attribute: 2.0631387639999996 seconds 1d grid, with grid as global variable: 2.136266148 seconds
原本认为将网格从类属性改为全局变量对性能无影响,但实际却产生了显著差异,该如何解释这种性能差异?
性能差异原因分析
核心原因是数据类型不一致,而非类属性/全局变量的存储方式:
- 2D场景中,类属性的网格明确指定了
dtype=np.float128,但全局变量的网格未指定 dtype,默认使用float64。float128的运算复杂度远高于float64,尤其是在2D广播乘法+余弦运算这种大规模数值计算场景下,性能差异会被大幅放大——这就是两者耗时差4倍的根本原因。 - 1D场景中,类属性和全局变量的网格都明确指定了
dtype=np.float128,运算的底层数据类型完全一致,因此性能几乎没有差异。甚至类属性版本略快,是因为实例属性的查找路径比全局变量更短(全局变量需要遍历全局命名空间查找,开销略高)。
你可以做个验证:把2D全局变量的网格定义改成和类属性一致的dtype=np.float128,再运行代码,两者的耗时就会基本持平。
另外补充:2D的广播运算对数据类型的敏感度更高,因为涉及到大量的矩阵级计算,float128的计算量是float64的数倍;而1D的单维度余弦运算,数据类型一致时,属性/变量的查找开销在总耗时中占比极低,所以体现不出明显差异。
内容的提问来源于stack exchange,提问作者IchKenneDeinenNamen
相关产品推荐
相关产品推荐

