寻求Numpy中计算密集型数学函数的优化方案
大型数组下Numpy向量化代码加速优化请求
我正在开发一个Python脚本,用于实现某一特定数学函数:

其中索引为周期性索引,满足1 <= j <= n。该函数灵感来源于此前的技术问题,主要用途是针对大型数组x(尺寸为5000或更大)计算该数学函数。以下是当前的Python实现代码:
import numpy as np import sys, os def compute_y(x, L, v): n = len(x) y = np.zeros(n) for k in range(L+1): # Generate the indices for the window around each j, with periodic wrapping indices = np.arange(-k, k+1) # Compute the weights weights_1 = k - np.abs(indices) weights_2 = k + 1 - np.abs(indices) weights_d = np.ones(2*k+1) # For each j, take the elements in the window around j and multiply by the weights x_matrix = np.take(x, (np.arange(n)[:, None] + indices) % n, mode='wrap') exp_1 = np.exp(-np.sum(weights_1[None, :] * x_matrix, axis=1)/v) exp_2 = np.exp(-np.sum(weights_2[None, :] * x_matrix, axis=1)/v) denom = np.sum(weights_d[None, :] * x_matrix, axis=1) # Compute the weighted sum for each j and add to the total y += (exp_1 - exp_2)/denom return y # Test the function x = np.random.rand(5000) L = len(x)//2 v = 1.4 y = compute_y(x, L, v)
目前这段代码可以正常运行且已实现向量化,但处理大型数组时速度远达不到预期。我认为速度慢的核心原因是循环中重复处理索引生成、权重计算、求和及指数运算的过程。
希望获得针对该代码的加速建议,尤其是针对5000或更大尺寸数组的优化方案,重点关注借助Numpy向量化实现加速的方法。
内容的提问来源于stack exchange,提问作者sam wolfe
相关产品推荐
相关产品推荐

