Python multiprocessing加速计算返回None值问题求助
问题:Multiprocessing计算Sky View Factor返回None值导致报错
核心问题定位
你的代码中calc_SVF函数存在致命逻辑错误:使用了if __name__ == '__SVF__':判断,这完全不符合Python多进程的启动规则——只有主进程的__name__为__main__,子进程导入模块时永远不会触发这个条件,导致calc_SVF默认返回None。这就是后续SVFs[i]触发TypeError: 'NoneType' object is not subscriptable的根本原因。
另外,SkyViewFactor函数未做异常捕获,若单个点计算出错会返回None,进一步导致结果列表出现无效值。
代码修正方案
1. 修复calc_SVF的执行逻辑
移除错误的__name__判断,直接执行并行逻辑,同时改用上下文管理器简化进程池管理:
def calc_SVF(coords, max_radius, blocklength): """ 计算天空视角因子(Sky View Factor) :param coords: 数据集所有坐标点 :param max_radius: 影响SVF的最大半径 :param blocklength: 需要计算的点数量 :return: 所有点的SVF列表 """ from multiprocessing import Pool from functools import partial points = [coords[i,:] for i in range(blocklength)] # 用上下文管理器自动管理进程池的关闭/回收 with Pool() as pool: SVF_par = partial(SkyViewFactor, coords=coords, max_radius=max_radius) # 用imap_unordered分块处理,降低内存占用,chunksize根据CPU核心数调整 SVF = list(pool.imap_unordered(SVF_par, points, chunksize=1000)) print(SVF) return SVF
2. 为SkyViewFactor添加异常捕获与逻辑优化
添加异常处理避免单个点计算失败返回None,同时优化角度跨2π的边界处理:
def SkyViewFactor(point, coords, max_radius): try: betas_lin = np.linspace(0, 2*np.pi, steps_beta) dome_area = max_radius**2 * 2 * np.pi dome_p = dome(point, coords, max_radius) # 无遮挡点时直接返回SVF=1.0 if dome_p.shape[0] == 0: return 1.0 # 向量化计算遮挡参数,替代低效while循环 psi = np.arctan((dome_p[:,2] - point[2]) / dome_p[:,3]) beta_half = np.arcsin(np.sqrt(2*gridboxsize**2)/2 / dome_p[:,3]) beta_min = -beta_half + dome_p[:,4] beta_max = beta_half + dome_p[:,4] betas = np.zeros(steps_beta) for p, b_min, b_max in zip(psi, beta_min, beta_max): # 处理角度跨0/2π的边界情况 b_min = np.mod(b_min, 2*np.pi) b_max = np.mod(b_max, 2*np.pi) if b_min > b_max: mask = np.logical_or( np.logical_and(betas < p, betas_lin >= b_min), np.logical_and(betas < p, betas_lin < b_max) ) else: mask = np.logical_and(betas < p, np.logical_and(betas_lin >= b_min, betas_lin < b_max)) betas[mask] = p areas = d_area(betas, steps_beta, max_radius) SVF = np.around((dome_area - np.sum(areas))/dome_area, 3) return SVF except Exception as e: print(f"处理点{point}出错: {e}") return np.nan # 用nan标记错误点,方便后续处理
3. 修正结果处理逻辑
在reshape_SVF中,先处理SVFs中的无效值,再进行矩阵重塑:
def reshape_SVF(data,coords,julianday,lat,long,LMT,reshape,save_CSV,save_Im): [x_len, y_len] = [int(data.shape[0]/2),int(data.shape[1]/2)] blocklength = int(x_len*y_len) SVFs = calc_SVF(coords,max_radius,blocklength) SFs = calc_SF(coords,julianday,lat,long,LMT,blocklength) # 替换nan为默认值(无遮挡时的SVF=1.0) SVFs = [x if not np.isnan(x) else 1.0 for x in SVFs] SFs = [x if not np.isnan(x) else 0.0 for x in SFs] if reshape == True: SVF_matrix = np.zeros([x_len,y_len]) SF_matrix = np.zeros([x_len,y_len]) for i in range(blocklength): # 修正坐标索引计算逻辑,避免越界 x_idx = int(coords[i,0]) y_idx = int(coords[i,1]) if 0 <= x_idx < x_len and 0 <= y_idx < y_len: SVF_matrix[x_idx, y_idx] = SVFs[i] SF_matrix[x_idx, y_idx] = SFs[i] if save_CSV == True: np.savetxt("SVFmatrix.csv", SVF_matrix, delimiter=",") np.savetxt("SFmatrix.csv", SF_matrix, delimiter=",") if save_Im == True: tf.imwrite('SVF_matrix.tif', SVF_matrix, photometric='minisblack') tf.imwrite('SF_matrix.tif', SF_matrix, photometric='minisblack') return SVF_matrix, SF_matrix elif reshape == False: np.savetxt("SVFs.csv", SVFs, delimiter=",") np.savetxt("SFs.csv", SFs, delimiter=",") return SVFs, SFs
针对125万数据点的额外加速建议
- 调整chunksize:根据CPU核心数设置
chunksize(比如核心数*1000),平衡进程调度开销与内存占用。 - 使用共享内存:对于大型
coords数组,用multiprocessing.Array或numpy.sharedmem共享内存,避免子进程重复拷贝数据。 - 尝试Dask替代multiprocessing:Dask支持分块处理大型数据集,自动管理内存与进程,适合超大规模数据计算。
- GPU加速:若有GPU资源,可使用CuPy替代Numpy,将向量化计算迁移到GPU,进一步提升速度。
内容的提问来源于stack exchange,提问作者Rosalie
相关产品推荐
相关产品推荐

