You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python multiprocessing加速计算返回None值问题求助

问题:Multiprocessing计算Sky View Factor返回None值导致报错

核心问题定位

你的代码中calc_SVF函数存在致命逻辑错误:使用了if __name__ == '__SVF__':判断,这完全不符合Python多进程的启动规则——只有主进程的__name__为__main__,子进程导入模块时永远不会触发这个条件,导致calc_SVF默认返回None。这就是后续SVFs[i]触发TypeError: 'NoneType' object is not subscriptable的根本原因。

另外,SkyViewFactor函数未做异常捕获,若单个点计算出错会返回None,进一步导致结果列表出现无效值。

代码修正方案

1. 修复calc_SVF的执行逻辑

移除错误的__name__判断,直接执行并行逻辑,同时改用上下文管理器简化进程池管理:

def calc_SVF(coords, max_radius, blocklength):
    """
    计算天空视角因子(Sky View Factor)
    :param coords: 数据集所有坐标点
    :param max_radius: 影响SVF的最大半径
    :param blocklength: 需要计算的点数量
    :return: 所有点的SVF列表
    """
    from multiprocessing import Pool
    from functools import partial

    points = [coords[i,:] for i in range(blocklength)]
    # 用上下文管理器自动管理进程池的关闭/回收
    with Pool() as pool:
        SVF_par = partial(SkyViewFactor, coords=coords, max_radius=max_radius)
        # 用imap_unordered分块处理,降低内存占用,chunksize根据CPU核心数调整
        SVF = list(pool.imap_unordered(SVF_par, points, chunksize=1000))
    print(SVF)
    return SVF

2. 为SkyViewFactor添加异常捕获与逻辑优化

添加异常处理避免单个点计算失败返回None,同时优化角度跨2π的边界处理:

def SkyViewFactor(point, coords, max_radius):
    try:
        betas_lin = np.linspace(0, 2*np.pi, steps_beta)
        dome_area = max_radius**2 * 2 * np.pi
        dome_p = dome(point, coords, max_radius)
        
        # 无遮挡点时直接返回SVF=1.0
        if dome_p.shape[0] == 0:
            return 1.0

        # 向量化计算遮挡参数,替代低效while循环
        psi = np.arctan((dome_p[:,2] - point[2]) / dome_p[:,3])
        beta_half = np.arcsin(np.sqrt(2*gridboxsize**2)/2 / dome_p[:,3])
        beta_min = -beta_half + dome_p[:,4]
        beta_max = beta_half + dome_p[:,4]

        betas = np.zeros(steps_beta)
        for p, b_min, b_max in zip(psi, beta_min, beta_max):
            # 处理角度跨0/2π的边界情况
            b_min = np.mod(b_min, 2*np.pi)
            b_max = np.mod(b_max, 2*np.pi)
            if b_min > b_max:
                mask = np.logical_or(
                    np.logical_and(betas < p, betas_lin >= b_min),
                    np.logical_and(betas < p, betas_lin < b_max)
                )
            else:
                mask = np.logical_and(betas < p, np.logical_and(betas_lin >= b_min, betas_lin < b_max))
            betas[mask] = p

        areas = d_area(betas, steps_beta, max_radius)
        SVF = np.around((dome_area - np.sum(areas))/dome_area, 3)
        return SVF
    except Exception as e:
        print(f"处理点{point}出错: {e}")
        return np.nan  # 用nan标记错误点,方便后续处理

3. 修正结果处理逻辑

在reshape_SVF中,先处理SVFs中的无效值,再进行矩阵重塑:

def reshape_SVF(data,coords,julianday,lat,long,LMT,reshape,save_CSV,save_Im):
    [x_len, y_len] = [int(data.shape[0]/2),int(data.shape[1]/2)]
    blocklength = int(x_len*y_len)

    SVFs = calc_SVF(coords,max_radius,blocklength)
    SFs = calc_SF(coords,julianday,lat,long,LMT,blocklength)
    
    # 替换nan为默认值(无遮挡时的SVF=1.0)
    SVFs = [x if not np.isnan(x) else 1.0 for x in SVFs]
    SFs = [x if not np.isnan(x) else 0.0 for x in SFs]

    if reshape == True:
        SVF_matrix = np.zeros([x_len,y_len])
        SF_matrix = np.zeros([x_len,y_len])
        for i in range(blocklength):
            # 修正坐标索引计算逻辑,避免越界
            x_idx = int(coords[i,0])
            y_idx = int(coords[i,1])
            if 0 <= x_idx < x_len and 0 <= y_idx < y_len:
                SVF_matrix[x_idx, y_idx] = SVFs[i]
                SF_matrix[x_idx, y_idx] = SFs[i]
        
        if save_CSV == True:
            np.savetxt("SVFmatrix.csv", SVF_matrix, delimiter=",")
            np.savetxt("SFmatrix.csv", SF_matrix, delimiter=",")
        if save_Im == True:
            tf.imwrite('SVF_matrix.tif', SVF_matrix, photometric='minisblack')
            tf.imwrite('SF_matrix.tif', SF_matrix, photometric='minisblack')
        return SVF_matrix, SF_matrix

    elif reshape == False:
        np.savetxt("SVFs.csv", SVFs, delimiter=",")
        np.savetxt("SFs.csv", SFs, delimiter=",")
        return SVFs, SFs

针对125万数据点的额外加速建议

  1. 调整chunksize:根据CPU核心数设置chunksize(比如核心数*1000),平衡进程调度开销与内存占用。
  2. 使用共享内存:对于大型coords数组,用multiprocessing.Array或numpy.sharedmem共享内存,避免子进程重复拷贝数据。
  3. 尝试Dask替代multiprocessing:Dask支持分块处理大型数据集,自动管理内存与进程,适合超大规模数据计算。
  4. GPU加速:若有GPU资源,可使用CuPy替代Numpy,将向量化计算迁移到GPU,进一步提升速度。

内容的提问来源于stack exchange,提问作者Rosalie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 12:55:18