Jupyter Notebook多进程Pool中metpy函数未定义问题排查
问题:多进程Pool调用MetPy函数时出现NameError
我在Jupyter Notebook中使用MetPy包的dewpoint_from_relative_humidity等函数定义了calc函数,但在多进程Pool中调用该函数时,出现以下错误:
NameError: name 'dewpoint_from_relative_humidity' is not defined
明明已提前导入该函数,怀疑问题和多进程使用或函数定义方式有关,相关代码如下:
import numpy as np import xarray as xr from time import time import matplotlib.pyplot as plt from scipy.interpolate import interp1d from metpy.units import units from metpy.calc import cape_cin, dewpoint_from_relative_humidity, parcel_profile, most_unstable_cape_cin, mixed_layer_cape_cin from metpy.calc import dewpoint_from_relative_humidity from multiprocessing import Pool from pytictoc import TicToc # conda install pytictoc -c ecf data1 = xr.open_dataset(r'new.nc') data = data1.sel(latitude = slice (50,45), longitude = slice(-110,-108)) data.r.values = data.r.values/100 data.t.values = data.t.values-273 data.t.attrs['units'] = 'degree_Celsius' data.r.attrs['units'] = 'dimensionless' data.level.attrs['units'] = 'hectopascal' ## reversing pressure levels descending data = data.isel(level=slice(None, None, -1)) p = data.level #Defining the calc function using the metpy functions: dewpointy_f_R_H and Most_unstabe_cape_cin def calc(idata): p1 = idata[0] t1 = idata[1] rh1 = idata[2] td1 = dewpoint_from_relative_humidity(t1,rh1) cape = most_unstable_cape_cin(p1, t1, td1) #print(cape) return cape SP_emp = data.drop_vars(['t']) SP_emp['cape'] = SP_emp['r']*0 SP_emp = SP_emp.drop_vars(['r']) #Dropping 'level' dimension SP_emp = SP_emp.mean(['level']) from multiprocess import Pool from pytictoc import TicToc if __name__ == '__main__': pool = Pool(processes = 8) #num_processors = 4 mucape_res = np.zeros((data.t.time.size, data.t.latitude.size * data.t.longitude.size)) # time * lat * lon print(mucape_res.shape) for lat in data.latitude.values: for lon in data.longitude.values: for tim in data.time.values: t = TicToc() t.tic() Temp = data.t.sel(time =tim, latitude=lat, longitude = lon) #print(Temp) RH = data.r.sel(time =tim, latitude=lat, longitude = lon) #print(RH) #print(TD) sets = p,Temp,RH out = pool.map(calc,sets) #out = calc(set) cape_mag = out.magnitude SP_emp.cape.loc[dict(time = tim,longitude = lon, latitude = lat)] = cape_mag t.toc() #pool.close()
问题原因
- 多进程子进程的独立环境:
Pool启动子进程时,会重新加载当前模块,但主进程中导入的MetPy函数不会自动传递到子进程,子进程无法找到dewpoint_from_relative_humidity的引用。 - 代码缩进错误:
if __name__ == '__main__':块内的代码缩进混乱,后续循环逻辑不在该块内,Jupyter环境下会导致多进程启动逻辑异常。 pool.map参数传递错误:sets = p,Temp,RH是三个独立数组,pool.map会把每个数组单独传给calc,但calc需要的是包含三个参数的单元素,导致参数不匹配(不过当前优先解决NameError)。
解决办法
1. 确保子进程能获取依赖函数
将MetPy的函数导入放到calc内部,或者确保子进程加载模块时能重新导入这些函数:
def calc(input_tuple): # 子进程内部重新导入依赖函数,确保能访问到 from metpy.calc import dewpoint_from_relative_humidity, most_unstable_cape_cin p1, t1, rh1 = input_tuple td1 = dewpoint_from_relative_humidity(t1,rh1) cape = most_unstable_cape_cin(p1, t1, td1) return cape
2. 修正多进程代码的缩进
所有多进程相关逻辑(创建Pool、循环、任务处理)必须放到if __name__ == '__main__':块内,避免Jupyter环境下子进程重复执行笔记本代码:
if __name__ == '__main__': pool = Pool(processes = 8) mucape_res = np.zeros((data.t.time.size, data.t.latitude.size * data.t.longitude.size)) print(mucape_res.shape) # 后续循环、任务处理代码全部缩进至此块内
3. 修正pool.map的参数传递
把每个时间/经纬度点的参数打包成元组,组成任务列表批量传入,提高多进程效率:
if __name__ == '__main__': pool = Pool(processes = 8) mucape_res = np.zeros((data.t.time.size, data.t.latitude.size * data.t.longitude.size)) print(mucape_res.shape) # 批量打包所有计算任务 tasks = [] for lat in data.latitude.values: for lon in data.longitude.values: for tim in data.time.values: Temp = data.t.sel(time=tim, latitude=lat, longitude=lon) RH = data.r.sel(time=tim, latitude=lat, longitude=lon) tasks.append((p, Temp, RH)) # 批量执行任务 results = pool.map(calc, tasks) # 将结果赋值回SP_emp idx = 0 for lat in data.latitude.values: for lon in data.longitude.values: for tim in data.time.values: cape_mag = results[idx].magnitude SP_emp.cape.loc[dict(time=tim, longitude=lon, latitude=lat)] = cape_mag idx += 1 pool.close() pool.join()
额外说明
- Jupyter Notebook中使用多进程必须严格遵循
if __name__ == '__main__':的规范,否则会出现子进程重复执行整个笔记本的问题。 - 函数内部导入依赖虽然有轻微开销,但能彻底解决子进程的NameError问题。
- 批量打包任务再处理比嵌套循环逐个调用效率更高,能充分发挥多进程的优势。
内容的提问来源于stack exchange,提问作者piyush
相关产品推荐
相关产品推荐

