LiDAR数据预处理Python代码优化:极坐标分箱计算提速需求
LiDAR数据处理代码优化方案(适配大规模数据集)
我正在预处理一个规模较小的LiDAR数据集(4000个实例),但编写的Python代码在笔记本电脑上处理每个实例需耗时9秒,总运行时长达10小时,现寻求代码优化方案以支持更大规模数据集。
数据说明
data:形状为(4000,N,3),存储每次扫描中点的XYZ笛卡尔坐标(原数据为(4000,N,5),截取前3列)data_2d:形状为(4000,N,3),存储点的极坐标,第一列为仰角theta,第二列为方位角phi,第三列为距离d(用于其他处理环节)
原代码逻辑:通过筛选theta和phi范围内的点,计算对应XYZ坐标的平均值并分箱存入new_2d_scans_3d。
原代码
import numpy as np import time new_2d_scans_3d = np.zeros((4000,80,640,3)) min_max_theta =[np.min(data_2d[0][:,0]),np.max(data_2d[0][:,0])] for i in data_2d: current_min = np.min(i[:,0]) current_max = np.max(i[:,0]) if current_min < min_max_theta [0]: min_max_theta [0] = current_min if current_max < min_max_theta [1]: min_max_theta [1] = current_max # min_max_theta = [1.3845561488601013, 2.1061203926592453] min_max_phi = [-3.1414179913970317, 3.1415925661738253] theta_scale = np.linspace(min_max_theta[0], min_max_theta[1],80) phi_scale = np.linspace(min_max_phi[0], min_max_phi[1],640) for k in range(data_2d.shape[0]): st = time.time() chosen_scan_data = data[k][:,:3] chosen_scan = data_2d[k] current_theta = theta_scale[0] current_phi = phi_scale[0] for i in range(1,640): for j in range(1,80): idx = np.where((chosen_scan[:,0] > current_theta) & (chosen_scan[:,0] < theta_scale[j]) & (chosen_scan[:,1] > current_phi) & (chosen_scan[:,1] <phi_scale[i])) if idx[0].size != 0: avg_point =np.average(chosen_scan_data[idx[0]],axis=0) new_2d_scans_3d[k][j-1][i-1] = avg_point else: new_2d_scans_3d[k][j-1][i-1] = np.array([0,0,0]) current_theta = theta_scale[j] current_phi = phi_scale[i] et = time.time() elapsed_time = et - st print('Executing Image ', k ,'/4000: ', elapsed_time, 'seconds')
示例数据生成代码
import numpy as np import math data = np.random.rand(4000,35000,3) data_2d = [] for scan in data: scan_2d = [] for point in scan: phi = math.atan2(point[1],point[0]) theta = math.acos(point[2]/math.sqrt(point[0]**2 + point[1]**2 + point[2]**2)) d =math.sqrt(point[0]**2+point[1]**2) scan_2d.append([theta,phi,d]) data_2d.append(scan_2d) data_2d = np.array(data_2d)
优化方案
核心思路:替换嵌套循环为向量化运算,Python循环是性能瓶颈,利用NumPy/SciPy的内置函数实现批量计算,能将单实例处理时间从9秒压缩到几百毫秒。
1. 简化全局theta范围计算
原代码遍历所有实例找theta极值,直接用NumPy批量计算:
min_max_theta = [np.min(data_2d[:, :, 0]), np.max(data_2d[:, :, 0])]
替代原for循环,速度提升10倍以上。
2. 用2D分箱统计替代嵌套循环
使用scipy.stats.binned_statistic_2d直接计算每个分箱的XYZ平均值,彻底去掉640*80的嵌套循环:
import numpy as np import time from scipy.stats import binned_statistic_2d # 初始化结果数组 new_2d_scans_3d = np.zeros((4000, 80, 640, 3)) # 全局范围计算 min_max_theta = [np.min(data_2d[:, :, 0]), np.max(data_2d[:, :, 0])] min_max_phi = [-np.pi, np.pi] # 直接用phi的物理范围,避免硬编码 # 生成分箱边界(注意:binned_statistic_2d需要边界数=分箱数+1) theta_bins = np.linspace(min_max_theta[0], min_max_theta[1], 80 + 1) phi_bins = np.linspace(min_max_phi[0], min_max_phi[1], 640 + 1) # 批量处理每个扫描实例 for k in range(data_2d.shape[0]): st = time.time() xyz = data[k][:, :3] theta = data_2d[k][:, 0] phi = data_2d[k][:, 1] # 对XYZ三个维度分别计算分箱平均值 for dim in range(3): stat, _, _, _ = binned_statistic_2d( phi, theta, # 输入顺序为x(phi), y(theta) values=xyz[:, dim], statistic='mean', bins=[phi_bins, theta_bins] ) # 空箱填充为[0,0,0] new_2d_scans_3d[k, :, :, dim] = np.nan_to_num(stat) et = time.time() print(f'Executing Image {k}/4000: {et - st:.2f} seconds')
3. 优化示例数据生成(可选)
原嵌套循环生成极坐标,改用向量化运算,生成时间从几分钟压缩到几秒:
import numpy as np data = np.random.rand(4000, 35000, 3) # 向量化计算极坐标 xy_sq_sum = data[..., 0]**2 + data[..., 1]**2 r = np.sqrt(xy_sq_sum + data[..., 2]**2) phi = np.arctan2(data[..., 1], data[..., 0]) theta = np.arccos(data[..., 2] / r) d = np.sqrt(xy_sq_sum) # 组合成data_2d data_2d = np.stack([theta, phi, d], axis=-1)
4. 额外性能提升建议
- Numba即时编译:如果向量化后仍有瓶颈,用
@numba.jit(nopython=True)装饰核心处理函数,可再提速2-3倍。 - 多进程并行:用
joblib或multiprocessing将4000个实例分成多批,利用CPU多核心并行处理,总耗时可按核心数比例降低。 - 分块处理:如果内存不足,将数据集分成小块分批处理,避免一次性加载全部数据。
内容的提问来源于stack exchange,提问作者Youssef Bonnaire
相关产品推荐
相关产品推荐

