You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LiDAR数据预处理Python代码优化:极坐标分箱计算提速需求

LiDAR数据处理代码优化方案(适配大规模数据集)

我正在预处理一个规模较小的LiDAR数据集(4000个实例),但编写的Python代码在笔记本电脑上处理每个实例需耗时9秒,总运行时长达10小时,现寻求代码优化方案以支持更大规模数据集。

数据说明

  • data:形状为(4000,N,3),存储每次扫描中点的XYZ笛卡尔坐标(原数据为(4000,N,5),截取前3列)
  • data_2d:形状为(4000,N,3),存储点的极坐标,第一列为仰角theta,第二列为方位角phi,第三列为距离d(用于其他处理环节)

原代码逻辑:通过筛选theta和phi范围内的点,计算对应XYZ坐标的平均值并分箱存入new_2d_scans_3d。


原代码

import numpy as np
import time

new_2d_scans_3d = np.zeros((4000,80,640,3)) 

min_max_theta =[np.min(data_2d[0][:,0]),np.max(data_2d[0][:,0])]

for i in data_2d:
    current_min = np.min(i[:,0])
    current_max = np.max(i[:,0])
    if  current_min < min_max_theta [0]:
        min_max_theta [0] = current_min
    if  current_max < min_max_theta [1]:
        min_max_theta [1] = current_max

# min_max_theta = [1.3845561488601013, 2.1061203926592453]
min_max_phi = [-3.1414179913970317, 3.1415925661738253]

theta_scale = np.linspace(min_max_theta[0], min_max_theta[1],80)
phi_scale = np.linspace(min_max_phi[0], min_max_phi[1],640)

for k in range(data_2d.shape[0]):
    st = time.time()
    chosen_scan_data = data[k][:,:3]
    chosen_scan = data_2d[k]
    current_theta = theta_scale[0]
    current_phi = phi_scale[0]
    for i in range(1,640):
        for j in range(1,80):
            idx = np.where((chosen_scan[:,0] > current_theta) & (chosen_scan[:,0] < theta_scale[j]) & (chosen_scan[:,1] > current_phi) & (chosen_scan[:,1] <phi_scale[i]))
            if idx[0].size != 0:
                avg_point =np.average(chosen_scan_data[idx[0]],axis=0)
                new_2d_scans_3d[k][j-1][i-1] = avg_point
            else:
                new_2d_scans_3d[k][j-1][i-1] = np.array([0,0,0])
            current_theta = theta_scale[j]
        current_phi = phi_scale[i]
    et = time.time()
    elapsed_time = et - st
    print('Executing Image ', k ,'/4000: ', elapsed_time, 'seconds')

示例数据生成代码

import numpy as np
import math

data = np.random.rand(4000,35000,3)
data_2d = []
for scan in data:
    scan_2d = []
    for point in scan:
        phi = math.atan2(point[1],point[0])
        theta = math.acos(point[2]/math.sqrt(point[0]**2 + point[1]**2 + point[2]**2))
        d =math.sqrt(point[0]**2+point[1]**2)
        scan_2d.append([theta,phi,d])
    data_2d.append(scan_2d)
data_2d = np.array(data_2d)

优化方案

核心思路:替换嵌套循环为向量化运算,Python循环是性能瓶颈,利用NumPy/SciPy的内置函数实现批量计算,能将单实例处理时间从9秒压缩到几百毫秒。

1. 简化全局theta范围计算

原代码遍历所有实例找theta极值,直接用NumPy批量计算:

min_max_theta = [np.min(data_2d[:, :, 0]), np.max(data_2d[:, :, 0])]

替代原for循环,速度提升10倍以上。

2. 用2D分箱统计替代嵌套循环

使用scipy.stats.binned_statistic_2d直接计算每个分箱的XYZ平均值,彻底去掉640*80的嵌套循环:

import numpy as np
import time
from scipy.stats import binned_statistic_2d

# 初始化结果数组
new_2d_scans_3d = np.zeros((4000, 80, 640, 3))

# 全局范围计算
min_max_theta = [np.min(data_2d[:, :, 0]), np.max(data_2d[:, :, 0])]
min_max_phi = [-np.pi, np.pi]  # 直接用phi的物理范围,避免硬编码

# 生成分箱边界(注意:binned_statistic_2d需要边界数=分箱数+1)
theta_bins = np.linspace(min_max_theta[0], min_max_theta[1], 80 + 1)
phi_bins = np.linspace(min_max_phi[0], min_max_phi[1], 640 + 1)

# 批量处理每个扫描实例
for k in range(data_2d.shape[0]):
    st = time.time()
    xyz = data[k][:, :3]
    theta = data_2d[k][:, 0]
    phi = data_2d[k][:, 1]
    
    # 对XYZ三个维度分别计算分箱平均值
    for dim in range(3):
        stat, _, _, _ = binned_statistic_2d(
            phi, theta,  # 输入顺序为x(phi), y(theta)
            values=xyz[:, dim],
            statistic='mean',
            bins=[phi_bins, theta_bins]
        )
        # 空箱填充为[0,0,0]
        new_2d_scans_3d[k, :, :, dim] = np.nan_to_num(stat)
    
    et = time.time()
    print(f'Executing Image {k}/4000: {et - st:.2f} seconds')

3. 优化示例数据生成(可选)

原嵌套循环生成极坐标,改用向量化运算,生成时间从几分钟压缩到几秒:

import numpy as np

data = np.random.rand(4000, 35000, 3)
# 向量化计算极坐标
xy_sq_sum = data[..., 0]**2 + data[..., 1]**2
r = np.sqrt(xy_sq_sum + data[..., 2]**2)
phi = np.arctan2(data[..., 1], data[..., 0])
theta = np.arccos(data[..., 2] / r)
d = np.sqrt(xy_sq_sum)
# 组合成data_2d
data_2d = np.stack([theta, phi, d], axis=-1)

4. 额外性能提升建议

  • Numba即时编译:如果向量化后仍有瓶颈,用@numba.jit(nopython=True)装饰核心处理函数,可再提速2-3倍。
  • 多进程并行:用joblib或multiprocessing将4000个实例分成多批,利用CPU多核心并行处理,总耗时可按核心数比例降低。
  • 分块处理:如果内存不足,将数据集分成小块分批处理,避免一次性加载全部数据。

内容的提问来源于stack exchange,提问作者Youssef Bonnaire

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 22:15:39