如何在Pandas中基于容差移除近似重复的XYZ坐标数据?
解决方案:基于阈值的空间邻近匹配移除重复点
取整分组的方法过于刚性,无法处理像149和151这种跨取整边界的偏移点。我们可以用KDTree空间邻近搜索实现基于自定义阈值的重复点判定,具体步骤如下:
实现思路
- 为两个DataFrame添加来源标识,方便后续区分;
- 对坐标进行缩放,将各维度的阈值统一为1,这样可以用L∞范数(切比雪夫距离)同时满足x/y/z的阈值条件;
- 利用KDTree高效查找满足阈值的邻近点,标记出所有互相匹配的重复点;
- 从合并后的DataFrame中移除这些重复点。
完整代码
import pandas as pd import numpy as np from scipy.spatial import KDTree # 初始化测试数据 df_test_1 = pd.DataFrame(np.array([[123, 449, 756.102], [406, 523, 543.089], [140, 856, 657.24], [151, 242, 124.42]]), columns=['x', 'y', 'z']) df_test_2 = pd.DataFrame(np.array([[123, 451, 756.099], [404, 521, 543.090], [139, 859, 657.23], [633, 176, 875.76]]), columns=['x', 'y', 'z']) # 添加来源标识,便于后续匹配 df_test_1['source'] = 'df1' df_test_2['source'] = 'df2' # 合并两个DataFrame df_combined = pd.concat([df_test_1, df_test_2], ignore_index=True) # 定义各维度的匹配阈值 x_threshold = 100 y_threshold = 100 z_threshold = 0.1 # 提取坐标数组并缩放,将各维度阈值统一为1 scaling_factor = np.array([x_threshold, y_threshold, z_threshold]) scaled_df1 = df_test_1[['x', 'y', 'z']].values / scaling_factor scaled_df2 = df_test_2[['x', 'y', 'z']].values / scaling_factor # 构建KDTree用于邻近搜索 tree = KDTree(scaled_df2) # 查找df1中所有在df2里满足阈值的点 df1_matches = tree.query_ball_point(scaled_df1, r=1, p=np.inf) df1_matched_idx = [i for i, matches in enumerate(df1_matches) if len(matches) > 0] df1_matched_rows = df_test_1.iloc[df1_matched_idx] # 查找df2中所有在df1里满足阈值的点 tree_rev = KDTree(scaled_df1) df2_matches = tree_rev.query_ball_point(scaled_df2, r=1, p=np.inf) df2_matched_idx = [i for i, matches in enumerate(df2_matches) if len(matches) > 0] df2_matched_rows = df_test_2.iloc[df2_matched_idx] # 合并所有需要移除的重复点 all_matched = pd.concat([df1_matched_rows, df2_matched_rows], ignore_index=True) # 处理浮点数精度问题,用保留3位小数的坐标进行匹配 for col in ['x', 'y', 'z']: df_combined[f'{col}_round'] = df_combined[col].round(3) all_matched[f'{col}_round'] = all_matched[col].round(3) # 找到合并DataFrame中需要移除的行 to_remove = df_combined.merge(all_matched, on=['x_round', 'y_round', 'z_round', 'source'], how='inner') # 移除重复点并清理辅助列 df_filtered = df_combined.drop(to_remove.index).drop(columns=['source', 'x_round', 'y_round', 'z_round']) print(df_filtered)
代码说明
- 坐标缩放:将x除以100,y除以100,z除以0.1,把原本的阈值条件转化为缩放后坐标的L∞距离≤1,KDTree可以高效完成这种邻近搜索;
- 双向匹配:同时查找df1在df2中的匹配点和df2在df1中的匹配点,确保所有互相重复的点都被标记;
- 浮点数精度处理:通过保留3位小数避免浮点数精度差异导致的匹配失败。
运行后,最终的df_filtered会保留无匹配的点:
x y z 3 151.0 242.0 124.42 7 633.0 176.0 875.76
内容的提问来源于stack exchange,提问作者Marcus K.
相关产品推荐
相关产品推荐

