You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中基于容差移除近似重复的XYZ坐标数据?

解决方案:基于阈值的空间邻近匹配移除重复点

取整分组的方法过于刚性,无法处理像149和151这种跨取整边界的偏移点。我们可以用KDTree空间邻近搜索实现基于自定义阈值的重复点判定,具体步骤如下:

实现思路

  1. 为两个DataFrame添加来源标识,方便后续区分;
  2. 对坐标进行缩放,将各维度的阈值统一为1,这样可以用L∞范数(切比雪夫距离)同时满足x/y/z的阈值条件;
  3. 利用KDTree高效查找满足阈值的邻近点,标记出所有互相匹配的重复点;
  4. 从合并后的DataFrame中移除这些重复点。

完整代码

import pandas as pd
import numpy as np
from scipy.spatial import KDTree

# 初始化测试数据
df_test_1 = pd.DataFrame(np.array([[123, 449, 756.102], [406, 523, 543.089], [140, 856, 657.24], [151, 242, 124.42]]), columns=['x', 'y', 'z'])
df_test_2 = pd.DataFrame(np.array([[123, 451, 756.099], [404, 521, 543.090], [139, 859, 657.23], [633, 176, 875.76]]), columns=['x', 'y', 'z'])

# 添加来源标识,便于后续匹配
df_test_1['source'] = 'df1'
df_test_2['source'] = 'df2'

# 合并两个DataFrame
df_combined = pd.concat([df_test_1, df_test_2], ignore_index=True)

# 定义各维度的匹配阈值
x_threshold = 100
y_threshold = 100
z_threshold = 0.1

# 提取坐标数组并缩放,将各维度阈值统一为1
scaling_factor = np.array([x_threshold, y_threshold, z_threshold])
scaled_df1 = df_test_1[['x', 'y', 'z']].values / scaling_factor
scaled_df2 = df_test_2[['x', 'y', 'z']].values / scaling_factor

# 构建KDTree用于邻近搜索
tree = KDTree(scaled_df2)

# 查找df1中所有在df2里满足阈值的点
df1_matches = tree.query_ball_point(scaled_df1, r=1, p=np.inf)
df1_matched_idx = [i for i, matches in enumerate(df1_matches) if len(matches) > 0]
df1_matched_rows = df_test_1.iloc[df1_matched_idx]

# 查找df2中所有在df1里满足阈值的点
tree_rev = KDTree(scaled_df1)
df2_matches = tree_rev.query_ball_point(scaled_df2, r=1, p=np.inf)
df2_matched_idx = [i for i, matches in enumerate(df2_matches) if len(matches) > 0]
df2_matched_rows = df_test_2.iloc[df2_matched_idx]

# 合并所有需要移除的重复点
all_matched = pd.concat([df1_matched_rows, df2_matched_rows], ignore_index=True)

# 处理浮点数精度问题,用保留3位小数的坐标进行匹配
for col in ['x', 'y', 'z']:
    df_combined[f'{col}_round'] = df_combined[col].round(3)
    all_matched[f'{col}_round'] = all_matched[col].round(3)

# 找到合并DataFrame中需要移除的行
to_remove = df_combined.merge(all_matched, on=['x_round', 'y_round', 'z_round', 'source'], how='inner')

# 移除重复点并清理辅助列
df_filtered = df_combined.drop(to_remove.index).drop(columns=['source', 'x_round', 'y_round', 'z_round'])

print(df_filtered)

代码说明

  • 坐标缩放:将x除以100,y除以100,z除以0.1,把原本的阈值条件转化为缩放后坐标的L∞距离≤1,KDTree可以高效完成这种邻近搜索;
  • 双向匹配:同时查找df1在df2中的匹配点和df2在df1中的匹配点,确保所有互相重复的点都被标记;
  • 浮点数精度处理:通过保留3位小数避免浮点数精度差异导致的匹配失败。

运行后,最终的df_filtered会保留无匹配的点:

x      y       z
3  151.0  242.0  124.42
7  633.0  176.0  875.76

内容的提问来源于stack exchange,提问作者Marcus K.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 03:54:55