You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何删除非精确/相似重复数据?以经纬度差值≤0.1场景为例

处理非精确经纬度重复数据的方案

当然可以实现这个需求,核心思路是先将经纬度转换为数值类型,再通过聚类或分组识别出经纬度差值在0.1范围内的记录,最后保留每组中的一条即可。以下是针对你提供的示例数据的具体实现:

步骤1:合并数据并转换数据类型

首先把需要处理的数据集合并,然后将Longitude和Latitude从字符串转为浮点型,方便后续计算:

import pandas as pd
import numpy as np

# 你的原始数据
site_df_1 = pd.DataFrame(np.array([["12345", "Wrexham Cwtch", "52.10", "-2.06"], ["12354", "Horse & Hound", "52.21", "-1.95"], ["12435", "Round Of Gras Badsey", "52.33", "-1.99"]]),
                   columns=['Site Number', 'Site Name', 'Longitude', 'Latitude'])

site_df_2 = pd.DataFrame(np.array([["52938", "Valkyrie Café Bar", "53.22", "-3.00"], ["12435", "Round Of Badsey", "52.33", "-1.99"], ["12345", "Cwtch", "52.11", "-2.00"]]),
                   columns=['Site Number', 'Site Name', 'Longitude', 'Latitude'])

# 合并两个数据集
combined_df = pd.concat([site_df_1, site_df_2], ignore_index=True)

# 将经纬度转为浮点型
combined_df['Longitude'] = combined_df['Longitude'].astype(float)
combined_df['Latitude'] = combined_df['Latitude'].astype(float)

步骤2:识别经纬度相近的重复组

这里提供两种实用方法:

方法1:基于分箱分组(简单高效)

把经纬度按0.1的间隔分箱,落在同一个箱子里的记录视为相似重复项:

# 按0.1间隔对经纬度分箱
combined_df['lon_bin'] = np.floor(combined_df['Longitude'] / 0.1) * 0.1
combined_df['lat_bin'] = np.floor(combined_df['Latitude'] / 0.1) * 0.1

# 按分箱列分组,保留每组第一条记录
deduped_df = combined_df.groupby(['lon_bin', 'lat_bin']).first().reset_index(drop=True)

方法2:基于距离计算(更精准)

如果需要严格判定“经纬度绝对差都≤0.1”,可以用遍历方式识别:

# 标记需要保留的记录
keep_indices = []

for idx, row in combined_df.iterrows():
    # 检查当前记录是否和已保留的记录经纬度差值都≤0.1
    is_duplicate = False
    for kept_idx in keep_indices:
        kept_row = combined_df.loc[kept_idx]
        lon_diff = abs(row['Longitude'] - kept_row['Longitude'])
        lat_diff = abs(row['Latitude'] - kept_row['Latitude'])
        if lon_diff <= 0.1 and lat_diff <= 0.1:
            is_duplicate = True
            break
    if not is_duplicate:
        keep_indices.append(idx)

# 获取去重后的数据集
deduped_df = combined_df.loc[keep_indices]

步骤3:调整与查看结果

你可以根据需求调整每组保留的记录——比如用last()代替first(),或者根据Site Number等字段优先级筛选。针对你的示例数据,最终会保留4条唯一记录(两组相近经纬度的重复项各留一条)。

内容的提问来源于stack exchange,提问作者Michael Amos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 01:35:19