如何删除非精确/相似重复数据?以经纬度差值≤0.1场景为例
处理非精确经纬度重复数据的方案
当然可以实现这个需求,核心思路是先将经纬度转换为数值类型,再通过聚类或分组识别出经纬度差值在0.1范围内的记录,最后保留每组中的一条即可。以下是针对你提供的示例数据的具体实现:
步骤1:合并数据并转换数据类型
首先把需要处理的数据集合并,然后将Longitude和Latitude从字符串转为浮点型,方便后续计算:
import pandas as pd import numpy as np # 你的原始数据 site_df_1 = pd.DataFrame(np.array([["12345", "Wrexham Cwtch", "52.10", "-2.06"], ["12354", "Horse & Hound", "52.21", "-1.95"], ["12435", "Round Of Gras Badsey", "52.33", "-1.99"]]), columns=['Site Number', 'Site Name', 'Longitude', 'Latitude']) site_df_2 = pd.DataFrame(np.array([["52938", "Valkyrie Café Bar", "53.22", "-3.00"], ["12435", "Round Of Badsey", "52.33", "-1.99"], ["12345", "Cwtch", "52.11", "-2.00"]]), columns=['Site Number', 'Site Name', 'Longitude', 'Latitude']) # 合并两个数据集 combined_df = pd.concat([site_df_1, site_df_2], ignore_index=True) # 将经纬度转为浮点型 combined_df['Longitude'] = combined_df['Longitude'].astype(float) combined_df['Latitude'] = combined_df['Latitude'].astype(float)
步骤2:识别经纬度相近的重复组
这里提供两种实用方法:
方法1:基于分箱分组(简单高效)
把经纬度按0.1的间隔分箱,落在同一个箱子里的记录视为相似重复项:
# 按0.1间隔对经纬度分箱 combined_df['lon_bin'] = np.floor(combined_df['Longitude'] / 0.1) * 0.1 combined_df['lat_bin'] = np.floor(combined_df['Latitude'] / 0.1) * 0.1 # 按分箱列分组,保留每组第一条记录 deduped_df = combined_df.groupby(['lon_bin', 'lat_bin']).first().reset_index(drop=True)
方法2:基于距离计算(更精准)
如果需要严格判定“经纬度绝对差都≤0.1”,可以用遍历方式识别:
# 标记需要保留的记录 keep_indices = [] for idx, row in combined_df.iterrows(): # 检查当前记录是否和已保留的记录经纬度差值都≤0.1 is_duplicate = False for kept_idx in keep_indices: kept_row = combined_df.loc[kept_idx] lon_diff = abs(row['Longitude'] - kept_row['Longitude']) lat_diff = abs(row['Latitude'] - kept_row['Latitude']) if lon_diff <= 0.1 and lat_diff <= 0.1: is_duplicate = True break if not is_duplicate: keep_indices.append(idx) # 获取去重后的数据集 deduped_df = combined_df.loc[keep_indices]
步骤3:调整与查看结果
你可以根据需求调整每组保留的记录——比如用last()代替first(),或者根据Site Number等字段优先级筛选。针对你的示例数据,最终会保留4条唯一记录(两组相近经纬度的重复项各留一条)。
内容的提问来源于stack exchange,提问作者Michael Amos
相关产品推荐
相关产品推荐

