如何基于经纬度条件将字典中的城市数据合并至Pandas DataFrame
实现方案
当然可以基于经纬度的近似匹配,把城市名合并到你的DataFrame中,下面是两种实用的方法:
方法一:四舍五入后精确合并(高效简洁)
先将城市坐标字典转换为DataFrame,再对经纬度做四舍五入处理,最后通过四舍五入后的经纬度字段合并:
import pandas as pd # 你的城市坐标字典 city_coords = { 'Rosario': [-60.63932, -32.946819], 'Concordia': [-74.448212, 40.31094], 'Avellaneda': [-58.367439, -34.660179], 'Corrientes': [-58.834099, -27.4806], 'Caballito': [-58.44104, -34.622639], 'Buenos Aires': [-78.497498, -9.12417], 'Paraná': [-60.5238, -31.73197], 'Santa Fé': [-78.14917, 8.65194], 'San Carlos de Bariloche': [-71.30822, -41.145569], 'Mendoza': [-68.827171, -32.890839] } # 将字典转为DataFrame,整理列名 city_df = pd.DataFrame.from_dict(city_coords, orient='index', columns=['lon', 'lat']) city_df = city_df.reset_index().rename(columns={'index': 'city'}) # 假设你的原始DataFrame名为df # 对经纬度四舍五入到4位小数(匹配你的示例精度) city_df['lon_round'] = city_df['lon'].round(4) city_df['lat_round'] = city_df['lat'].round(4) df['lon_round'] = df['lon'].round(4) df['lat_round'] = df['lat'].round(4) # 合并数据 merged_df = df.merge(city_df[['city', 'lon_round', 'lat_round']], on=['lon_round', 'lat_round'], how='left') # 删除临时的四舍五入列 merged_df = merged_df.drop(['lon_round', 'lat_round'], axis=1)
方法二:基于容差的近似匹配(灵活可控)
如果需要更灵活的精度控制,可以用numpy.isclose设置绝对容差,逐行匹配城市:
import pandas as pd import numpy as np # 你的城市坐标字典和原始DataFrame(同上) def match_city(row): for city, (target_lon, target_lat) in city_coords.items(): # 设置容差为0.0001,可根据实际情况调整 if np.isclose(row['lon'], target_lon, atol=1e-4) and np.isclose(row['lat'], target_lat, atol=1e-4): return city return None # 添加城市列 df['city'] = df.apply(match_city, axis=1)
说明
- 方法一适合数据量较大的场景,合并效率更高;
- 方法二可以通过调整
atol参数(绝对容差)来适配不同的近似精度需求; - 因为你提到DataFrame仅包含字典中存在的坐标,所以两种方法都能得到完整的城市匹配结果。
内容的提问来源于stack exchange,提问作者Tomás Jullier
相关产品推荐
相关产品推荐

