如何基于DataFrame经纬度获取国家名并提取欧盟GROW传感器数据
问题修复说明
你的代码存在7处核心错误,会直接导致运行失败、结果不准:
- 欧盟国家集合定义错误:将
Greece(希腊)和Hungary(匈牙利)误写为同一个字符串Greece Hungary,会导致这两个国家的记录全部匹配失败 - 语法缩进错误:
for循环下的所有执行代码没有缩进,Python运行时会直接报语法错 - 空值处理无效:
locations.dropna(axis=0, how='any')没有将结果赋值回原变量,这行代码执行后不会对原数据产生任何修改 - 坐标传参顺序错误:
geolocator.reverse要求传入坐标的顺序为(纬度, 经度),原代码传入的是(经度, 纬度),会导致所有坐标定位结果完全错误 - 坐标过滤逻辑缺陷:GROW传感器数据集普遍存在经纬度字段值互换的问题,仅做数值范围过滤会保留大量坐标错误的记录,同时漏掉有效数据
- 反向查询无容错:没有处理网络超时、坐标无匹配结果的异常情况,遇到异常时代码会直接崩溃;且
user_agent设置为http极易被Nominatim服务封禁 - 结果存储逻辑错误:每次匹配到欧盟国家的记录时,都会覆盖
countries变量的原有值,最终只能得到最后一条匹配的记录,且没有关联原始传感器的其他字段,无法输出完整的传感器数据集
修正后代码
import pandas as pd import time from geopy.geocoders import Nominatim # 修正欧盟国家集合,补全希腊、匈牙利 EU = { 'Austria', 'Belgium', 'Bulgaria', 'Croatia', 'Cyprus', 'Czechia', 'Denmark', 'Estonia', 'Finland', 'France', 'Germany', 'Greece', 'Hungary', 'Ireland', 'Italy', 'Latvia', 'Lithuania', 'Luxembourg', 'Malta', 'Netherlands', 'Poland', 'Portugal', 'Romania', 'Slovakia', 'Slovenia', 'Spain', 'Sweden' } # 初始化地理编码器,设置合规的user_agent geolocator = Nominatim(user_agent="grow_sensor_eu_filter", timeout=10) # 读取全量数据,保留所有传感器字段 locations = pd.read_csv('GrowLocations.csv') # 去重、去空值 locations = locations.drop_duplicates(keep='first') locations = locations.dropna(axis=0, how='any', subset=['Latitude', 'Longitude']) # 修正经纬度互换问题:纬度值必然在[-90,90]区间,经度值必然在[-180,180]区间 def correct_coord(row): lat, lon = row['Latitude'], row['Longitude'] if not (-90 <= lat <= 90) and (-90 <= lon <=90) and (-180 <= lat <=180): return lon, lat return lat, lon locations[['Latitude', 'Longitude']] = locations.apply(correct_coord, axis=1, result_type='expand') # 过滤合法经纬度范围的记录 locations = locations[(locations['Latitude'].between(-90, 90)) & (locations['Longitude'].between(-180, 180))] # 存储符合条件的欧盟境内传感器记录 eu_sensors = [] for idx, row in locations.iterrows(): lat = row['Latitude'] lon = row['Longitude'] try: location = geolocator.reverse((lat, lon), language='en') # 无匹配结果直接跳过 if not location: time.sleep(1) continue country = location.raw.get('address', {}).get('country', '') if country in EU: record = row.to_dict() record['country'] = country eu_sensors.append(record) except Exception as e: print(f"坐标({lat},{lon})查询失败: {str(e)}") # 遵守Nominatim免费接口调用频率限制 time.sleep(1) # 转换为DataFrame输出 eu_sensors_df = pd.DataFrame(eu_sensors) # 保存结果 eu_sensors_df.to_csv('EU_GROW_sensors.csv', index=False) print(f"共提取到欧盟境内GROW传感器记录{len(eu_sensors_df)}条,已保存至EU_GROW_sensors.csv")
注意:如果你的数据集记录量超过1000条,不建议使用Nominatim免费接口做批量查询,会触发访问频率限制,可替换为本地部署的离线地理编码库(如GeoPandas结合Natural Earth的国家边界数据)做空间匹配,效率会提升数百倍。
内容的提问来源于stack exchange,提问作者Aigerim
相关产品推荐
相关产品推荐

