You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于DataFrame经纬度获取国家名并提取欧盟GROW传感器数据

问题修复说明

你的代码存在7处核心错误,会直接导致运行失败、结果不准:

  • 欧盟国家集合定义错误:将Greece(希腊)和Hungary(匈牙利)误写为同一个字符串Greece Hungary,会导致这两个国家的记录全部匹配失败
  • 语法缩进错误:for循环下的所有执行代码没有缩进,Python运行时会直接报语法错
  • 空值处理无效:locations.dropna(axis=0, how='any')没有将结果赋值回原变量,这行代码执行后不会对原数据产生任何修改
  • 坐标传参顺序错误:geolocator.reverse要求传入坐标的顺序为(纬度, 经度),原代码传入的是(经度, 纬度),会导致所有坐标定位结果完全错误
  • 坐标过滤逻辑缺陷:GROW传感器数据集普遍存在经纬度字段值互换的问题,仅做数值范围过滤会保留大量坐标错误的记录,同时漏掉有效数据
  • 反向查询无容错:没有处理网络超时、坐标无匹配结果的异常情况,遇到异常时代码会直接崩溃;且user_agent设置为http极易被Nominatim服务封禁
  • 结果存储逻辑错误:每次匹配到欧盟国家的记录时,都会覆盖countries变量的原有值,最终只能得到最后一条匹配的记录,且没有关联原始传感器的其他字段,无法输出完整的传感器数据集
修正后代码
import pandas as pd
import time
from geopy.geocoders import Nominatim
# 修正欧盟国家集合,补全希腊、匈牙利
EU = {
    'Austria', 'Belgium', 'Bulgaria', 'Croatia', 'Cyprus', 'Czechia', 'Denmark',
    'Estonia', 'Finland', 'France', 'Germany', 'Greece', 'Hungary', 'Ireland', 'Italy',
    'Latvia', 'Lithuania', 'Luxembourg', 'Malta', 'Netherlands', 'Poland', 'Portugal',
    'Romania', 'Slovakia', 'Slovenia', 'Spain', 'Sweden'
}
# 初始化地理编码器,设置合规的user_agent
geolocator = Nominatim(user_agent="grow_sensor_eu_filter", timeout=10)
# 读取全量数据,保留所有传感器字段
locations = pd.read_csv('GrowLocations.csv')
# 去重、去空值
locations = locations.drop_duplicates(keep='first')
locations = locations.dropna(axis=0, how='any', subset=['Latitude', 'Longitude'])
# 修正经纬度互换问题:纬度值必然在[-90,90]区间,经度值必然在[-180,180]区间
def correct_coord(row):
    lat, lon = row['Latitude'], row['Longitude']
    if not (-90 <= lat <= 90) and (-90 <= lon <=90) and (-180 <= lat <=180):
        return lon, lat
    return lat, lon
locations[['Latitude', 'Longitude']] = locations.apply(correct_coord, axis=1, result_type='expand')
# 过滤合法经纬度范围的记录
locations = locations[(locations['Latitude'].between(-90, 90)) & (locations['Longitude'].between(-180, 180))]
# 存储符合条件的欧盟境内传感器记录
eu_sensors = []
for idx, row in locations.iterrows():
    lat = row['Latitude']
    lon = row['Longitude']
    try:
        location = geolocator.reverse((lat, lon), language='en')
        # 无匹配结果直接跳过
        if not location:
            time.sleep(1)
            continue
        country = location.raw.get('address', {}).get('country', '')
        if country in EU:
            record = row.to_dict()
            record['country'] = country
            eu_sensors.append(record)
    except Exception as e:
        print(f"坐标({lat},{lon})查询失败: {str(e)}")
    # 遵守Nominatim免费接口调用频率限制
    time.sleep(1)
# 转换为DataFrame输出
eu_sensors_df = pd.DataFrame(eu_sensors)
# 保存结果
eu_sensors_df.to_csv('EU_GROW_sensors.csv', index=False)
print(f"共提取到欧盟境内GROW传感器记录{len(eu_sensors_df)}条,已保存至EU_GROW_sensors.csv")

注意:如果你的数据集记录量超过1000条,不建议使用Nominatim免费接口做批量查询,会触发访问频率限制,可替换为本地部署的离线地理编码库(如GeoPandas结合Natural Earth的国家边界数据)做空间匹配,效率会提升数百倍。

内容的提问来源于stack exchange,提问作者Aigerim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 16:18:25