使用Lambda函数从DataFrame城市名提取经纬度的问题解决
问题描述
我有一个包含1100行迁徙数据的DataFrame,包含出发地、目的地的城市及国家信息。需要将城市名(比如Portland, Oregon)发送到Nominatim搜索接口提取经纬度。单条查询代码可以正常运行,但用apply结合Lambda遍历DataFrame列时,出现ValueError: Expected a 1D array, got an array with shape (1100, 17)的错误。制作的可复现示例没有报错,但返回全NA值,需要解决这个问题来获取目标数据。
可正常运行的单条查询代码
import requests import urllib.parse address = 'Portland, Oregon' url = 'https://nominatim.openstreetmap.org/search/' + urllib.parse.quote(address) +'?format=json' response = requests.get(url).json() print(response[0]["lat"]) print(response[0]["lon"])
存在问题的代码
原问题代码
segment1 = 'https://nominatim.openstreetmap.org/search/' segment3 = '?format=json' df1['json_location_data'] = df1.apply(lambda x: requests.get(segment1 + urllib.parse.quote(str(df1['Origin'])) + segment3).json())
可复现示例代码
import pandas as pd locations = ['Portland, Oregon', 'Seattle, Washington','New York, New York','Texas, United States'] df = pd.DataFrame(locations, columns=['locations']) segment1 = 'https://nominatim.openstreetmap.org/search/' segment3 = '?format=json' df['json_location_data'] = df.apply(lambda x: requests.get(segment1 + urllib.parse.quote(str(df['locations'])) + segment3).json())
错误原因分析
- apply调用逻辑错误:Lambda函数里直接引用
df['locations'](或df1['Origin'])是调用整个列的所有数据,而非当前行的单个城市名。这导致每次请求都把整列城市名拼接成一个字符串发送给接口,接口无法识别,返回空结果,最终出现全NA或维度不匹配的错误。 - 未指定apply处理轴:默认情况下
apply按列(axis=0)处理,而我们需要按行(axis=1)处理每一行的单个城市名,否则会触发维度不匹配的ValueError。
修正后的代码
基础修正版本
import pandas as pd import requests import urllib.parse locations = ['Portland, Oregon', 'Seattle, Washington','New York, New York','Texas, United States'] df = pd.DataFrame(locations, columns=['locations']) segment1 = 'https://nominatim.openstreetmap.org/search/' segment3 = '?format=json' # 按行处理,用x['locations']获取当前行的城市名 df['json_location_data'] = df.apply(lambda x: requests.get(segment1 + urllib.parse.quote(str(x['locations'])) + segment3).json(), axis=1) # 从返回的JSON中提取经纬度 def get_lat_lon(json_data): if json_data: return pd.Series([json_data[0]['lat'], json_data[0]['lon']]) return pd.Series([None, None]) df[['latitude', 'longitude']] = df['json_location_data'].apply(get_lat_lon) print(df)
进阶优化(避免被Nominatim封禁)
Nominatim有请求频率限制,直接批量请求容易被封禁,建议添加请求头和延迟:
import pandas as pd import requests import urllib.parse import time locations = ['Portland, Oregon', 'Seattle, Washington','New York, New York','Texas, United States'] df = pd.DataFrame(locations, columns=['locations']) segment1 = 'https://nominatim.openstreetmap.org/search/' segment3 = '?format=json' def fetch_location_data(address): # 添加请求头模拟浏览器访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } url = segment1 + urllib.parse.quote(str(address)) + segment3 response = requests.get(url, headers=headers) # 添加1秒延迟,符合Nominatim使用规范 time.sleep(1) return response.json() # 直接对列使用apply,无需指定axis df['json_location_data'] = df['locations'].apply(fetch_location_data) # 提取经纬度 def get_lat_lon(json_data): if json_data: return pd.Series([json_data[0]['lat'], json_data[0]['lon']]) return pd.Series([None, None]) df[['latitude', 'longitude']] = df['json_location_data'].apply(get_lat_lon) print(df)
内容的提问来源于stack exchange,提问作者user2813606
相关产品推荐
相关产品推荐

