使用Python的GeoPy Geocoders处理邮编地址列分类错误的求助
问题描述
我是Python新手,尝试用GeoPy处理CSV中的账单邮编数据,提取城市、州/省、国家信息并存入对应列。目前代码可运行,但仅按逗号拆分地址字符串,导致分类错误:如邮编93049对应的Bayern被误分到城市列(应属州/省),邮编95684对应的California被误分到城市列(应属州/省)。
当前代码
import csv import pandas as pd from geopy.geocoders import Nominatim test_filename='test.csv' df=pd.read_csv(test_filename, dtype=object, index_col = 1) df.head() geolocator = Nominatim(user_agent="user_me") from geopy.extra.rate_limiter import RateLimiter geocode = RateLimiter(geolocator.geocode, min_delay_seconds=1) Address = geolocator.geocode("Billing Zip") df['Address'] = df["Billing Zip"].apply(geocode) df['Address'] = df['Address'].astype(str).str.replace('.', '') df[['City', 'County', 'State', 'Zip Code', 'Country', ' ', ' ']] = df.Address.str.split(',' , expand = True) print(df.head())
当前输出
| 账单邮编 | 地址 | 城市 | 州/省 | 国家 |
|---|---|---|---|---|
| 10006 | Manhattan, City of New York, New York, United States | New York | New York | United States |
| 93049 | Bayern, Deutschland | Bayern | Deutschland | |
| 95684 | California, United States | California | United States |
期望输出
| 账单邮编 | 地址 | 城市 | 州/省 | 国家 |
|---|---|---|---|---|
| 10006 | Manhattan, City of New York, New York, United States | New York | New York | United States |
| 93049 | Bayern, Deutschland | Bayern | Deutschland | |
| 95684 | California, United States | California | United States |
最优解决方案
不要通过拆分字符串提取地址组件,GeoPy返回的Location对象本身包含结构化的地址数据(存放在raw属性的字典中),直接从字典按类型提取字段,能彻底避免字符串拆分的格式错误。
核心思路
- 保留
Location原始对象,不转成字符串 - 从对象的
raw['address']字典中,按地址类型(城市、州/省、国家)提取对应值 - 适配不同国家的地址字段命名差异(比如部分国家用
province而非state表示州/省)
完整修改代码
import pandas as pd from geopy.geocoders import Nominatim from geopy.extra.rate_limiter import RateLimiter test_filename='test.csv' df = pd.read_csv(test_filename, dtype=object) geolocator = Nominatim(user_agent="user_me") geocode = RateLimiter(geolocator.geocode, min_delay_seconds=1) # 保留Location原始对象,不转字符串 df['Location'] = df["Billing Zip"].apply(geocode) # 定义地址组件提取函数 def parse_address(location): if not location: return pd.Series([None, None, None]) addr = location.raw.get('address', {}) # 提取城市:优先取city,无则取town,再无留空 city = addr.get('city') or addr.get('town') or None # 提取州/省:优先取state,无则取province state = addr.get('state') or addr.get('province') or None # 提取国家 country = addr.get('country') or None return pd.Series([city, state, country]) # 生成目标列 df[['City', 'State', 'Country']] = df['Location'].apply(parse_address) # 可选:生成地址字符串列(如果需要展示) df['Address'] = df['Location'].astype(str).str.replace('.', '') # 输出结果 print(df[['Billing Zip', 'Address', 'City', 'State', 'Country']].head())
方案优势
- 准确性:依赖Nominatim返回的结构化数据,完全避免字符串拆分的格式误判
- 兼容性:自动适配不同国家的地址字段命名规则
- 扩展性:后续需提取县、街道等其他字段时,只需在函数中添加对应字典键即可
内容的提问来源于stack exchange,提问作者newbiezz
相关产品推荐
相关产品推荐

