You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的GeoPy Geocoders处理邮编地址列分类错误的求助

问题描述

我是Python新手,尝试用GeoPy处理CSV中的账单邮编数据,提取城市、州/省、国家信息并存入对应列。目前代码可运行,但仅按逗号拆分地址字符串,导致分类错误:如邮编93049对应的Bayern被误分到城市列(应属州/省),邮编95684对应的California被误分到城市列(应属州/省)。

当前代码

import csv
import pandas as pd
from geopy.geocoders import Nominatim

test_filename='test.csv'
df=pd.read_csv(test_filename, dtype=object, index_col = 1)
df.head()

geolocator = Nominatim(user_agent="user_me")

from geopy.extra.rate_limiter import RateLimiter
geocode = RateLimiter(geolocator.geocode, min_delay_seconds=1)
Address = geolocator.geocode("Billing Zip")
df['Address'] = df["Billing Zip"].apply(geocode)
df['Address'] = df['Address'].astype(str).str.replace('.', '')

df[['City', 'County', 'State', 'Zip Code', 'Country', ' ', ' ']] = df.Address.str.split(',' , expand = True)
print(df.head())

当前输出

账单邮编地址城市州/省国家
10006Manhattan, City of New York, New York, United StatesNew YorkNew YorkUnited States
93049Bayern, DeutschlandBayernDeutschland
95684California, United StatesCaliforniaUnited States

期望输出

账单邮编地址城市州/省国家
10006Manhattan, City of New York, New York, United StatesNew YorkNew YorkUnited States
93049Bayern, DeutschlandBayernDeutschland
95684California, United StatesCaliforniaUnited States

最优解决方案

不要通过拆分字符串提取地址组件,GeoPy返回的Location对象本身包含结构化的地址数据(存放在raw属性的字典中),直接从字典按类型提取字段,能彻底避免字符串拆分的格式错误。

核心思路

  1. 保留Location原始对象,不转成字符串
  2. 从对象的raw['address']字典中,按地址类型(城市、州/省、国家)提取对应值
  3. 适配不同国家的地址字段命名差异(比如部分国家用province而非state表示州/省)

完整修改代码

import pandas as pd
from geopy.geocoders import Nominatim
from geopy.extra.rate_limiter import RateLimiter

test_filename='test.csv'
df = pd.read_csv(test_filename, dtype=object)

geolocator = Nominatim(user_agent="user_me")
geocode = RateLimiter(geolocator.geocode, min_delay_seconds=1)

# 保留Location原始对象,不转字符串
df['Location'] = df["Billing Zip"].apply(geocode)

# 定义地址组件提取函数
def parse_address(location):
    if not location:
        return pd.Series([None, None, None])
    addr = location.raw.get('address', {})
    # 提取城市:优先取city,无则取town,再无留空
    city = addr.get('city') or addr.get('town') or None
    # 提取州/省:优先取state,无则取province
    state = addr.get('state') or addr.get('province') or None
    # 提取国家
    country = addr.get('country') or None
    return pd.Series([city, state, country])

# 生成目标列
df[['City', 'State', 'Country']] = df['Location'].apply(parse_address)

# 可选:生成地址字符串列(如果需要展示)
df['Address'] = df['Location'].astype(str).str.replace('.', '')

# 输出结果
print(df[['Billing Zip', 'Address', 'City', 'State', 'Country']].head())

方案优势

  • 准确性:依赖Nominatim返回的结构化数据,完全避免字符串拆分的格式误判
  • 兼容性:自动适配不同国家的地址字段命名规则
  • 扩展性:后续需提取县、街道等其他字段时,只需在函数中添加对应字典键即可

内容的提问来源于stack exchange,提问作者newbiezz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 11:45:40